Fetching the paper…
Reading the bibliography…
Classifier-free guidance (CFG) is a key technique for improving conditional generation in diffusion models, enabling more accurate control while enhancing sample quality.
Animating rotation with quaternion curves
Shoemake, K · 1985
Earlier work this paper cites.
Linear and convex aggregation of density estimators
Rigollet, P. and Tsybakov, A. B · 2007
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Optimal exponential bounds for aggregation of density estimators
Bellec, P. C · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J. and Zisserman, A · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I · 2017
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
Stereo magnification: Learning view synthesis using multiplane images
Zhou, T., Tucker, R., Flynn, J., Fyffe, G., and Snavely, N · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
Nichol, A. Q. and Dhariwal, P · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Earlier work this paper cites.
Latent video diffusion models for high-fidelity long video generation
He, Y., Yang, T., Zhang, Y., Shan, Y., and Chen, Q · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2022
Cited alongside, same era.
Compositional visual generation with composable diffusion models
Liu, N., Li, S., Du, Y., Torralba, A., and Tenenbaum, J. B · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Progressive distillation for fast sampling of diffusion models
Salimans, T. and Ho, J · 2022
Cited alongside, same era.
Temporally consistent transformers for video generation
Yan, W., Hafner, D., James, S., and Abbeel, P · 2023
Later among the works it cites.
Video generation models as world simulators
Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., et al · 2024
Later among the works it cites.
Tutorial on diffusion models for imaging and vision
Chan, S. et al · 2024
Later among the works it cites.
Diffusion forcing: Next-token prediction meets full-sequence diffusion
Chen, B., Monso, D. M., Du, Y., Simchowitz, M., Tedrake, R., and Sitzmann, V · 2024
Later among the works it cites.
Diffusion is spectral autoregression, 2024
Dieleman, S · 2024
Later among the works it cites.
Compositional generative modeling: A single model is not all you need
Du, Y. and Kaelbling, L · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al · 2022
Cited alongside, same era.
All are worth words: A vit backbone for diffusion models
Bao, F., Nie, S., Xue, K., Cao, Y., Li, C., Su, H., and Zhu, J · 2023
Cited alongside, same era.
On the importance of noise scheduling for diffusion models
Chen, T · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Xu, Z., Feng, S., Cousineau, E., Du, Y., Burchfiel, B., Tedrake, R., and Song, S · 2023
Cited alongside, same era.
Reduce, reuse, recycle: Compositional generation with energy-based diffusion models and mcmc
Du, Y., Durkan, C., Strudel, R., Tenenbaum, J. B., Dieleman, S., Fergus, R., Sohl-Dickstein, J., Doucet, A., and Grathwohl, W. S · 2023
Cited alongside, same era.
Act3d: 3d feature field transformers for multi-task robotic manipulation
Gervet, T., Xian, Z., Gkanatsios, N., and Fragkiadaki, K · 2023
Cited alongside, same era.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y., Yang, C., Rao, A., Liang, Z., Wang, Y., Qiao, Y., Agrawala, M., Lin, D., and Dai, B · 2023
Cited alongside, same era.
Later among the works it cites.
Cat3d: Create anything in 3d with multi-view diffusion models
Gao, R., Holynski, A., Henzler, P., Brussee, A., Martin-Brualla, R., Srinivasan, P. P., Barron, J. T., and Poole, B · 2024
Later among the works it cites.
Photorealistic video generation with diffusion models
Gupta, A., Yu, L., Sohn, K., Gu, X., Hahn, M., Li, F.-F., Essa, I., Jiang, L., and Lezama, J · 2024
Later among the works it cites.
Simpler diffusion (sid2): 1.5 fid on imagenet512 with pixel-space diffusion
Hoogeboom, E., Mensink, T., Heek, J., Lamerigts, K., Gao, R., and Salimans, T · 2024
Later among the works it cites.
Vbench: Comprehensive benchmark suite for video generative models
Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., et al · 2024
Later among the works it cites.
Pyramidal flow matching for efficient video generative modeling
Jin, Y., Sun, Z., Li, N., Xu, K., Jiang, H., Zhuang, N., Huang, Q., Song, Y., Mu, Y., and Lin, Z · 2024
Later among the works it cites.
Analyzing and improving the training dynamics of diffusion models
Karras, T., Aittala, M., Lehtinen, J., Hellsten, J., Aila, T., and Laine, S · 2024
Later among the works it cites.
Hunyuanvideo: A systematic framework for large video generative models
Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., et al · 2024
Later among the works it cites.
Latte: Latent diffusion transformer for video generation
Ma, X., Wang, Y., Jia, G., Chen, X., Liu, Z., Li, Y.-F., Chen, C., and Qiao, Y · 2024
Later among the works it cites.
Rolling diffusion models
Ruhe, D., Heek, J., Salimans, T., and Hoogeboom, E · 2024
Later among the works it cites.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., et al · 2024
Later among the works it cites.
From slow bidirectional to fast causal video generators
Yin, T., Zhang, Q., Zhang, R., Freeman, W. T., Durand, F., Shechtman, E., and Huang, X · 2024
Later among the works it cites.
Open-sora: Democratizing efficient video production for all
Zheng, Z., Peng, X., Yang, T., Shen, C., Li, S., Liu, H., Zhou, Y., Li, T., and You, Y · 2024
Later among the works it cites.
Controlling space and time with diffusion models
Watson, D., Saxena, S., Li, L., Tagliasacchi, A., and Fleet, D. J · 2025
Closest in time.