Fetching the paper…
Reading the bibliography…
Diffusion models are the main driver of progress in image and video synthesis, but suffer from slow inference speed.
Estimation of non-normalized statistical models by score matching
A. Hyvärinen and P. Dayan · 2005
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
P. Vincent · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. N. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Improved techniques for training gans, 2016
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness
R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models, 2020
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Scaling laws for neural language models, 2020
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Generative adversarial networks are special cases of artificial curiosity (1990) and also closely related to predictability minimization (1991), 2020
J. Schmidhuber · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. N. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis, 2021
P. Dhariwal and A. Nichol · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Projected gans converge faster
A. Sauer, K. Chitta, J. Müller, and A. Geiger · 2021
Earlier work this paper cites.
ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers
Y. Balaji, S. Nah, X. Huang, A. Vahdat, J. Song, Q. Zhang, K. Kreis, M. Aittala, T. Aila, S. Laine, et al · 2022
Earlier work this paper cites.
Genie: Higher-order denoising diffusion solvers, 2022
T. Dockhorn, A. Vahdat, and K. Kreis · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models, 2022
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, and T. Salimans · 2022
Earlier work this paper cites.
Training compute-optimal large language models, 2022
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
T. Karras, M. Aittala, T. Aila, and S. Laine · 2022
Earlier work this paper cites.
Flow straight and fast: Learning to generate and transfer data with rectified flow, 2022
X. Liu, C. Gong, and Q. Liu · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents, 2022
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Cited alongside, same era.
Progressive distillation for fast sampling of diffusion models, 2022
T. Salimans and J. Ho · 2022
Cited alongside, same era.
Stylegan-xl: Scaling stylegan to large diverse datasets
Scalable diffusion models with transformers, 2023
W. Peebles and S. Xie · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach · 2023
Later among the works it cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Later among the works it cites.
Bespoke solvers for generative flow models, 2023
N. Shaul, J. Perez, R. T. Q. Chen, A. Thabet, A. Pumarola, and Y. Lipman · 2023
Later among the works it cites.
Emu edit: Precise image editing via recognition and generation tasks
S. Sheynin, A. Polyak, U. Singer, Y. Kirstain, A. Zohar, O. Ashual, D. Parikh, and Y. Taigman · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Sauer, K. Schwarz, and A. Geiger · 2022
Cited alongside, same era.
Make-a-video: Text-to-video generation without text-video data, 2022
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, D. Parikh, S. Gupta, and Y. Taigman · 2022
Cited alongside, same era.
Denoising diffusion implicit models, 2022
J. Song, C. Meng, and S. Ermon · 2022
Cited alongside, same era.
Resolution-robust large mask inpainting with fourier convolutions
R. Suvorov, E. Logacheva, A. Mashikhin, A. Remizova, A. Ashukha, A. Silvestrov, N. Kong, H. Goka, K. Park, and V. Lempitsky · 2022
Cited alongside, same era.
Scaling autoregressive models for content-rich text-to-image generation
J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid, Z. Wang, V. Vasudevan, A. Ku, Y. Yang, B. K. Ayan, et al · 2022
Cited alongside, same era.
Tract: Denoising diffusion models with transitive closure time-distillation
D. Berthelot, A. Autef, J. Lin, D. A. Yap, S. Zhai, S. Hu, D. Zheng, W. Talbott, and E. Gu · 2023
Cited alongside, same era.
Instructpix2pix: Learning to follow image editing instructions
T. Brooks, A. Holynski, and A. A. Efros · 2023
Cited alongside, same era.
Later among the works it cites.
Improved techniques for training consistency models
Y. Song and P. Dhariwal · 2023
Later among the works it cites.
Consistency models
Y. Song, P. Dhariwal, M. Chen, and I. Sutskever · 2023
Later among the works it cites.
Diffusion Model Alignment Using Direct Preference Optimization
B. Wallace, M. Dang, R. Rafailov, L. Zhou, A. Lou, S. Purushwalkam, S. Ermon, C. Xiong, S. Joty, and N. Naik · 2023
Later among the works it cites.
Ufogen: You forward once large scale text-to-image generation via diffusion gans, 2023
Y. Xu, Y. Zhao, Z. Xiao, and T. Hou · 2023
Later among the works it cites.
One-step diffusion with distribution matching distillation, 2023
T. Yin, M. Gharbi, R. Zhang, E. Shechtman, F. Durand, W. T. Freeman, and T. Park · 2023
Later among the works it cites.
Fast sampling of diffusion models with exponential integrator, 2023
Q. Zhang and Y. Chen · 2023
Later among the works it cites.
Hive: Harnessing human feedback for instructional visual editing
S. Zhang, X. Yang, Y. Feng, C. Qin, C.-C. Chen, N. Yu, Z. Chen, H. Wang, S. Savarese, S. Ermon, et al · 2023
Later among the works it cites.
Lumiere: A space-time diffusion model for video generation, 2024
O. Bar-Tal, H. Chefer, O. Tov, C. Herrmann, R. Paiss, S. Zada, A. Ephrat, J. Hur, Y. Li, T. Michaeli, O. Wang, D. Sun, T. Dekel, and I. Mosseri · 2024
Closest in time.
Improving image editing models with generative data refinement, 2024
F. Boesel and R. Rombach · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis, 2024
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, D. Podell, T. Dockhorn, Z. English, K. Lacey, A. Goodwin, Y. Marek, and R. Rombach · 2024
Closest in time.
J. Heek, E. Hoogeboom, and T. Salimans · 2024
Closest in time.
Sdxl-lightning: Progressive adversarial diffusion distillation, 2024
S. Lin, A. Wang, and X. Yang · 2024
Closest in time.
Diffusion hyperfeatures: Searching through time and space for semantic correspondence
G. Luo, L. Dunlap, D. H. Park, A. Holynski, and T. Darrell · 2024
Closest in time.
Magicbrush: A manually annotated dataset for instruction-guided image editing
K. Zhang, L. Mo, W. Chen, H. Sun, and Y. Su · 2024
Closest in time.
Trajectory consistency distillation
J. Zheng, M. Hu, Z. Fan, C. Wang, C. Ding, D. Tao, and T.-J. Cham · 2024
Closest in time.