Fetching the paper…
Reading the bibliography…
Diffusion models create data from noise by inverting the forward paths of data towards noise and have emerged as a powerful generative modeling technique for high-dimensional, perceptual data such as images and videos.
Logistic-normal distributions: Some properties and uses
Atchison, J. and Shen, S. M · 1980
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A · 2005
Earlier work this paper cites.
Optimal transport: Old and new
Villani, C · 2008
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context , pp. 740–755
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M. S., Berg, A. C., and Fei-Fei, L · 2014
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation , pp. 234–241
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J. N., Weiss, E. A., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium, 2017
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Fixing weight decay regularization in adam
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention is all you need, 2017
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset, 2018
Carreira, J. and Zisserman, A · 2018
Earlier work this paper cites.
Neural ordinary differential equations
Chen, T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K · 2018
Earlier work this paper cites.
Bfloat16: The secret to high performance on cloud tpus, 2019
Chen, D., Chou, C., Xu, Y., and Hseu, J · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2019
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Earlier work this paper cites.
Root mean square layer normalization, 2019
Zhang, B. and Sennrich, R · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models, 2020
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution, 2020
Song, Y. and Ermon, S · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J. N., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2020
Earlier work this paper cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., Sharma, P. K., Ding, N., and Soricut, R · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis, 2021
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
Score-based generative modeling with critically-damped langevin diffusion
Dockhorn, T., Vahdat, A., and Kreis, K · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J., Holtzman, A., Forbes, M., Le Bras, R., and Choi, Y · 2021
Earlier work this paper cites.
Deduplicating training data makes language models better
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models, 2021
Nichol, A. and Dhariwal, P · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
Projected gans converge faster
Sauer, A., Chitta, K., Müller, J., and Geiger, A · 2021
Cited alongside, same era.
Building normalizing flows with stochastic interpolants, 2022
Albergo, M. S. and Vanden-Eijnden, E · 2022
Cited alongside, same era.
ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers, 2022
Balaji, Y., Nah, S., Huang, X., Vahdat, A., Song, J., Zhang, Q., Kreis, K., Aittala, M., Aila, T., Laine, S., Catanzaro, B., Karras, T., and Liu, M.-Y · 2022
Cited alongside, same era.
Genie: Higher-order denoising diffusion solvers, 2022
Dockhorn, T., Vahdat, A., and Kreis, K · 2022
Cited alongside, same era.
Classifier-free diffusion guidance, 2022
Ho, J. and Salimans, T · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models, 2022
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., and Salimans, T · 2022
Boosting latent diffusion with flow matching
Fischer, J. S., Gui, M., Ma, P., Stracke, N., Baumann, S. A., and Ommer, B · 2023
Later among the works it cites.
Geneval: An object-focused framework for evaluating text-to-image alignment
Ghosh, D., Hajishirzi, H., and Schmidt, L · 2023
Later among the works it cites.
Photorealistic video generation with diffusion models, 2023
Gupta, A., Yu, L., Sohn, K., Gu, X., Hahn, M., Fei-Fei, L., Essa, I., Jiang, L., and Lezama, J · 2023
Later among the works it cites.
Simple diffusion: End-to-end diffusion for high resolution images, 2023
Hoogeboom, E., Heek, J., and Salimans, T · 2023
Later among the works it cites.
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation
Huang, K., Sun, K., Xie, E., Li, Z., and Liu, X · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Cited alongside, same era.
Flow straight and fast: Learning to generate and transfer data with rectified flow, 2022
Liu, X., Gong, C., and Liu, Q · 2022
Cited alongside, same era.
Dall-e 2 pre-training mitigations
Nichol, A · 2022
Cited alongside, same era.
Novelai improvements on stable diffusion, 2022
NovelAI · 2022
Cited alongside, same era.
A self-supervised descriptor for image copy detection
Pizzi, E., Roy, S. D., Ravindra, S. N., Goyal, P., and Douze, M · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents, 2022
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Cited alongside, same era.
Later among the works it cites.
Analyzing and improving the training dynamics of diffusion models
Karras, T., Aittala, M., Lehtinen, J., Hellsten, J., Aila, T., and Laine, S · 2023
Later among the works it cites.
Understanding diffusion objectives as the elbo with simple data augmentation
Kingma, D. P. and Gao, R · 2023
Later among the works it cites.
Minimizing trajectory curvature of ode-based generative models, 2023
Lee, S., Kim, B., and Ye, J. C · 2023
Later among the works it cites.
Flow matching for generative modeling
Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M · 2023
Later among the works it cites.
Instaflow: One step is enough for high-quality diffusion-based text-to-image generation, 2023
Liu, X., Zhang, X., Ma, J., Peng, J., and Liu, Q · 2023
Later among the works it cites.
Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models, 2023
Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J · 2023
Later among the works it cites.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Later among the works it cites.
Wuerstchen: An efficient architecture for large-scale text-to-image diffusion models, 2023
Pernias, P., Rampas, D., Richter, M. L., Pal, C. J., and Aubreville, M · 2023
Later among the works it cites.
State of the art on diffusion models for visual computing
Po, R., Yifan, W., Golyanik, V., Aberman, K., Barron, J. T., Bermano, A. H., Chan, E. R., Dekel, T., Holynski, A., Kanazawa, A., et al · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Later among the works it cites.
Multisample flow matching: Straightening flows with minibatch couplings, 2023
Pooladian, A.-A., Ben-Hamu, H., Domingo-Enrich, C., Amos, B., Lipman, Y., and Chen, R. T. Q · 2023
Later among the works it cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Adversarial diffusion distillation
Sauer, A., Lorenz, D., Blattmann, A., and Rombach, R · 2023
Later among the works it cites.
Emu edit: Precise image editing via recognition and generation tasks
Sheynin, S., Polyak, A., Singer, U., Kirstain, Y., Zohar, A., Ashual, O., Parikh, D., and Taigman, Y · 2023
Later among the works it cites.
Improving and generalizing flow-based generative models with minibatch optimal transport, 2023
Tong, A., Malkin, N., Huguet, G., Zhang, Y., Rector-Brooks, J., Fatras, K., Wolf, G., and Bengio, Y · 2023
Later among the works it cites.
Diffusion Model Alignment Using Direct Preference Optimization
Wallace, B., Dang, M., Rafailov, R., Zhou, L., Lou, A., Purushwalkam, S., Ermon, S., Xiong, C., Joty, S., and Naik, N · 2023
Later among the works it cites.
Cogvlm: Visual expert for pretrained language models
Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., Song, X., et al · 2023
Later among the works it cites.
Small-scale proxies for large-scale transformer training instabilities, 2023
Wortsman, M., Liu, P. J., Xiao, L., Everett, K., Alemi, A., Adlam, B., Co-Reyes, J. D., Gur, I., Kumar, A., Novak, R., Pennington, J., Sohl-dickstein, J., Xu, K., Lee, J., Gilmer, J., and Kornblith, S · 2023
Later among the works it cites.
URL https://about.ideogram.ai/1.0
Ideogram v1.0 announcement, 2024 · 2024
Closest in time.
URL https://blog.playgroundai.com/playground-v2-5/
Playground v2.5 announcement, 2024 · 2024
Closest in time.
Common diffusion noise schedules and sample steps are flawed
Lin, S., Liu, B., Li, J., and Yang, X · 2024
Closest in time.
Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers, 2024
Ma, N., Goldstein, M., Albergo, M. S., Boffi, N. M., Vanden-Eijnden, E., and Xie, S · 2024
Closest in time.