Fetching the paper…
Reading the bibliography…
Diffusion models (DMs) have achieved significant success in generating imaginative images given textual descriptions.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Generative Modeling by Estimating Gradients of the Data Distribution
Song, Y.; and Ermon, S. 2019 · 2019
Earlier work this paper cites.
WaveGrad: Estimating Gradients for Waveform Generation
Chen, N.; Zhang, Y.; Zen, H.; Weiss, R. J.; Norouzi, M.; and Chan, W. 2020 · 2020
Earlier work this paper cites.
Retinaface: Single-shot multi-level face localisation in the wild
Deng, J.; Guo, J.; Ververas, E.; Kotsia, I.; and Zafeiriou, S. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Score-Based Generative Modeling through Stochastic Differential Equations
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2020 · 2020
Earlier work this paper cites.
SER-FIQ: Unsupervised estimation of face image quality based on stochastic embedding robustness
Terhorst, P.; Kolf, J. N.; Damer, N.; Kirchbuchner, F.; and Kuijper, A. 2020 · 2020
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P.; Rombach, R.; and Ommer, B. 2021 · 2021
Earlier work this paper cites.
Classifier-Free Diffusion Guidance
Ho, J.; and Salimans, T. 2021 · 2021
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al. 2021 · 2021
Earlier work this paper cites.
DiffWave: A Versatile Diffusion Model for Audio Synthesis
Kong, Z.; Ping, W.; Huang, J.; Zhao, K.; and Catanzaro, B. 2021 · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
Nichol, A. Q.; and Dhariwal, P. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Blended diffusion for text-driven editing of natural images
Avrahami, O.; Lischinski, D.; and Fried, O. 2022 · 2022
Earlier work this paper cites.
Imagen Video: High Definition Video Generation with Diffusion Models
Ho, J.; Chan, W.; Saharia, C.; Whang, J.; Gao, R.; Gritsenko, A.; Kingma, D. P.; Poole, B.; Norouzi, M.; Fleet, D. J.; and Salimans, T. 2022 · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022 · 2022
Cited alongside, same era.
Repaint: Inpainting using denoising diffusion probabilistic models
Lugmayr, A.; Danelljan, M.; Romero, A.; Yu, F.; Timofte, R.; and Van Gool, L. 2022 · 2022
Cited alongside, same era.
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
Nichol, A. Q.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; Mcgrew, B.; Sutskever, I.; and Chen, M. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Cited alongside, same era.
Kandinsky 3.0 Technical Report
Arkhipkin, V.; Filatov, A.; Vasilev, V.; Maltseva, A.; Azizov, S.; Pavlov, I.; Agafonova, J.; Kuznetsov, A.; and Dimitrov, D. 2024 · 2024
Closest in time.
Training Diffusion Models with Reinforcement Learning
Black, K.; Janner, M.; Du, Y.; Kostrikov, I.; and Levine, S. 2024 · 2024
Closest in time.
PixArt- α \alpha : Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Chen, J.; Jincheng, Y.; Chongjian, G.; Yao, L.; Xie, E.; Wang, Z.; Kwok, J.; Luo, P.; Lu, H.; and Li, Z. 2024 · 2024
Closest in time.
Directly Fine-Tuning Diffusion Models on Differentiable Rewards
Clark, K.; Vicol, P.; Swersky, K.; and Fleet, D. J. 2024 · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P.; Kulal, S.; Blattmann, A.; Entezari, R.; Müller, J.; Saini, H.; Levi, Y.; Lorenz, D.; Sauer, A.; Boesel, F.; et al. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. 2022 · 2022
Cited alongside, same era.
Blended latent diffusion
Avrahami, O.; Fried, O.; and Lischinski, D. 2023 · 2023
Cited alongside, same era.
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Blattmann, A.; Dockhorn, T.; Kulal, S.; Mendelevitch, D.; Kilian, M.; Lorenz, D.; Levi, Y.; English, Z.; Voleti, V.; Letts, A.; Jampani, V.; and Rombach, R. 2023 · 2023
Cited alongside, same era.
Instructpix2pix: Learning to follow image editing instructions
Brooks, T.; Holynski, A.; and Efros, A. A. 2023 · 2023
Cited alongside, same era.
Photorealistic Video Generation with Diffusion Models
Gupta, A.; Yu, L.; Sohn, K.; Gu, X.; Hahn, M.; Fei-Fei, L.; Essa, I.; Jiang, L.; and Lezama, J. 2023 · 2023
Cited alongside, same era.
Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes
Ju, X.; Zeng, A.; Wang, J.; Xu, Q.; and Zhang, L. 2023 · 2023
Cited alongside, same era.
Reinforcement learning for fine-tuning text-to-image diffusion models
Fan, Y.; Watkins, O.; Du, Y.; Liu, H.; Ryu, M.; Boutilier, C.; Abbeel, P.; Ghavamzadeh, M.; Lee, K.; and Lee, K. 2024 · 2024
Closest in time.
Fang, G.; Yan, W.; Guo, Y.; Han, J.; Jiang, Z.; Xu, H.; Liao, S.; and Liang, X. 2024 · 2024
Closest in time.
Kolors: Effective Training of Diffusion Model for Photorealistic Text-to-Image Synthesis
Kuaishou. 2024 · 2024
Closest in time.
Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpainting
Lu, W.; Xu, Y.; Zhang, J.; Wang, C.; and Tao, D. 2023 · 2024
Closest in time.
Würstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion Models
Pernias, P.; Rampas, D.; Richter, M. L.; Pal, C.; and Aubreville, M. 2024 · 2024
Closest in time.
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; Müller, J.; Penna, J.; and Rombach, R. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2024 · 2024
Closest in time.
Finetuning Text-to-Image Diffusion Models for Fairness
Shen, X.; Du, C.; Pang, T.; Lin, M.; Wong, Y.; and Kankanhalli, M. 2024 · 2024
Closest in time.
Diffusion model alignment using direct preference optimization
Wallace, B.; Dang, M.; Rafailov, R.; Zhou, L.; Lou, A.; Purushwalkam, S.; Ermon, S.; Xiong, C.; Joty, S.; and Naik, N. 2024 · 2024
Closest in time.
Imagereward: Learning and evaluating human preferences for text-to-image generation
Xu, J.; Liu, X.; Wu, Y.; Tong, Y.; Li, Q.; Ding, M.; Tang, J.; and Dong, Y. 2024 · 2024
Closest in time.