Fetching the paper…
Reading the bibliography…
Diffusion models, as a type of generative model, have achieved impressive results in generating images and videos conditioned on textual conditions.
Perplexity—a measure of the difficulty of speech recognition tasks
Jelinek, F.; Mercer, R. L.; Bahl, L. R.; and Baker, J. K. 1977 · 1977
Earlier work this paper cites.
Making a “completely blind” image quality analyzer
Mittal, A.; Soundararajan, R.; and Bovik, A. C. 2012 · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Improved techniques for training gans
Salimans, T.; Goodfellow, I.; Zaremba, W.; Cheung, V.; Radford, A.; and Chen, X. 2016 · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J.; Mei, T.; Yao, T.; and Rui, Y. 2016 · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017 · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P.; and Nichol, A. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Denoising Diffusion Implicit Models
Song, J.; Meng, C.; and Ermon, S. 2021 · 2021
Earlier work this paper cites.
Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps
Lu, C.; Zhou, Y.; Bao, F.; Chen, J.; Li, C.; and Zhu, J. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Earlier work this paper cites.
LaionCOCO
Schuhmann, C.; Köpf, A.; Vencu, R.; Coombes, T.; and Beaumont, R. 2022 · 2022
Earlier work this paper cites.
Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models
Wang, Z. J.; Montoya, E.; Munechika, D.; Yang, H.; Hoover, B.; and Chau, D. H. 2022 · 2022
Cited alongside, same era.
Training diffusion models with reinforcement learning
Black, K.; Janner, M.; Du, Y.; Kostrikov, I.; and Levine, S. 2023 · 2023
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A.; Dockhorn, T.; Kulal, S.; Mendelevitch, D.; Kilian, M.; Lorenz, D.; Levi, Y.; English, Z.; Voleti, V.; Letts, A.; et al. 2023 · 2023
Cited alongside, same era.
DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models
Fan, Y.; Watkins, O.; Du, Y.; Liu, H.; Ryu, M.; Boutilier, C.; Abbeel, P.; Ghavamzadeh, M.; Lee, K.; and Lee, K. 2023 · 2023
Cited alongside, same era.
Human Preference Score: Better Aligning Text-to-Image Models with Human Preference
Wu, X.; Sun, K.; Zhu, F.; Zhao, R.; and Li, H. 2023 · 2023
Closest in time.
Diffir: Efficient diffusion model for image restoration
Xia, B.; Zhang, Y.; Wang, S.; Wang, Y.; Wu, X.; Tian, Y.; Yang, W.; and Van Gool, L. 2023 · 2023
Closest in time.
Imagereward: Learning and evaluating human preferences for text-to-image generation
Xu, J.; Liu, X.; Wu, Y.; Tong, Y.; Li, Q.; Ding, M.; Tang, J.; and Dong, Y. 2023 · 2023
Closest in time.
UniPC: A Unified Predictor-Corrector Framework for Fast Sampling of Diffusion Models
Zhao, W.; Bai, L.; Rao, Y.; Zhou, J.; and Lu, J. 2023 · 2023
Closest in time.
Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models
Bao, F.; Xiang, C.; Yue, G.; He, G.; Zhu, H.; Zheng, K.; Zhao, M.; Liu, S.; Wang, Y.; and Zhu, J. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Structural pruning for diffusion models
Fang, G.; Ma, X.; and Wang, X. 2023 · 2023
Cited alongside, same era.
Implicit diffusion models for continuous super-resolution
Gao, S.; Liu, X.; Zeng, B.; Xu, S.; Li, Y.; Luo, X.; Liu, J.; Zhen, X.; and Zhang, B. 2023 · 2023
Cited alongside, same era.
BK-SDM: Architecturally Compressed Stable Diffusion for Efficient Text-to-Image Generation
Kim, B.-K.; Song, H.-K.; Castells, T.; and Choi, S. 2023 · 2023
Cited alongside, same era.
Faster diffusion: Rethinking the role of unet encoder in diffusion models
Li, S.; Hu, T.; Khan, F. S.; Li, L.; Yang, S.; Wang, Y.; Cheng, M.-M.; and Yang, J. 2023 · 2023
Cited alongside, same era.
Latent consistency models: Synthesizing high-resolution images with few-step inference
Luo, S.; Tan, Y.; Huang, L.; Li, J.; and Zhao, H. 2023 · 2023
Cited alongside, same era.
On distillation of guided diffusion models
Meng, C.; Rombach, R.; Gao, R.; Kingma, D.; Ermon, S.; Ho, J.; and Salimans, T. 2023 · 2023
Cited alongside, same era.
Adversarial diffusion distillation
Sauer, A.; Lorenz, D.; Blattmann, A.; and Rombach, R. 2023 · 2023
Cited alongside, same era.
Segmind Stable Diffusion 1B
Segmend. 2023 · 2023
Cited alongside, same era.
Closest in time.
PixArt- α \alpha : Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Chen, J.; Yu, J.; Ge, C.; Yao, L.; Xie, E.; Wu, Y.; Wang, Z.; Kwok, J.; Luo, P.; Lu, H.; et al. 2024 · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P.; Kulal, S.; Blattmann, A.; Entezari, R.; Müller, J.; Saini, H.; Levi, Y.; Lorenz, D.; Sauer, A.; Boesel, F.; et al. 2024 · 2024
Closest in time.
SDXL-Lightning: Progressive Adversarial Diffusion Distillation
Lin, S.; Wang, A.; and Yang, X. 2024 · 2024
Closest in time.
Deepcache: Accelerating diffusion models for free
Ma, X.; Fang, G.; and Wang, X. 2024 · 2024
Closest in time.
Video generation models as world simulators
OpenAI. 2024 · 2024
Closest in time.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; Müller, J.; Penna, J.; and Rombach, R. 2024 · 2024
Closest in time.
Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis
Ren, Y.; Xia, X.; Lu, Y.; Zhang, J.; Wu, J.; Xie, P.; Wang, X.; and Xiao, X. 2024 · 2024
Closest in time.
Cache me if you can: Accelerating diffusion models through block caching
Wimbauer, F.; Wu, B.; Schoenfeld, E.; Dai, X.; Hou, J.; He, Z.; Sanakoyeu, A.; Zhang, P.; Tsai, S.; Kohler, J.; et al. 2024 · 2024
Closest in time.