Fetching the paper…
Reading the bibliography…
The diffusion models are widely used for image and video generation, but their iterative generation process is slow and expansive.
Generative adversarial nets
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. C., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Stabilizing training of generative adversarial networks through regularization
Roth, K., Lucchi, A., Nowozin, S., and Hofmann, T · 2017
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2018
Earlier work this paper cites.
Which training methods for gans do actually converge?
Mescheder, L. M., Geiger, A., and Nowozin, S · 2018
Earlier work this paper cites.
Large scale GAN training for high fidelity natural image synthesis
Brock, A., Donahue, J., and Simonyan, K · 2019
Earlier work this paper cites.
Efficient video generation on complex datasets
Clark, A., Donahue, J., and Simonyan, K · 2019
Earlier work this paper cites.
Analyzing and improving the image quality of stylegan
Karras, T., Laine, S., Aittala, M., Hellsten, J., Lehtinen, J., and Aila, T · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J. N., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2020
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2021
Earlier work this paper cites.
Alias-free generative adversarial networks
Karras, T., Aittala, M., Laine, S., Härkönen, E., Hellsten, J., Lehtinen, J., and Aila, T · 2021
Earlier work this paper cites.
On buggy resizing libraries and surprising subtleties in fid calculation
Parmar, G., Zhang, R., and Zhu, J.-Y · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2021
Earlier work this paper cites.
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
Skorokhodov, I., Tulyakov, S., and Elhoseiny, M · 2021
Earlier work this paper cites.
A good image generator is what you need for high-resolution video synthesis
Tian, Y., Ren, J., Chai, M., Olszewski, K., Peng, X., Metaxas, D. N., and Tulyakov, S · 2021
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M · 2022
Cited alongside, same era.
Progressive distillation for fast sampling of diffusion models
Salimans, T. and Ho, J · 2022
Cited alongside, same era.
Diffusers: State-of-the-art diffusion models
von Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., Nair, D., Paul, S., Berman, W., Xu, Y., Liu, S., and Wolf, T · 2022
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Muller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., Podell, D., Dockhorn, T., English, Z., Lacey, K., Goodwin, A., Marek, Y., and Rombach, R · 2024
Later among the works it cites.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y., Yang, C., Rao, A., Liang, Z., Wang, Y., Qiao, Y., Agrawala, M., Lin, D., and Dai, B · 2024
Later among the works it cites.
Distilling Diffusion Models into Conditional GANs
Kang, M., Zhang, R., Barnes, C., Paris, S., Kwak, S., Park, J., Shechtman, E., Zhu, J.-Y., and Park, T · 2024
Later among the works it cites.
Guiding a diffusion model with a bad version of itself
Karras, T., Aittala, M., Kynkäänniemi, T., Lehtinen, J., Aila, T., and Laine, S · 2024
Later among the works it cites.
Imagine flash: Accelerating emu diffusion models with backward distillation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Diffusion-gan: Training gans with diffusion
Wang, Z., Zheng, H., He, P., Chen, W., and Zhou, M · 2022
Cited alongside, same era.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., Hutchinson, B., Han, W., Parekh, Z., Li, X., Zhang, H., Baldridge, J., and Wu, Y · 2022
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Cited alongside, same era.
Scaling up gans for text-to-image synthesis
Kang, M., Zhu, J.-Y., Zhang, R., Park, J., Shechtman, E., Paris, S., and Park, T · 2023
Cited alongside, same era.
Diffusion model with perceptual loss
Lin, S. and Yang, X · 2023
Cited alongside, same era.
Common diffusion noise schedules and sample steps are flawed
Lin, S., Liu, B., Li, J., and Yang, X · 2023
Cited alongside, same era.
Flow matching for generative modeling
Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M · 2023
Cited alongside, same era.
Kohler, J., Pumarola, A., Schönfeld, E., Sanakoyeu, A., Sumbaly, R., Vajda, P., and Thabet, A. K · 2024
Later among the works it cites.
Animatediff-lightning: Cross-model diffusion distillation
Lin, S. and Yang, X · 2024
Later among the works it cites.
Sdxl-lightning: Progressive adversarial diffusion distillation
Lin, S., Wang, A., and Yang, X · 2024
Later among the works it cites.
Simplifying, stabilizing and scaling continuous-time consistency models
Lu, C. and Song, Y · 2024
Later among the works it cites.
Osv: One step is enough for high-quality image to video generation
Mao, X., Jiang, Z., Wang, F.-Y., Zhang, J., Chen, H., Chi, M., Wang, Y., and Luo, W · 2024
Later among the works it cites.
Hyper-SD: Trajectory segmented consistency model for efficient image synthesis
Ren, Y., Xia, X., Lu, Y., Zhang, J., Wu, J., Xie, P., Wang, X., and Xiao, X · 2024
Later among the works it cites.
Fast high-resolution image synthesis with latent adversarial diffusion distillation
Sauer, A., Boesel, F., Dockhorn, T., Blattmann, A., Esser, P., and Rombach, R · 2024
Later among the works it cites.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T · 2024
Later among the works it cites.
Improved techniques for training consistency models
Song, Y. and Dhariwal, P · 2024
Later among the works it cites.
Ufogen: You forward once large scale text-to-image generation via diffusion gans
Xu, Y., Zhao, Y., Xiao, Z., and Hou, T · 2024
Later among the works it cites.
Ufogen: You forward once large scale text-to-image generation via diffusion gans
Xu, Y., Zhao, Y., Xiao, Z., and Hou, T · 2024
Later among the works it cites.
Perflow: Piecewise rectified flow as universal plug-and-play accelerator
Yan, H., Liu, X., Pan, J., Liew, J. H., Liu, Q., and Feng, J · 2024
Later among the works it cites.
Zhai, Y., Lin, K. Q., Yang, Z., Li, L., Wang, J., Lin, C.-C., Doermann, D., Yuan, J., and Wang, L · 2024
Later among the works it cites.
Sf-v: Single forward video generation model
Zhang, Z., Li, Y., Wu, Y., Kag, A., Skorokhodov, I., Menapace, W., Siarohin, A., Cao, J., Metaxas, D., Tulyakov, S., et al · 2024
Later among the works it cites.
Adversarial diffusion distillation
Sauer, A., Lorenz, D., Blattmann, A., and Rombach, R · 2025
Closest in time.
Seaweed-7b: Cost-effective training of video generation foundation model
Seawead, T., Yang, C., Lin, Z., Zhao, Y., Lin, S., Ma, Z., Guo, H., Chen, H., Qi, L., Wang, S., et al · 2025
Closest in time.