Fetching the paper…
Reading the bibliography…
Latent diffusion models have emerged as the leading approach for generating high-quality images and videos, utilizing compressed latent representations to reduce the computational burden of the diffusion process.
The jpeg still picture compression standard
Wallace, G. K · 1991
Earlier work this paper cites.
Origins of scaling in natural images
Ruderman, D. L · 1997
Earlier work this paper cites.
Discrete cosine transform
Ahmed, N., Natarajan, T., and Rao, K. R · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations, 2021
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
End-to-end optimized image compression
Ballé, J., Laparra, V., and Simoncelli, E. P · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I · 2017
Earlier work this paper cites.
Variational image compression with a scale hyperprior
Ballé, J., Minnen, D., Singh, S., Hwang, S. J., and Johnston, N · 2018
Earlier work this paper cites.
Which training methods for gans do actually converge?
Mescheder, L., Geiger, A., and Nowozin, S · 2018
Earlier work this paper cites.
Joint autoregressive and hierarchical priors for learned image compression
Minnen, D., Ballé, J., and Toderici, G. D · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
A short note on the kinetics-700 human action dataset
Carreira, J., Noland, E., Hillier, C., and Zisserman, A · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2019
Earlier work this paper cites.
Learned image compression with discretized gaussian mixture likelihoods and attention modules
Cheng, Z., Sun, H., Takeuchi, M., and Katto, J · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Deep contextual video compression
Li, J., Li, B., and Lu, Y · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
Nichol, A. Q. and Dhariwal, P · 2021
Earlier work this paper cites.
Score-based generative modeling in latent space. 2021
Vahdat, A., Kreis, K., and Kautz, J · 2021
Earlier work this paper cites.
Building normalizing flows with stochastic interpolants
Albergo, M. S. and Vanden-Eijnden, E · 2022
Earlier work this paper cites.
Coyo-700m: Image-text pair dataset
Byeon, M., Park, B., Kim, H., Lee, S., Baek, W., and Kim, S · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Earlier work this paper cites.
Flow straight and fast: Learning to generate and transfer data with rectified flow
Liu, X., Gong, C., and Liu, Q · 2022
Cited alongside, same era.
Vct: A video compression transformer
Mentzer, F., Toderici, G., Minnen, D., Hwang, S.-J., Caelles, S., Lucic, M., and Agustsson, E · 2022
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al · 2022
Hunyuanvideo: A systematic framework for large video generative models
Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., et al · 2024
Later among the works it cites.
Applying guidance in a limited interval improves sample and distribution quality in diffusion models
Kynkäänniemi, T., Aittala, M., Karras, T., Laine, S., Aila, T., and Lehtinen, J · 2024
Later among the works it cites.
On error propagation of diffusion models
Li, Y. and van der Schaar, M · 2024
Later among the works it cites.
Open-sora plan: Open-source large video generation model
Lin, B., Ge, Y., Cheng, X., Li, Z., Zhu, B., Wang, S., He, X., Ye, Y., Yuan, S., Chen, L., et al · 2024
Later among the works it cites.
Instaflow: One step is enough for high-quality diffusion-based text-to-image generation, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Temporal context mining for learned video compression
Sheng, X., Li, J., Li, B., Li, L., Liu, D., and Lu, Y · 2022
Cited alongside, same era.
Diffusers: State-of-the-art diffusion models
von Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., Nair, D., Paul, S., Berman, W., Xu, Y., Liu, S., and Wolf, T · 2022
Cited alongside, same era.
Improving image generation with better captions
Betker, J., Goh, G., Jing, L., Brooks†, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo†, Y., Manassra, W., Dhariwal, P., Chu, C., Jiao†, Y., and Ramesh, A · 2023
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al · 2023
Cited alongside, same era.
Fit: Far-reaching interleaved transformers
Chen, T. and Li, L · 2023
Cited alongside, same era.
Dao, Q., Phung, H., Nguyen, B., and Tran, A · 2023
Cited alongside, same era.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y., Yang, C., Rao, A., Wang, Y., Qiao, Y., Lin, D., and Dai, B · 2023
Cited alongside, same era.
Liu, X., Zhang, X., Ma, J., Peng, J., and Liu, Q · 2024
Later among the works it cites.
Snap video: Scaled spatiotemporal transformers for text-to-video synthesis
Menapace, W., Siarohin, A., Skorokhodov, I., Deyneka, E., Chen, T.-S., Kag, A., Fang, Y., Stoliar, A., Ricci, E., Ren, J., et al · 2024
Later among the works it cites.
Dctdiff: Intriguing properties of image generative modeling in the dct space
Ning, M., Li, M., Su, J., Jia, H., Liu, L., Beneš, M., Salah, A. A., and Ertugrul, I. O · 2024
Later among the works it cites.
Movie gen: A cast of media foundation models
Polyak, A., Zohar, A., Brown, A., Tjandra, A., Sinha, A., Lee, A., Vyas, A., Shi, B., Ma, C.-Y., Chuang, C.-Y., et al · 2024
Later among the works it cites.
Litevae: Lightweight and efficient variational autoencoders for latent diffusion models
Sadat, S., Buhmann, J., Bradley, D., Hilliges, O., and Weber, R. M · 2024
Later among the works it cites.
Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models
Stein, G., Cresswell, J., Hosseinzadeh, R., Sui, Y., Ross, B., Villecroze, V., Liu, Z., Caterini, A. L., Taylor, E., and Loaiza-Ganem, G · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Later among the works it cites.
Vidtok: A versatile and open-source video tokenizer
Tang, A., He, T., Guo, J., Cheng, X., Song, L., and Bian, J · 2024
Later among the works it cites.
Reducio! generating 1024x1024 video within 16 seconds using extremely compressed motion latents
Tian, R., Dai, Q., Bao, J., Qiu, K., Yang, Y., Luo, C., Wu, Z., and Jiang, Y.-G · 2024
Later among the works it cites.
Omnitokenizer: A joint image-video tokenizer for visual generation
Wang, J., Jiang, Y., Yuan, Z., Peng, B., Wu, Z., and Jiang, Y.-G · 2024
Later among the works it cites.
Improved video vae for latent video diffusion model
Wu, P., Zhu, K., Liu, Y., Zhao, L., Zhai, W., Cao, Y., and Zha, Z.-J · 2024
Later among the works it cites.
Sana: Efficient high-resolution image synthesis with linear diffusion transformers
Xie, E., Chen, J., Chen, J., Cai, H., Tang, H., Lin, Y., Zhang, Z., Li, M., Zhu, L., Lu, Y., et al · 2024
Later among the works it cites.
Large motion video autoencoding with cross-modal video vae
Xing, Y., Fei, Y., He, Y., Chen, J., Xie, J., Chi, X., and Chen, Q · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., et al · 2024
Later among the works it cites.
Cv-vae: A compatible video vae for latent generative video models
Zhao, S., Zhang, Y., Cun, X., Yang, S., Niu, M., Li, X., Hu, W., and Shan, Y · 2024
Later among the works it cites.
Open-sora: Democratizing efficient video production for all, 2024
Zheng, Z., Peng, X., Yang, T., Shen, C., Li, S., Liu, H., Zhou, Y., Li, T., and You, Y · 2024
Later among the works it cites.
Allegro: Open the black box of commercial-level video generation model
Zhou, Y., Wang, Q., Cai, Y., and Yang, H · 2024
Later among the works it cites.
Cosmos world foundation model platform for physical ai
Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., et al · 2025
Closest in time.
Deep compression autoencoder for efficient high-resolution diffusion models
Chen, J., Cai, H., Chen, J., Xie, E., Yang, S., Tang, H., Li, M., Lu, Y., and Han, S · 2025
Closest in time.
Learnings from scaling visual tokenizers for reconstruction and generation
Hansen-Estruch, P., Yan, D., Chung, C.-Y., Zohar, O., Wang, J., Hou, T., Xu, T., Vishwanath, S., Vajda, P., and Chen, X · 2025
Closest in time.
Eq-vae: Equivariance regularized latent space for improved generative image modeling
Kouzelis, T., Ioannis, K., Spyros, G., and Nikos, K · 2025
Closest in time.
Wan: Open and advanced large-scale video generative models
Wan, T., Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.-W., Chen, D., Yu, F., Zhao, H., Yang, J., et al · 2025
Closest in time.
Alias-free latent diffusion models: Improving fractional shift equivariance of diffusion latent space
Zhou, Y., Xiao, Z., Yang, S., and Pan, X · 2025
Closest in time.