Fetching the paper…
Reading the bibliography…
Diffusion Transformers (DiTs) with 3D full attention power state-of-the-art video generation, but suffer from prohibitive compute cost -- when generating just a 5-second 720P video, attention alone takes 800 out of 945 seconds of total inference time.
Longformer: The long-document transformer, 2020
Beltagy, I., Peters, M. E., and Cohan, A · 2004
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2015
Lin, T.-Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C. L., and Dollár, P · 2015
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Earlier work this paper cites.
Gotta go fast when generating data with score-based models
Jolicoeur-Martineau, A., Li, K., Piché-Taillefer, R., Kachman, T., and Mitliagkas, I · 2021
Earlier work this paper cites.
Learned queries for efficient local attention
Arar, M., Shamir, A., and Bermano, A. H · 2022
Earlier work this paper cites.
Flow matching for generative modeling
Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M · 2022
Earlier work this paper cites.
Sdedit: Guided image synthesis and editing with stochastic differential equations, 2022
Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S · 2022
Earlier work this paper cites.
Make-a-video: Text-to-video generation without text-video data, 2022
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., Parikh, D., Gupta, S., and Taigman, Y · 2022
Earlier work this paper cites.
Videocrafter1: Open diffusion models for high-quality video generation, 2023
Chen, H., Xia, M., He, Y., Zhang, Y., Cun, X., Yang, S., Xing, J., Liu, Y., Chen, Q., Wang, X., Weng, C., and Shan, Y · 2023
Earlier work this paper cites.
Neighborhood attention transformer
Hassani, A., Walton, S., Li, J., Li, S., and Shi, H · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Earlier work this paper cites.
Consistency trajectory models: Learning probability flow ode trajectory of diffusion
Kim, D., Lai, C.-H., Liao, W.-H., Murata, N., Takida, Y., Uesaka, T., He, Y., Mitsufuji, Y., and Ermon, S · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Adversarial diffusion distillation
Sauer, A., Lorenz, D., Blattmann, A., and Rombach, R · 2023
Cited alongside, same era.
Song, Y., Dhariwal, P., Chen, M., and Sutskever, I · 2023
Cited alongside, same era.
Lavie: High-quality video generation with cascaded latent diffusion models
Wang, Y., Chen, X., Ma, X., Zhou, S., Huang, Z., Wang, Y., Yang, C., He, Y., Yu, J., Yang, P., et al · 2023
Cited alongside, same era.
One-step diffusion with distribution matching distillation
Yin, T., Gharbi, M., Zhang, R., Shechtman, E., Durand, F., Freeman, W. T., and Park, T · 2023
Latte: Latent diffusion transformer for video generation
Ma, X., Wang, Y., Jia, G., Chen, X., Liu, Z., Li, Y.-F., Chen, C., and Qiao, Y · 2024
Later among the works it cites.
Sora, 2024
OpenAI · 2024
Later among the works it cites.
Movie gen: A cast of media foundation models
Polyak, A., Zohar, A., Brown, A., Tjandra, A., Sinha, A., Lee, A., Vyas, A., Shi, B., Ma, C.-Y., Chuang, C.-Y., et al · 2024
Later among the works it cites.
Multistep distillation of diffusion models via moment matching
Salimans, T., Mensink, T., Heek, J., and Hoogeboom, E · 2024
Later among the works it cites.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision, 2024
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Delta-dit: A training-free acceleration method tailored for diffusion transformers
Chen, P., Shen, M., Ye, P., Cao, J., Tu, C., Bouganis, C.-S., Zhao, Y., and Chen, T · 2024
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2024
Cited alongside, same era.
Flex attention: A programming model for generating optimized attention kernels, 2024
Dong, J., Feng, B., Guessous, D., Liang, Y., and He, H · 2024
Cited alongside, same era.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al · 2024
Cited alongside, same era.
On the content bias in frechet video distance
Ge, S., Mahapatra, A., Parmar, G., Zhu, J.-Y., and Huang, J.-B · 2024
Cited alongside, same era.
Vbench: Comprehensive benchmark suite for video generative models
Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., et al · 2024
Cited alongside, same era.
Open-sora plan: Open-source large video generation model
Lin, B., Ge, Y., Cheng, X., Li, Z., Zhu, B., Wang, S., He, X., Ye, Y., Yuan, S., Chen, L., et al · 2024
Cited alongside, same era.
Later among the works it cites.
Thunderkittens: Simple, fast, and adorable ai kernels, 2024
Spector, B. F., Arora, S., Singhal, A., Fu, D. Y., and Ré, C · 2024
Later among the works it cites.
Mlcm: Multistep consistency distillation of latent diffusion model
Xie, Q., Liao, Z., Deng, Z., Tang, S., Lu, H., et al · 2024
Later among the works it cites.
One-step diffusion with distribution matching distillation
Yin, T., Gharbi, M., Zhang, R., Shechtman, E., Durand, F., Freeman, W. T., and Park, T · 2024
Later among the works it cites.
Sageattention: Accurate 8-bit attention for plug-and-play inference acceleration
Zhang, J., Wei, J., Huang, H., Zhang, P., Zhu, J., and Chen, J · 2024
Later among the works it cites.
Real-time video generation with pyramid attention broadcast
Zhao, X., Jin, X., Wang, K., and You, Y · 2024
Later among the works it cites.
Open-sora: Democratizing efficient video production for all, March 2024
Zheng, Z., Peng, X., Yang, T., Shen, C., Li, S., Liu, H., Zhou, Y., Li, T., and You, Y · 2024
Later among the works it cites.
Hunyuanvideo: A systematic framework for large video generative models, 2025
HunyuanVideo-Team · 2025
Closest in time.
Sparse videogen: Accelerating video diffusion transformers with spatial-temporal sparsity
Xi, H., Yang, S., Zhao, Y., Xu, C., Li, M., Li, X., Lin, Y., Cai, H., Zhang, J., Li, D., et al · 2025
Closest in time.