Fetching the paper…
Reading the bibliography…
Diffusion Transformers (DiTs) have shown remarkable performance in generating high-quality videos.
UCF101: A dataset of 101 human actions classes from videos in the wild
K Soomro. 2012 · 2012
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Attention is all you need. In NIPS
A Waswani, N Shazeer, N Parmar, J Uszkoreit, L Jones, A Gomez, L Kaiser, and I Polosukhin. 2017 · 2017
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. 2018 · 2018
Earlier work this paper cites.
Information aggregation for multi-head attention with routing-by-agreement
Jian Li, Baosong Yang, Zi-Yi Dou, Xing Wang, Michael R Lyu, and Zhaopeng Tu. 2019 · 2019
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov. 2019 · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models. In Proc. NeurIPS
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval. In Proc. IEEE/CVF ICCV
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021 · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness. In Advances in Neural Information Processing Systems
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022 · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. 2022 · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. 2022 · 2022
Earlier work this paper cites.
Flow matching for generative modeling
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2022 · 2022
Earlier work this paper cites.
Pixartalpha : Fast training of diffusion transformer for photorealistic text-to-image synthesis
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al · 2023
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao. 2023 · 2023
Earlier work this paper cites.
Neighborhood attention transformer. In Proc. IEEE/CVF CVPR
Ali Hassani, Steven Walton, Jiachen Li, Shen Li, and Humphrey Shi. 2023 · 2023
Earlier work this paper cites.
Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Shuaiwen Leon Song, Samyam Rajbhandari, and Yuxiong He. 2023 · 2023
Earlier work this paper cites.
Ring attention with blockwise transformers for near-infinite context
Hao Liu, Matei Zaharia, and Pieter Abbeel. 2023 · 2023
Earlier work this paper cites.
OpenAI. 2023 · 2023
Cited alongside, same era.
Scalable diffusion models with transformers. In Proc. IEEE/CVF ICCV
William Peebles and Saining Xie. 2023 · 2023
Cited alongside, same era.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2023 · 2023
Cited alongside, same era.
H2o: Heavy-hitter oracle for efficient generative inference of large language models. In Proc. NeurIPS
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al · 2023
Cited alongside, same era.
Pytorch fsdp: experiences on scaling fully sharded data parallel
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al · 2023
Vbench: Comprehensive benchmark suite for video generative models. In Proc. IEEE/CVF CVPR
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al · 2024
Later among the works it cites.
Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention
Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Zhenhua Han, Amir H Abdi, Dongsheng Li, Chin-Yew Lin, et al · 2024
Later among the works it cites.
Moh: Multi-head attention as mixture-of-head attention
Peng Jin, Bo Zhu, Li Yuan, and Shuicheng Yan. 2024 · 2024
Later among the works it cites.
Adaptive caching for faster video generation with diffusion transformers
Kumara Kahatapitiya, Haozhe Liu, Sen He, Ding Liu, Menglin Jia, Michael S Ryoo, and Tian Xie. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
FusedAttention
2024 · 2024
Cited alongside, same era.
HunyuanVideo
2024 · 2024
Cited alongside, same era.
Llama 3.1
2024 · 2024
Cited alongside, same era.
Mistral-7B
2024 · 2024
Cited alongside, same era.
Open-Sora
2024 · 2024
Cited alongside, same era.
Openai Triton
2024 · 2024
Cited alongside, same era.
stabilityai-stable-diffusion-2-1-base
2024 · 2024
Cited alongside, same era.
Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. 2024b · 2024
Later among the works it cites.
Openvid-1m: A large-scale high-quality dataset for text-to-video generation
Kepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan, Zhenheng Yang, Zhijie Chen, Xiang Li, Jian Yang, and Ying Tai. 2024 · 2024
Later among the works it cites.
Movie gen: A cast of media foundation models
Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, et al · 2024
Later among the works it cites.
Efficient diffusion transformer with step-wise dynamic attention mediators. In Proc. ECCV
Yifan Pu, Zhuofan Xia, Jiayi Guo, Dongchen Han, Qixiu Li, Duo Li, Yuhui Yuan, Ji Li, Yizeng Han, Shiji Song, et al · 2024
Later among the works it cites.
Unveiling Redundancy in Diffusion Transformers (DiTs): A Systematic Study
Xibo Sun, Jiarui Fang, Aoyu Li, and Jinzhe Pan. 2024 · 2024
Later among the works it cites.
Vidgen-1m: A large-scale dataset for text-to-video generation
Zhiyu Tan, Xiaomeng Yang, Luozheng Qin, and Hao Li. 2024 · 2024
Later among the works it cites.
Qihoo-t2x: An efficiency-focused diffusion transformer via proxy tokens for text-to-any-task
Jing Wang, Ao Ma, Jiasong Feng, Dawei Leng, Yuhui Yin, and Xiaodan Liang. 2024 · 2024
Later among the works it cites.
EasyAnimate: A High-Performance Long Video Generation Method based on Transformer Architecture
Jiaqi Xu, Xinyi Zou, Kunzhe Huang, Yunkuo Chen, Bo Liu, MengLi Cheng, Xing Shi, and Jun Huang. 2024 · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al · 2024
Later among the works it cites.
DiTFastAttn: Attention Compression for Diffusion Transformer Models
Zhihang Yuan, Hanling Zhang, Pu Lu, Xuefei Ning, Linfeng Zhang, Tianchen Zhao, Shengen Yan, Guohao Dai, and Yu Wang. 2024 · 2024
Later among the works it cites.
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
Haocheng Xi, Shuo Yang, Yilong Zhao, Chenfeng Xu, Muyang Li, Xiuyu Li, Yujun Lin, Han Cai, Jintao Zhang, Dacheng Li, et al · 2025
Closest in time.
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, YX Wei, Lean Wang, Zhiping Xiao, et al · 2025
Closest in time.
Fast Video Generation with Sliding Tile Attention
Peiyuan Zhang, Yongqi Chen, Runlong Su, Hangliang Ding, Ion Stoica, Zhenghong Liu, and Hao Zhang. 2025 · 2025
Closest in time.