Fetching the paper…
Reading the bibliography…
Diffusion Transformers (DiTs) dominate video generation but their high computational cost severely limits real-world applicability, usually requiring tens of minutes to generate a few seconds of video even on high-performance GPUs.
Image quality metrics: Psnr vs. ssim
Horé, A. and Ziou, D · 2010
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
Learning fast algorithms for linear transforms using butterfly factorizations
Dao, T., Gu, A., Eichhorn, M., Rudra, A., and Ré, C · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Tillet, P., Kung, H.-T., and Cox, D. D · 2019
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Earlier work this paper cites.
Vivit: A video vision transformer
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., and Schmid, C · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Flow straight and fast: Learning to generate and transfer data with rectified flow
Liu, X., Gong, C., and Liu, Q · 2022
Earlier work this paper cites.
SDEdit: Guided image synthesis and editing with stochastic differential equations
Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S · 2022
Earlier work this paper cites.
Metaformer is actually what you need for vision
Yu, W., Luo, M., Zhou, P., Si, C., Zhou, Y., Wang, X., Feng, J., and Yan, S · 2022
Earlier work this paper cites.
Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction
Cai, H., Li, J., Hu, M., Gan, C., and Han, S · 2023
Earlier work this paper cites.
Lm-infinite: Simple on-the-fly length generalization for large language models
Han, C., Wang, Q., Xiong, W., Chen, Y., Ji, H., and Wang, S · 2023
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2023
Cited alongside, same era.
Vbench: Comprehensive benchmark suite for video generative models, 2023
Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., Wang, Y., Chen, X., Wang, L., Lin, D., Qiao, Y., and Liu, Z · 2023
Cited alongside, same era.
Streamdiffusion: A pipeline-level solution for real-time interactive generation
Kodaira, A., Xu, C., Hazama, T., Yoshimoto, T., Ohno, K., Mitsuhori, S., Sugano, S., Cho, H., Liu, Z., and Keutzer, K · 2023
Cited alongside, same era.
Q-diffusion: Quantizing diffusion models
Distrifusion: Distributed parallel inference for high-resolution diffusion models
Li, M., Cai, T., Cao, J., Zhang, Q., Cai, H., Bai, J., Jia, Y., Liu, M.-Y., Li, K., and Han, S · 2024
Later among the works it cites.
Looking backward: Streaming video-to-video translation with feature banks
Liang, F., Kodaira, A., Xu, C., Tomizuka, M., Keutzer, K., and Marculescu, D · 2024
Later among the works it cites.
Fastercache: Training-free video diffusion model acceleration with high quality
Lv, Z., Si, C., Song, J., Yang, Z., Qiao, Y., Liu, Z., and Wong, K.-Y. K · 2024
Later among the works it cites.
Deepcache: Accelerating diffusion models for free
Ma, X., Fang, G., and Wang, X · 2024
Later among the works it cites.
Longvu: Spatiotemporal adaptive compression for long video-language understanding, 2024
Shen, X., Xiong, Y., Zhao, C., Wu, L., Chen, J., Zhu, C., Liu, Z., Xiao, F., Varadarajan, B., Bordes, F., Liu, Z., Xu, H., Kim, H. J., Soran, B., Krishnamoorthi, R., Elhoseiny, M., and Chandra, V · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, X., Liu, Y., Lian, L., Yang, H., Dong, Z., Kang, D., Zhang, S., and Keutzer, K · 2023
Cited alongside, same era.
Latent consistency models: Synthesizing high-resolution images with few-step inference
Luo, S., Tan, Y., Huang, L., Li, J., and Zhao, H · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Consistency models
Song, Y., Dhariwal, P., Chen, M., and Sutskever, I · 2023
Cited alongside, same era.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Cited alongside, same era.
Sparsetir: Composable abstractions for sparse compilation in deep learning, 2023
Ye, Z., Lai, R., Shao, J., Chen, T., and Ceze, L · 2023
Cited alongside, same era.
Zheng, N., Jiang, H., Zhang, Q., Han, Z., Yang, Y., Ma, L., Yang, F., Zhang, C., Qiu, L., Yang, M., and Zhou, L · 2023
Cited alongside, same era.
Condition-aware neural network for controlled image generation
Cai, H., Li, M., Zhang, Q., Liu, M.-Y., and Han, S · 2024
Cited alongside, same era.
Later among the works it cites.
Quest: Query-aware sparsity for efficient long-context llm inference, 2024
Tang, J., Zhao, Y., Zhu, K., Xiao, G., Kasikci, B., and Han, S · 2024
Later among the works it cites.
Sana: Efficient high-resolution image synthesis with linear diffusion transformers
Xie, E., Chen, J., Chen, J., Cai, H., Tang, H., Lin, Y., Zhang, Z., Li, M., Zhu, L., Lu, Y., et al · 2024
Later among the works it cites.
Ditfastattn: Attention compression for diffusion transformer models, 2024
Yuan, Z., Zhang, H., Lu, P., Ning, X., Zhang, L., Zhao, T., Yan, S., Dai, G., and Wang, Y · 2024
Later among the works it cites.
Zhang, J., Huang, H., Zhang, P., Wei, J., Zhu, J., and Chen, J · 2024
Later among the works it cites.
Pixart-sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation
Chen, J., Ge, C., Xie, E., Wu, Y., Yao, L., Ren, X., Wang, Z., Luo, P., Lu, H., and Li, Z · 2025
Closest in time.
Svdquant: Absorbing outliers by low-rank components for 4-bit diffusion models
Li*, M., Lin*, Y., Zhang*, Z., Cai, T., Li, X., Guo, J., Xie, E., Meng, C., Zhu, J.-Y., and Han, S · 2025
Closest in time.
Li, Y., Jiang, H., Zhang, C., Wu, Q., Luo, X., Ahn, S., Abdi, A. H., Li, D., Gao, J., Yang, Y., et al · 2025
Closest in time.
Akvq-vl: Attention-aware kv cache adaptive 2-bit quantization for vision-language models, 2025
Su, Z., Shen, W., Li, L., Chen, Z., Wei, H., Yu, H., and Yuan, K · 2025
Closest in time.
Wan: Open and advanced large-scale video generative models
Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.-W., Chen, D., Yu, F., Zhao, H., Yang, J., Zeng, J., et al · 2025
Closest in time.
Xattention: Block sparse attention with antidiagonal scoring
Xu, R., Xiao, G., Huang, H., Guo, J., and Han, S · 2025
Closest in time.
Flashinfer: Efficient and customizable attention engine for llm inference serving, 2025
Ye, Z., Chen, L., Lai, R., Lin, W., Zhang, Y., Wang, S., Chen, T., Kasikci, B., Grover, V., Krishnamurthy, A., and Ceze, L · 2025
Closest in time.