Fetching the paper…
Reading the bibliography…
We present a neural network structure, FramePack, to train next-frame (or next-frame-section) prediction models for video generation.
Retinaface: Single-stage dense face localisation in the wild, 2019b
J. Deng, J. Guo, Y. Zhou, J. Yu, I. Kotsia, and S. Zafeiriou · 1905
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow, 2020
Z. Teed and J. Deng · 2003
Earlier work this paper cites.
Rethinking attention with performers
K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Davis, A. Mohiuddin, L. Kaiser, et al · 2020
Earlier work this paper cites.
Perceptual quality assessment of smartphone photography
Y. Fang, H. Zhu, Y. Zeng, K. Ma, and Z. Wang · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Earlier work this paper cites.
Musiq: Multi-scale image quality transformer, 2021
J. Ke, Q. Wang, Y. Wang, P. Milanfar, and F. Yang · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Earlier work this paper cites.
Flexible diffusion modeling of long videos, 2022
W. Harvey, S. Naderiparizi, V. Masrani, C. Weilbach, and F. Wood · 2022
Earlier work this paper cites.
Latent video diffusion models for high-fidelity long video generation
Y. He, T. Yang, Y. Zhang, Y. Shan, and Q. Chen · 2022
Earlier work this paper cites.
Real-time intermediate flow estimation for video frame interpolation
Z. Huang, T. Zhang, W. Heng, B. Shi, and S. Zhou · 2022
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Earlier work this paper cites.
Phenaki: Variable length video generation from open domain textual description
R. Villegas, M. Babaeizadeh, P.-J. Kindermans, H. Moraldo, H. Zhang, M. T. Saffar, S. Castro, J. Kunze, and D. Erhan · 2022
Earlier work this paper cites.
Metaformer is actually what you need for vision
W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, and S. Yan · 2022
Earlier work this paper cites.
Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction
H. Cai, J. Li, M. Hu, C. Gan, and S. Han · 2023
Earlier work this paper cites.
Q-diffusion: Quantizing diffusion models
X. Li, Y. Liu, L. Lian, H. Yang, Z. Dong, D. Kang, S. Zhang, and K. Keutzer · 2023
Earlier work this paper cites.
Latent consistency models: Synthesizing high-resolution images with few-step inference
S. Luo, Y. Tan, L. Huang, J. Li, and H. Zhao · 2023
Earlier work this paper cites.
Freenoise: Tuning-free longer video diffusion via noise rescheduling
H. Qiu, M. Xia, Y. Zhang, Y. He, X. Wang, Y. Shan, and Z. Liu · 2023
Earlier work this paper cites.
Consistency models
Y. Song, P. Dhariwal, M. Chen, and I. Sutskever · 2023
Earlier work this paper cites.
Art•v: Auto-regressive text-to-video generation with diffusion models
W. Weng, R. Feng, Y. Wang, Q. Dai, C. Wang, D. Yin, Z. Zhao, K. Qiu, J. Bao, Y. Yuan, C. Luo, Y. Zhang, and Z. Xiong · 2023
Cited alongside, same era.
Nuwa-xl: Diffusion over diffusion for extremely long video generation
S. Yin, C. Wu, H. Yang, J. Wang, X. Wang, M. Ni, Z. Yang, L. Li, S. Liu, F. Yang, et al · 2023
Cited alongside, same era.
Talc: Time-aligned captions for multi-scene text-to-video generation
H. Bansal, Y. Bitton, M. Yarom, I. Szpektor, A. Grover, and K.-W. Chang · 2024
Cited alongside, same era.
M. Cai, X. Cun, X. Li, W. Liu, Z. Zhang, Y. Zhang, Y. Shan, and X. Yue · 2024
Cited alongside, same era.
Diffusion models are real-time game engines, 2024
D. Valevski, Y. Leviathan, M. Arar, and S. Fruchter · 2024
Later among the works it cites.
3d reconstruction with spatial memory
H. Wang and L. Agapito · 2024
Later among the works it cites.
Mind the time: Temporally-controlled multi-event video generation
Z. Wu, A. Siarohin, W. Menapace, I. Skorokhodov, Y. Fang, V. Chordia, I. Gilitschenski, and S. Tulyakov · 2024
Later among the works it cites.
Synchronized video storytelling: Generating video narrations with structured storyline
D. Yang, C. Zhan, Z. Wang, B. Wang, T. Ge, B. Zheng, and Q. Jin · 2024
Later among the works it cites.
Wonderjourney: Going from anywhere to everywhere
H.-X. Yu, H. Duan, J. Hur, K. Sargent, M. Rubinstein, W. T. Freeman, F. Cole, D. Sun, N. Snavely, J. Wu, and C. Herrmann · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Gao, J. Shi, H. Zhang, C. Wang, and J. Xiao · 2024
Cited alongside, same era.
Ltx-video: Realtime video latent diffusion
Y. HaCohen, N. Chiprut, B. Brazowski, D. Shalem, D. Moshe, E. Richardson, E. Levin, G. Shiran, N. Zabari, O. Gordon, et al · 2024
Cited alongside, same era.
Streamingt2v: Consistent, dynamic, and extendable long video generation from text
R. Henschel, L. Khachatryan, D. Hayrapetyan, H. Poghosyan, V. Tadevosyan, Z. Wang, S. Navasardyan, and H. Shi · 2024
Cited alongside, same era.
Slowfast-vgen: Slow-fast learning for action-driven long video generation, 2024
Y. Hong, B. Liu, M. Wu, Y. Zhai, K.-W. Chang, L. Li, K. Lin, C.-C. Lin, J. Wang, Z. Yang, Y. Wu, and L. Wang · 2024
Cited alongside, same era.
Storyagent: Customized storytelling video generation via multi-agent collaboration
P. Hu, J. Jiang, J. Chen, M. Han, S. Liao, X. Chang, and X. Liang · 2024
Cited alongside, same era.
VBench: Comprehensive benchmark suite for video generative models
Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, Y. Wang, X. Chen, L. Wang, D. Lin, Y. Qiao, and Z. Liu · 2024
Cited alongside, same era.
Pyramidal flow matching for efficient video generative modeling, 2024
Y. Jin, Z. Sun, N. Li, K. Xu, K. Xu, H. Jiang, N. Zhuang, Q. Huang, Y. Song, Y. Mu, and Z. Lin · 2024
Cited alongside, same era.
Hunyuanvideo: A systematic framework for large video generative models
W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, et al · 2024
Cited alongside, same era.
Later among the works it cites.
J. Zhang, H. Huang, P. Zhang, J. Wei, J. Zhu, and J. Chen · 2024
Later among the works it cites.
FlexTok: Resampling images into 1d token sequences of flexible length, 2025
R. Bachmann, J. Allardice, D. Mizrahi, E. Fini, O. F. Kar, E. Amirloo, A. El-Nouby, A. Zamir, and A. Dehghan · 2025
Closest in time.
One-minute video generation with test-time training, 2025
K. Dalal, D. Koceja, G. Hussein, J. Xu, Y. Zhao, Y. Song, S. Han, K. C. Cheung, J. Kautz, C. Guestrin, T. Hashimoto, S. Koyejo, Y. Choi, Y. Sun, and X. Wang · 2025
Closest in time.
Long-context autoregressive video modeling with next-frame prediction, 2025
Y. Gu, W. Mao, and M. Z. Shou · 2025
Closest in time.
Svdquant: Absorbing outliers by low-rank components for 4-bit diffusion models
M. Li*, Y. Lin*, Z. Zhang*, T. Cai, X. Li, J. Guo, E. Xie, C. Meng, J.-Y. Zhu, and S. Han · 2025
Closest in time.
Freelong: Training-free long video generation with spectralblend temporal attention
Y. Lu, Y. Liang, L. Zhu, and Y. Yang · 2025
Closest in time.
Gen3c: 3d-informed world-consistent video generation with precise camera control, 2025
X. Ren, T. Shen, J. Huang, H. Ling, Y. Lu, M. Nimier-David, T. Müller, A. Keller, S. Fidler, and J. Gao · 2025
Closest in time.
History-guided video diffusion
K. Song, B. Chen, M. Simchowitz, Y. Du, R. Tedrake, and V. Sitzmann · 2025
Closest in time.
Sparse videogen: Accelerating video diffusion transformers with spatial-temporal sparsity
H. Xi, S. Yang, Y. Zhao, C. Xu, M. Li, X. Li, Y. Lin, H. Cai, J. Zhang, D. Li, et al · 2025
Closest in time.
Training-free and adaptive sparse attention for efficient long video generation, 2025
Y. Xia, S. Ling, F. Fu, Y. Wang, H. Li, X. Xiao, and B. Cui · 2025
Closest in time.
Worldmem: Long-term consistent world simulation with memory, 2025
Z. Xiao, Y. Lan, Y. Zhou, W. Ouyang, S. Yang, Y. Zeng, and X. Pan · 2025
Closest in time.
Wonderworld: Interactive 3d scene generation from a single image
H.-X. Yu, H. Duan, C. Herrmann, W. T. Freeman, and J. Wu · 2025
Closest in time.
VBench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness
D. Zheng, Z. Huang, H. Liu, K. Zou, Y. He, F. Zhang, Y. Zhang, J. He, W.-S. Zheng, Y. Qiao, and Z. Liu · 2025
Closest in time.
Z. Zhou, Y. Yang, Y. Yang, T. He, H. Peng, K. Qiu, Q. Dai, L. Qiu, C. Luo, and L. Liu · 2025
Closest in time.