Fetching the paper…
Reading the bibliography…
Future Frame Synthesis (FFS), the task of generating subsequent video frames from context, represents a core challenge in machine intelligence and a cornerstone for developing predictive world models.
Learning internal representations by error propagation, 1985
David E Rumelhart, Geoffrey E Hinton, Ronald J Williams, et al · 1985
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Recognizing human actions: a local svm approach
Christian Schuldt, Ivan Laptev, and Barbara Caputo · 2004
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli · 2004
Earlier work this paper cites.
Dtvnet+: A high-resolution scenic dataset for dynamic time-lapse video generation
Jiangning Zhang, Chao Xu, Yong Liu, and Yunliang Jiang · 2008
Earlier work this paper cites.
Pedestrian detection: An evaluation of the state of the art
Piotr Dollar, Christian Wojek, Bernt Schiele, and Pietro Perona · 2011
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
Towards understanding action recognition
Hueihan Jhuang, Juergen Gall, Silvia Zuffi, Cordelia Schmid, and Michael J Black · 2013
Earlier work this paper cites.
The sjtu 4k video sequence dataset
Li Song, Xun Tang, Wei Zhang, Xiaokang Yang, and Pingjian Xia · 2013
Earlier work this paper cites.
From actemes to action: A strongly-supervised representation for detailed action understanding
Weiyu Zhang, Menglong Zhu, and Konstantinos G Derpanis · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Earlier work this paper cites.
Generating images with perceptual similarity metrics based on deep networks
Alexey Dosovitskiy and Thomas Brox · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei · 2016
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun · 2016
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks
Tianfan Xue, Jiajun Wu, Katherine Bouman, and Bill Freeman · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Unsupervised learning of disentangled representations from video
Emily L Denton et al · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
Chelsea Finn and Sergey Levine · 2017
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Beyond face rotation: Global and local perception gan for photorealistic and identity preserving frontal view synthesis
Rui Huang, Shu Zhang, Tianyu Li, and Ran He · 2017
Earlier work this paper cites.
Video pixel networks
Nal Kalchbrenner, Aäron Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Video frame synthesis using deep voxel flow
Ziwei Liu, Raymond A Yeh, Xiaoou Tang, Yiming Liu, and Aseem Agarwala · 2017
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning
William Lotter, Gabriel Kreiman, and David Cox · 2017
Earlier work this paper cites.
Predicting deeper into the future of semantic segmentation
Pauline Luc, Natalia Neverova, Camille Couprie, Jakob Verbeek, and Yann LeCun · 2017
Earlier work this paper cites.
Video frame interpolation via adaptive separable convolution
Simon Niklaus, Long Mai, and Feng Liu · 2017
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alex Sorkine-Hornung, and Luc Van Gool · 2017
Earlier work this paper cites.
Deep learning for precipitation nowcasting: A benchmark and a new model
Xingjian Shi, Zhihan Gao, Leonard Lausen, Hao Wang, Dit-Yan Yeung, Wai-kin Wong, and Wang-chun Woo · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Learning to generate long-term future via hierarchical prediction
Ruben Villegas, Jimei Yang, Yuliang Zou, Sungryull Sohn, Xunyu Lin, and Honglak Lee · 2017
Earlier work this paper cites.
The pose knows: Video forecasting by generating pose futures
Jacob Walker, Kenneth Marino, Abhinav Gupta, and Martial Hebert · 2017
Earlier work this paper cites.
Predrnn: Recurrent neural networks for predictive learning using spatiotemporal lstms
Yunbo Wang, Mingsheng Long, Jianmin Wang, Zhifeng Gao, and Philip S Yu · 2017
Earlier work this paper cites.
Stochastic variational video prediction
Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H. Campbell, and Sergey Levine · 2018
Earlier work this paper cites.
The perception-distortion tradeoff
Yochai Blau and Tomer Michaeli · 2018
Earlier work this paper cites.
Deep video generation, prediction and completion of human action sequences
Haoye Cai, Chunyan Bai, Yu-Wing Tai, and Chi-Keung Tang · 2018
Earlier work this paper cites.
Stochastic video generation with a learned prior
Emily Denton and Rob Fergus · 2018
Earlier work this paper cites.
Shapestacks: Learning vision-based physical intuition for generalised object stacking
Oliver Groth, Fabian B Fuchs, Ingmar Posner, and Andrea Vedaldi · 2018
Earlier work this paper cites.
Few-shot human motion prediction via meta-learning
Liang-Yan Gui, Yu-Xiong Wang, Deva Ramanan, and José MF Moura · 2018
Earlier work this paper cites.
Video prediction with appearance and motion conditions
Yunseok Jang, Gunhee Kim, and Yale Song · 2018
Earlier work this paper cites.
Super slomo: High quality estimation of multiple intermediate frames for video interpolation
Huaizu Jiang, Deqing Sun, Varun Jampani, Ming-Hsuan Yang, Erik Learned-Miller, and Jan Kautz · 2018
Earlier work this paper cites.
Stochastic adversarial video prediction
Alex X Lee, Richard Zhang, Frederik Ebert, Pieter Abbeel, Chelsea Finn, and Sergey Levine · 2018
Earlier work this paper cites.
Future frame prediction for anomaly detection–a new baseline
Wen Liu, Weixin Luo, Dongze Lian, and Shenghua Gao · 2018
Earlier work this paper cites.
Sdc-net: Video prediction using spatially-displaced convolution
Fitsum A Reda, Guilin Liu, Kevin J Shih, Robert Kirby, Jon Barker, David Tarjan, Andrew Tao, and Bryan Catanzaro · 2018
Earlier work this paper cites.
Mocogan: Decomposing motion and content for video generation
Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz · 2018
Earlier work this paper cites.
Video-to-video synthesis
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, and Bryan Catanzaro · 2018
Earlier work this paper cites.
Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks
Wei Xiong, Wenhan Luo, Lin Ma, Wei Liu, and Jiebo Luo · 2018
Earlier work this paper cites.
Predcnn: Predictive learning with cascade convolutions
Ziru Xu, Yunbo Wang, Mingsheng Long, Jianmin Wang, and M KLiss · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
A short note on the kinetics-700 human action dataset
Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman · 2019
Earlier work this paper cites.
Improved conditional vrnns for video prediction
Lluis Castrejon, Nicolas Ballas, and Aaron Courville · 2019
Earlier work this paper cites.
D2-city: a large-scale dashcam video dataset of diverse traffic scenarios
Zhengping Che, Guangyu Li, Tracy Li, Bo Jiang, Xuefeng Shi, Xinsheng Zhang, Ying Lu, Guobin Wu, Yan Liu, and Jieping Ye · 2019
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
Sudeep Dasari, Frederik Ebert, Stephen Tian, Suraj Nair, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh, Sergey Levine, and Chelsea Finn · 2019
Earlier work this paper cites.
Predicting future frames using retrospective cycle gan
Yong-Hoon Kwon and Min-Gyu Park · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Eidetic 3d LSTM: A model for video prediction and beyond
Yunbo Wang, Lu Jiang, Ming-Hsuan Yang, Li-Jia Li, Mingsheng Long, and Li Fei-Fei · 2019
Earlier work this paper cites.
Video enhancement with task-oriented flow
Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei, and William T Freeman · 2019
Earlier work this paper cites.
Compositional video prediction
Yufei Ye, Maneesh Singh, Abhinav Gupta, and Shubham Tulsiani · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Earlier work this paper cites.
Channel attention is all you need for video frame interpolation
Myungsub Choi, Heewon Kim, Bohyung Han, Ning Xu, and Kyoung Mu Lee · 2020
Cited alongside, same era.
Stochastic latent residual video prediction
Jean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier, and Patrick Gallinari · 2020
Cited alongside, same era.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2020
Cited alongside, same era.
Disentangling physical dynamics from unknown factors for unsupervised video prediction
Vincent Le Guen and Nicolas Thome · 2020
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Tell me what happened: Unifying text-guided video completion via multimodal masked video generation
Tsu-Jui Fu, Licheng Yu, Ning Zhang, Cheng-Yang Fu, Jong-Chyi Su, William Yang Wang, and Sean Bell · 2023
Later among the works it cites.
Vip3d: End-to-end visual trajectory prediction via 3d agent queries
Junru Gu, Chenxu Hu, Tianyuan Zhang, Xuanyao Chen, Yilun Wang, Yue Wang, and Hang Zhao · 2023
Later among the works it cites.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Later among the works it cites.
Flavr: Flow-agnostic video representations for fast frame interpolation
Tarun Kalluri, Deepak Pathak, Manmohan Chandraker, and Du Tran · 2023
Later among the works it cites.
On efficient transformer-based image pre-training for low-level vision
Wenbo Li, Xin Lu, Shengju Qian, and Jiangbo Lu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Probabilistic future prediction for video scene understanding
Anthony Hu, Fergal Cotter, Nikhil Mohan, Corina Gurau, and Alex Kendall · 2020
Cited alongside, same era.
Training generative adversarial networks with limited data
Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila · 2020
Cited alongside, same era.
Softmax splatting for video frame interpolation
Simon Niklaus and Feng Liu · 2020
Cited alongside, same era.
A review on deep learning techniques for video prediction
Sergiu Oprea, Pablo Martinez-Gonzalez, Alberto Garcia-Garcia, John Alejandro Castro-Vargas, Sergio Orts-Escolano, Jose Garcia-Rodriguez, and Antonis Argyros · 2020
Cited alongside, same era.
Deep learning for vision-based prediction: A survey
Amir Rasouli · 2020
Cited alongside, same era.
Raft: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng · 2020
Cited alongside, same era.
Flow matching for generative modeling
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le · 2023
Later among the works it cites.
Meta-auxiliary learning for future depth prediction in videos
Huan Liu, Zhixiang Chi, Yuanhao Yu, Yang Wang, Jun Chen, and Jin Tang · 2023
Later among the works it cites.
Wayformer: Motion forecasting via simple & efficient attention networks
Nigamaa Nayakanti, Rami Al-Rfou, Aurick Zhou, Kratarth Goel, Khaled S Refaat, and Benjamin Sapp · 2023
Later among the works it cites.
Conditional image-to-video generation with latent flow diffusion models
Haomiao Ni, Changhao Shi, Kai Li, Sharon X Huang, and Martin Renqiang Min · 2023
Later among the works it cites.
Biformer: Learning bilateral motion estimation via bilateral transformer for 4k video frame interpolation
Junheum Park, Jintae Kim, and Chang-Su Kim · 2023
Later among the works it cites.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman · 2023
Later among the works it cites.
Convnets match vision transformers at scale
Samuel L Smith, Andrew Brock, Leonard Berrada, and Soham De · 2023
Later among the works it cites.
Moso: Decomposing motion, scene and object for video prediction
Mingzhen Sun, Weining Wang, Xinxin Zhu, and Jing Liu · 2023
Later among the works it cites.
Openstl: A comprehensive benchmark of spatio-temporal predictive learning
Cheng Tan, Siyuan Li, Zhangyang Gao, Wenfei Guan, Zedong Wang, Zicheng Liu, Lirong Wu, and Stan Z Li · 2023
Later among the works it cites.
Object-centric video prediction via decoupling of object dynamics and interactions
Angel Villar-Corrales, Ismail Wahdan, and Sven Behnke · 2023
Later among the works it cites.
Slotformer: Unsupervised visual dynamics simulation with object-centric models
Ziyi Wu, Nikita Dvornik, Klaus Greff, Thomas Kipf, and Animesh Garg · 2023
Later among the works it cites.
A survey on video diffusion models
Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, and Yu-Gang Jiang · 2023
Later among the works it cites.
Diffusion probabilistic modeling for video generation
Ruihan Yang, Prakhar Srivastava, and Stephan Mandt · 2023
Later among the works it cites.
Video prediction by efficient transformers
Xi Ye and Guillaume-Alexandre Bilodeau · 2023
Later among the works it cites.
Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory
Shengming Yin, Chenfei Wu, Jian Liang, Jie Shi, Houqiang Li, Gong Ming, and Nan Duan · 2023
Later among the works it cites.
I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models
Shiwei Zhang, Jiayu Wang, Yingya Zhang, Kang Zhao, Hangjie Yuan, Zhiwu Qin, Xiang Wang, Deli Zhao, and Jingren Zhou · 2023
Later among the works it cites.
Mmvp: Motion-matrix-based video prediction
Yiqi Zhong, Luming Liang, Ilya Zharkov, and Ulrich Neumann · 2023
Later among the works it cites.
Sequential modeling enables scalable learning for large vision models
Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan L Yuille, Trevor Darrell, Jitendra Malik, and Alexei A Efros · 2024
Closest in time.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh · 2024
Closest in time.
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation
Chi-Lam Cheang, Guangzeng Chen, Ya Jing, Tao Kong, Hang Li, Yifeng Li, Yuxiao Liu, Hongtao Wu, Jiafeng Xu, Yichu Yang, et al · 2024
Closest in time.
Sportsslomo: A new benchmark and baselines for human-centric video frame interpolation
Jiaben Chen and Huaizu Jiang · 2024
Closest in time.
Memflow: Optical flow estimation and prediction with memory
Qiaole Dong and Yanwei Fu · 2024
Closest in time.
Video prediction models as rewards for reinforcement learning
Alejandro Escontrela, Ademi Adeniji, Wilson Yan, Ajay Jain, Xue Bin Peng, Ken Goldberg, Youngwoon Lee, Danijar Hafner, and Pieter Abbeel · 2024
Closest in time.
WorldGPT: Empowering LLM as multimodal world model
Zhiqi Ge, Hongzhe Huang, Mingze Zhou, Juncheng Li, Guoming Wang, Siliang Tang, and Yueting Zhuang · 2024
Closest in time.
Factorizing text-to-video generation by explicit image conditioning
Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Duval, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin, Devi Parikh, and Ishan Misra · 2024
Closest in time.
Seer: Language instructed video prediction with latent diffusion models
Xianfan Gu, Chuan Wen, Weirui Ye, Jiaming Song, and Yang Gao · 2024
Closest in time.
Sparsectrl: Adding sparse controls to text-to-video diffusion models
Yuwei Guo, Ceyuan Yang, Anyi Rao, Maneesh Agrawala, Dahua Lin, and Bo Dai · 2024
Closest in time.
Vbench: Comprehensive benchmark suite for video generative models
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al · 2024
Closest in time.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Closest in time.
Peekaboo: Interactive video generation via masked-diffusion
Yash Jain, Anshul Nasery, Vibhav Vineet, and Harkirat Behl · 2024
Closest in time.
Videopoet: A large language model for zero-shot video generation
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming Chang Chiu, et al · 2024
Closest in time.
Hunyuanvideo: A systematic framework for large video generative models
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al · 2024
Closest in time.
Mosca: Dynamic gaussian fusion from casual videos via 4d motion scaffolds
Jiahui Lei, Yijia Weng, Adam Harley, Leonidas Guibas, and Kostas Daniilidis · 2024
Closest in time.
A survey on long video generation: Challenges, methods, and prospects
Chengxuan Li, Di Huang, Zeyu Lu, Yang Xiao, Qingqi Pei, and Lei Bai · 2024
Closest in time.
Movideo: Motion-aware video generation with diffusion model
Jingyun Liang, Yuchen Fan, Kai Zhang, Radu Timofte, Luc Van Gool, and Rakesh Ranjan · 2024
Closest in time.
Vdt: General-purpose video diffusion transformers via mask modeling
Haoyu Lu, Guoxing Yang, Nanyi Fei, Yuqi Huo, Zhiwu Lu, Ping Luo, and Mingyu Ding · 2024
Closest in time.
Video diffusion models: A survey
Andrew Melnik, Michal Ljubljanac, Cong Lu, Qi Yan, Weiming Ren, and Helge Ritter · 2024
Closest in time.
Advancing auto-regressive continuation for video frames
Ruibo Ming, Jingwei Wu, Zhewei Huang, Zhuoxuan Ju, Jianming Hu, Lihui Peng, and Shuchang Zhou · 2024
Closest in time.
Multi-modal auto-regressive modeling via visual tokens
Tianshuo Peng, Zuchao Li, Lefei Zhang, Hai Zhao, Ping Wang, and Bo Du · 2024
Closest in time.
Decouple content and motion for conditional image-to-video generation
Cuifeng Shen, Yulu Gan, Chen Chen, Xiongwei Zhu, Lele Cheng, Tingting Gao, and Jinzhi Wang · 2024
Closest in time.
Vidgen-1m: A large-scale dataset for text-to-video generation
Zhiyu Tan, Xiaomeng Yang, Luozheng Qin, and Hao Li · 2024
Closest in time.
Art-v: Auto-regressive text-to-video generation with diffusion models
Wenming Weng, Ruoyu Feng, Yanhui Wang, Qi Dai, Chunyu Wang, Dacheng Yin, Zhiyuan Zhao, Kai Qiu, Jianmin Bao, Yuhui Yuan, et al · 2024
Closest in time.
Dynamicrafter: Animating open-domain images with video diffusion priors
Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang, Ying Shan, and Tien-Tsin Wong · 2024
Closest in time.
Magicanimate: Temporally consistent human image animation using diffusion model
Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou · 2024
Closest in time.
Generalized predictive model for autonomous driving
Jiazhi Yang, Shenyuan Gao, Yihang Qiu, Li Chen, Tianyu Li, Bo Dai, Kashyap Chitta, Penghao Wu, Jia Zeng, Ping Luo, et al · 2024
Closest in time.
Iam-vfi: Interpolate any motion for video frame interpolation with motion complexity map
Kihwan Yoon, Yong Han Kim, Sungjei Kim, and Jinwoo Jeong · 2024
Closest in time.
Language model beats diffusion - tokenizer is key to visual generation
Lijun Yu, Jose Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A Ross, and Lu Jiang · 2024
Closest in time.
VFIMamba: Video frame interpolation with state space models
Guozhen Zhang, Chunxu Liu, Yutao Cui, Xiaotong Zhao, Kai Ma, and Limin Wang · 2024
Closest in time.
Gaussianprediction: Dynamic 3d gaussian prediction for motion extrapolation and free view synthesis
Boming Zhao, Yuan Li, Ziyu Sun, Lin Zeng, Yujun Shen, Rui Ma, Yinda Zhang, Hujun Bao, and Zhaopeng Cui · 2024
Closest in time.
Is sora a world simulator? a comprehensive survey on general world models and beyond
Zheng Zhu, Xiaofeng Wang, Wangbo Zhao, Chen Min, Nianchen Deng, Min Dou, Yuqi Wang, Botian Shi, Kai Wang, Chi Zhang, et al · 2024
Closest in time.
Cosmos world foundation model platform for physical ai
Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al · 2025
Closest in time.
Go-with-the-flow: Motion-controllable video diffusion models using real-time warped noise
Ryan Burgert, Yuancheng Xu, Wenqi Xian, Oliver Pilarski, Pascal Clausen, Mingming He, Li Ma, Yitong Deng, Lingxiao Li, Mohsen Mousavi, Michael Ryoo, Paul Debevec, and Ning Yu · 2025
Closest in time.
Skyreels-v2: Infinite-length film generative model
Guibin Chen, Dixuan Lin, Jiangping Yang, Chunze Lin, Juncheng Zhu, Mingyuan Fan, Hao Zhang, Sheng Chen, Zheng Chen, Chengchen Ma, et al · 2025
Closest in time.
Skyreels-a2: Compose anything in video diffusion transformers
Zhengcong Fei, Debang Li, Di Qiu, Jiahua Wang, Yikun Dou, Rui Wang, Jingtao Xu, Mingyuan Fan, Guibin Chen, Yang Li, et al · 2025
Closest in time.
Photorealistic video generation with diffusion models
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Fei-Fei Li, Irfan Essa, Lu Jiang, and José Lezama · 2025
Closest in time.
Step-video-ti2v technical report: A state-of-the-art text-driven image-to-video generation model
Haoyang Huang, Guoqing Ma, Nan Duan, Xing Chen, Changyi Wan, Ranchen Ming, Tianyu Wang, Bo Wang, Zhiying Lu, Aojie Li, et al · 2025
Closest in time.
Vace: All-in-one video creation and editing
Zeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang, Yulin Pan, and Yu Liu · 2025
Closest in time.
Step-video-t2v technical report: The practice, challenges, and future of video foundation model
Guoqing Ma, Haoyang Huang, Kun Yan, Liangyu Chen, Nan Duan, Shengming Yin, Changyi Wan, Ranchen Ming, Xiaoniu Song, Xing Chen, et al · 2025
Closest in time.
Magi-1: Autoregressive video generation at scale, 2025
Sand-AI · 2025
Closest in time.
Seaweed-7b: Cost-effective training of video generation foundation model
Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, et al · 2025
Closest in time.
Diffusion models are real-time game engines
Dani Valevski, Yaniv Leviathan, Moab Arar, and Shlomi Fruchter · 2025
Closest in time.
VibiDSampler: Enhancing video interpolation using bidirectional diffusion sampler
Serin Yang, Taesung Kwon, and Jong Chul Ye · 2025
Closest in time.
Magic 1-for-1: Generating one minute video clips within one minute
Hongwei Yi, Shitong Shao, Tian Ye, Jiantong Zhao, Qingyu Yin, Michael Lingelbach, Li Yuan, Yonghong Tian, Enze Xie, and Daquan Zhou · 2025
Closest in time.
Dropletvideo: A dataset and approach to explore integral spatio-temporal consistent video generation
Runze Zhang, Guoguang Du, Xiaochuan Li, Qi Jia, Liang Jin, Lu Liu, Jingjing Wang, Cong Xu, Zhenhua Guo, Yaqian Zhao, et al · 2025
Closest in time.
Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness
Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Yuanhan Zhang, Jingwen He, Wei-Shi Zheng, Yu Qiao, et al · 2025
Closest in time.
Vistorybench: Comprehensive benchmark suite for story visualization
Cailin Zhuang, Ailin Huang, Wei Cheng, Jingwei Wu, Yaoqi Hu, Jiaqi Liao, Zhewei Huang, Hongyuan Wang, Xinyao Liao, Weiwei Cai, et al · 2025
Closest in time.
Survey on ai-generated media detection: From non-mllm to mllm
Yueying Zou, Peipei Li, Zekun Li, Huaibo Huang, Xing Cui, Xuannan Liu, Chenghanyu Zhang, and Ran He · 2025
Closest in time.