Fetching the paper…
Reading the bibliography…
Video generation has advanced rapidly, improving evaluation methods, yet assessing video's motion remains a major challenge.
“Distributed hierarchical processing in the primate cerebral cortex.”
Daniel. Felleman and David. van Essen · 1991
Earlier work this paper cites.
“Recommendation 500-10: Methodology for the subjective assessment of the quality of television pictures”
I Telecom · 2000
Earlier work this paper cites.
“Image quality assessment: from error visibility to structural similarity”
Zhou Wang, Alan Bovik, Hamid Sheikh and Eero Simoncelli · 2004
Earlier work this paper cites.
“Spearman rank correlation”
Jerrold Zar · 2005
Earlier work this paper cites.
“Influences of weather phenomena on automotive laser radar systems”
Ralph Rasshofer, Martin Spies and Hans Spies · 2011
Earlier work this paper cites.
“Biological motion processing as a hallmark of social cognition”
Marina Pavlova · 2012
Earlier work this paper cites.
“Entangling mechanical motion with microwave fields”
TA Palomaki, JD Teufel, RW Simmonds and Konrad Lehnert · 2013
Earlier work this paper cites.
“The ecology of collective behavior”
Deborah Gordon · 2014
Earlier work this paper cites.
“Likert-scale questionnaires”
Tomoko Nemoto and David Beglar · 2014
Earlier work this paper cites.
“Improved techniques for training gans”
Tim Salimans et al · 2016
Earlier work this paper cites.
“Msr-vtt: A large video description dataset for bridging video and language”
Jun Xu, Tao Mei, Ting Yao and Yong Rui · 2016
Earlier work this paper cites.
“Gans trained by a two time-scale update rule converge to a local nash equilibrium”
Martin Heusel et al · 2017
Earlier work this paper cites.
“Equilibrium propagation: Bridging the gap between energy-based models and backpropagation”
Benjamin Scellier and Yoshua Bengio · 2017
Earlier work this paper cites.
“Places: A 10 million Image Database for Scene Recognition”
Bolei Zhou et al · 2017
Earlier work this paper cites.
“Localizing moments in video with natural language”
Lisa Anne et al · 2017
Earlier work this paper cites.
“Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes”
Yu Xiang, Tanner Schmidt, Venkatraman Narayanan and Dieter Fox · 2017
Earlier work this paper cites.
“Towards accurate generative models of video: A new metric & challenges”
Thomas Unterthiner et al · 2018
Earlier work this paper cites.
“Video generation from text”
Yitong Li et al · 2018
Earlier work this paper cites.
“Learning blind video temporal consistency”
Wei-Sheng Lai et al · 2018
Earlier work this paper cites.
“A short note on the kinetics-700 human action dataset”
Joao Carreira, Eric Noland, Chloe Hillier and Andrew Zisserman · 2019
Earlier work this paper cites.
“IRC-GAN: Introspective Recurrent Convolutional GAN for Text-to-video Generation.”
Kangle Deng, Tianyi Fei, Xin Huang and Yuxin Peng · 2019
Earlier work this paper cites.
“Raft: Recurrent all-pairs field transforms for optical flow”
Zachary Teed and Jia Deng · 2020
Earlier work this paper cites.
“Tivgan: Text to image to video generation with step-by-step evolutionary generator”
Doyeon Kim, Donggyu Joo and Junmo Kim · 2020
Earlier work this paper cites.
“Hierarchical patch vae-gan: Generating diverse videos from a single sample”
Shir Gur, Sagie Benaim and Lior Wolf · 2020
Earlier work this paper cites.
“Generative pretraining from pixels”
Mark Chen et al · 2020
Earlier work this paper cites.
“Blind video temporal consistency via deep video prior”
Chenyang Lei, Yazhou Xing and Qifeng Chen · 2020
Earlier work this paper cites.
“OpenMMLab Pose Estimation Toolbox and Benchmark”, https://github.com/open-mmlab/mmpose , 2020
MMPose Contributors · 2020
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale”
Alexey Dosovitskiy et al · 2020
Earlier work this paper cites.
“Learning transferable visual models from natural language supervision”
Alec Radford et al · 2021
Earlier work this paper cites.
“Frozen in time: A joint video and image encoder for end-to-end retrieval”
Max Bain, Arsha Nagrani, Gül Varol and Andrew Zisserman · 2021
Earlier work this paper cites.
“Emerging properties in self-supervised vision transformers”
Mathilde Caron et al · 2021
Earlier work this paper cites.
“Godiva: Generating open-domain videos from natural descriptions”
Chenfei Wu et al · 2021
Earlier work this paper cites.
“Video diffusion models”
Jonathan Ho et al · 2022
Earlier work this paper cites.
“CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers”
Wenyi Hong et al · 2022
Cited alongside, same era.
“Make-a-video: Text-to-video generation without text-video data”
Uriel Singer et al · 2022
Cited alongside, same era.
“Imagen Video: High Definition Video Generation with Diffusion Models”, 2022
Jonathan Ho et al · 2022
Cited alongside, same era.
“Magicvideo: Efficient video generation with latent diffusion models”
Daquan Zhou et al · 2022
Cited alongside, same era.
“Animal kingdom: A large and diverse dataset for animal behavior understanding”
Xun Ng et al · 2022
“T2v-compbench: A comprehensive benchmark for compositional text-to-video generation”
Kaiyue Sun et al · 2024
Later among the works it cites.
“VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models”
Wenhao Wang and Yi Yang · 2024
Later among the works it cites.
“Video instruction tuning with synthetic data”
Yuanhan Zhang et al · 2024
Later among the works it cites.
“VTM-GAN: video-text matcher based generative adversarial network for generating videos from textual description”
Rayeesa Mehmood, Rumaan Bashir and Kaiser Giri · 2024
Later among the works it cites.
“VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models”, 2024
Haoxin Chen et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Conditional image-to-video generation with latent flow diffusion models”
Haomiao Ni et al · 2023
Cited alongside, same era.
“Disco: Disentangled control for referring human dance generation in real world”
Tan Wang et al · 2023
Cited alongside, same era.
“Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets”, 2023
Andreas Blattmann et al · 2023
Cited alongside, same era.
“Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video generation”
Yuanxin Liu et al · 2023
Cited alongside, same era.
“Stable fluids”
Jos Stam · 2023
Cited alongside, same era.
“Exploring video quality assessment on user generated contents from aesthetic and technical perspectives”
Haoning Wu et al · 2023
Cited alongside, same era.
“Magvit: Masked generative video transformer”
Lijun Yu et al · 2023
Cited alongside, same era.
Later among the works it cites.
“MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation”, 2024
Weimin Wang et al · 2024
Later among the works it cites.
“Motionclone: Training-free motion cloning for controllable video generation”
Pengyang Ling et al · 2024
Later among the works it cites.
Jiachen Li et al · 2024
Later among the works it cites.
“Videophy: Evaluating physical commonsense for video generation”
Hritik Bansal et al · 2024
Later among the works it cites.
“Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels”
Haoning Wu et al · 2024
Later among the works it cites.
“Grounding dino: Marrying dino with grounded pre-training for open-set object detection”
Shilong Liu et al · 2024
Later among the works it cites.
“Grounded sam: Assembling open-world models for diverse visual tasks”
Tianhe Ren et al · 2024
Later among the works it cites.
“CoTracker3: Simpler and better point tracking by pseudo-labelling real videos”
Nikita Karaev et al · 2024
Later among the works it cites.
“Sam 2: Segment anything in images and videos”
Nikhila Ravi et al · 2024
Later among the works it cites.
“Qwen2.5: A Party of Foundation Models”, 2024
Qwen Team · 2024
Later among the works it cites.
“Panda-70m: Captioning 70m videos with multiple cross-modality teachers”
Tsai-Shien Chen et al · 2024
Later among the works it cites.
“Minicpm-v: A gpt-4v level mllm on your phone”
Yuan Yao et al · 2024
Later among the works it cites.
Zhe Chen et al · 2024
Later among the works it cites.
URL: https://klingai.com/
“kling”, Accessed February 25, 2025 [Online] https://klingai.com/ , 2025 · 2025
Closest in time.
URL: https://openai.com/index/sora/
“Sora”, Accessed February 25, 2025 [Online] https://openai.com/index/sora/ , 2025 · 2025
Closest in time.
“Evaluation of text-to-video generation models: A dynamics perspective”
Mingxiang Liao et al · 2025
Closest in time.
“DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”, 2025
DeepSeek-AI · 2025
Closest in time.
“InternVideo2. 5: Empowering Video MLLMs with Long and Rich Context Modeling”
Yi Wang et al · 2025
Closest in time.
“Qwen2. 5-VL Technical Report”
Shuai Bai et al · 2025
Closest in time.
“VideoWorld: Exploring Knowledge Learning from Unlabeled Videos”
Zhongwei Ren et al · 2025
Closest in time.
“Movie Gen: A Cast of Media Foundation Models”, 2025
Adam Polyak et al · 2025
Closest in time.
URL: https://app.runwayml.com/
“Runway Gen3”, Accessed February 25, 2025 [Online] https://app.runwayml.com/ , 2025 · 2025
Closest in time.
URL: https://deepmind.google/technologies/veo/veo-2/
“Veo 2”, Accessed February 25, 2025 [Online] https://deepmind.google/technologies/veo/veo-2/ , 2025 · 2025
Closest in time.
“Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model”
Guoqing Ma et al · 2025
Closest in time.
URL: https://www.pika.art/
“Pika Labs”, Accessed February 25, 2025 [Online] https://www.pika.art/ , 2025 · 2025
Closest in time.
“Improving Video Generation with Human Feedback”
Jie Liu et al · 2025
Closest in time.
“Wan: Open and Advanced Large-Scale Video Generative Models”, 2025
Wan Team · 2025
Closest in time.