Fetching the paper…
Reading the bibliography…
AI video generation is undergoing a revolution, with quality and realism advancing rapidly.
Possible principles underlying the transformation of sensory messages
Horace B Barlow et al · 1961
Earlier work this paper cites.
Telling more than we can know: Verbal reports on mental processes
Richard E Nisbett and Timothy D Wilson · 1977
Earlier work this paper cites.
Curvilinear motion in the absence of external forces: Naive beliefs about the motion of objects
Michael McCloskey, Alfonso Caramazza, and Bert Green · 1980
Earlier work this paper cites.
Intuitive physics
Michael McCloskey · 1983
Earlier work this paper cites.
Perception of partly occluded objects in infancy
Philip J Kellman and Elizabeth S Spelke · 1983
Earlier work this paper cites.
The origins of physical knowledge
Elizabeth S Spelke · 1988
Earlier work this paper cites.
Origins of knowledge
Elizabeth S Spelke, Karen Breinlinger, Janet Macomber, and Kristen Jacobson · 1992
Earlier work this paper cites.
Spatiotemporal continuity, smoothness of motion and object identity in infancy
Elizabeth S Spelke, Roberta Kestenbaum, Daniel J Simons, and Debra Wein · 1995
Earlier work this paper cites.
The acquisition of physical knowledge in infancy: A summary in eight lessons
Renée Baillargeon · 2002
Earlier work this paper cites.
A theory of causal learning in children: causal maps and bayes nets
Alison Gopnik, Clark Glymour, David M Sobel, Laura E Schulz, Tamar Kushnir, and David Danks · 2004
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli · 2004
Earlier work this paper cites.
A theory of cortical responses
Karl Friston · 2005
Earlier work this paper cites.
The perception of causality in infancy
Rebecca Saxe and Susan Carey · 2006
Earlier work this paper cites.
Image quality metrics: Psnr vs. ssim
Alain Hore and Djemel Ziou · 2010
Earlier work this paper cites.
How to grow a mind: Statistics, structure, and abstraction
Joshua B Tenenbaum, Charles Kemp, Thomas L Griffiths, and Noah D Goodman · 2011
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
Pulkit Agrawal, Ashvin V Nair, Pieter Abbeel, Jitendra Malik, and Sergey Levine · 2016
Earlier work this paper cites.
A compositional object-based approach to learning physical dynamics
Michael B Chang, Tomer Ullman, Antonio Torralba, and Joshua B Tenenbaum · 2016
Earlier work this paper cites.
Intuitive physics: Current research and controversies
James R Kubricht, Keith J Holyoak, and Hongjing Lu · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Earlier work this paper cites.
Generalisation in humans and deep neural networks
Robert Geirhos, Carlos RM Temme, Jonas Rauber, Heiko H Schütt, Matthias Bethge, and Felix A Wichmann · 2018
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2018
Earlier work this paper cites.
IntPhys: A framework and benchmark for visual intuitive physics reasoning
Ronan Riochet, Mario Ynocente Castro, Mathieu Bernard, Adam Lerer, Rob Fergus, Véronique Izard, and Emmanuel Dupoux · 2018
Cited alongside, same era.
Towards accurate generative models of video: A new metric & challenges
Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly · 2018
Cited alongside, same era.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Cited alongside, same era.
Cophy: Counterfactual learning of physical dynamics
Fabien Baradel, Natalia Neverova, Julien Mille, Greg Mori, and Christian Wolf · 2019
Cited alongside, same era.
Clevrer: Collision events for video representation and reasoning
Sora: OpenAI’s Multimodal Agent
OpenAI · 2024
Later among the works it cites.
Meta Movie Gen: AI-powered movie generation
Meta AI · 2024
Later among the works it cites.
Lumiere: A space-time diffusion model for video generation, 2024
Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Herrmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Guanghui Liu, Amit Raj, Yuanzhen Li, Michael Rubinstein, Tomer Michaeli, Oliver Wang, Deqing Sun, Tali Dekel, and Inbar Mosseri · 2024
Later among the works it cites.
VideoPoet: A large language model for zero-shot video generation
Dan Kondratyuk, Lijun Yu, Xiuye Gu, Jose Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Joshua V. Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Xuan Yang, Hartwig Adam, Ming-Hsuan Yang, Irfan Essa, Huisheng Wang, David A Ross, Bryan Seybold, and Lu Jiang · 2024
Later among the works it cites.
Physion++: Evaluating physical scene understanding that requires online inference of different physical properties
Hsiao-Yu Tung, Mingyu Ding, Zhenfang Chen, Daniel Bear, Chuang Gan, Josh Tenenbaum, Dan Yamins, Judith Fan, and Kevin Smith · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B Tenenbaum · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning · 2019
Cited alongside, same era.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann · 2020
Cited alongside, same era.
Craft: A benchmark for causal reasoning about forces and interactions
Tayfun Ates, M Samil Atesoglu, Cagatay Yigit, Ilker Kesen, Mert Kobas, Erkut Erdem, Aykut Erdem, Tilbe Goksun, and Deniz Yuret · 2020
Cited alongside, same era.
Esprit: Explaining solutions to physical reasoning tasks
Nazneen Fatema Rajani, Rui Zhang, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, and Dragomir Radev · 2020
Cited alongside, same era.
Physion: Evaluating physical prediction from vision in humans and machines, 2021
Daniel M. Bear, Elias Wang, Damian Mrowca, Felix J. Binder, Hsiao-Yu Fish Tung, R. T. Pramod, Cameron Holdaway, Sirui Tao, Kevin Smith, Fan-Yun Sun, Li Fei-Fei, Nancy Kanwisher, Joshua B. Tenenbaum, Daniel L. K. Yamins, and Judith E. Fan · 2021
Cited alongside, same era.
Tiered reasoning for intuitive physics: Toward verifiable commonsense language understanding
Shane Storks, Qiaozi Gao, Yichi Zhang, and Joyce Chai · 2021
Cited alongside, same era.
What tool representation, intuitive physics, and action have in common: The brain’s first-person physics engine
Jason Fischer and Bradford Z Mahon · 2021
Cited alongside, same era.
Later among the works it cites.
Videophy: Evaluating physical commonsense for video generation, 2024
Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang, and Aditya Grover · 2024
Later among the works it cites.
Towards world simulator: Crafting physical commonsense-based benchmark for video generation, 2024
Fanqing Meng, Jiaqi Liao, Xinyu Tan, Wenqi Shao, Quanfeng Lu, Kaipeng Zhang, Yu Cheng, Dianqi Li, Yu Qiao, and Ping Luo · 2024
Later among the works it cites.
Physgame: Uncovering physical commonsense violations in gameplay videos
Meng Cao, Haoran Tang, Haoze Zhao, Hangyu Guo, Jiaheng Liu, Ge Zhang, Ruyang Liu, Qiang Sun, Ian Reid, and Xiaodan Liang · 2024
Later among the works it cites.
LLMPhy: Complex physical reasoning using large language models and world models
Anoop Cherian, Radu Corcodel, Siddarth Jain, and Diego Romeres · 2024
Later among the works it cites.
Videoscore: Building automatic metrics to simulate fine-grained human feedback for video generation
Xuan He, Dongfu Jiang, Ge Zhang, Max Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, et al · 2024
Later among the works it cites.
On the content bias in frechet video distance
Songwei Ge, Aniruddha Mahapatra, Gaurav Parmar, Jun-Yan Zhu, and Jia-Bin Huang · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team Google: Petko Georgiev and 1133 other authors · 2024
Later among the works it cites.
Vision language models are blind
Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri, and Anh Totti Nguyen · 2024
Later among the works it cites.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Later among the works it cites.
Scaling LLM test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
Veo2: Our state-of-the-art video generation model
DeepMind · 2025
Closest in time.
Cosmos world foundation model platform for physical AI
Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al · 2025
Closest in time.
Generative physical AI in vision: A survey
Daochang Liu, Junyu Zhang, Anh-Dung Dinh, Eunbyung Park, Shichao Zhang, and Chang Xu · 2025
Closest in time.
Visual cognition in multimodal large language models
Luca M Schulze Buschoff, Elif Akata, Matthias Bethge, and Eric Schulz · 2025
Closest in time.
Blending simulation and abstraction for physical reasoning
Felix A Sosa, Samuel J Gershman, and Tomer D Ullman · 2025
Closest in time.
Inference-time scaling for diffusion models beyond scaling denoising steps
Nanye Ma, Shangyuan Tong, Haolin Jia, Hexiang Hu, Yu-Chuan Su, Mingda Zhang, Xuan Yang, Yandong Li, Tommi Jaakkola, Xuhui Jia, et al · 2025
Closest in time.