Fetching the paper…
Reading the bibliography…
We investigate the emergence of intuitive physics understanding in general-purpose deep neural network models trained to predict masked regions in natural videos.
Core knowledge
Elizabeth S. Spelke · 1935
Earlier work this paper cites.
The violation-of-expectation paradigm: A conceptual overview
Francesco Margoni, Luca Surian, and Renée Baillargeon · 1939
Earlier work this paper cites.
Origins of knowledge
Elizabeth S. Spelke, Karen Breinlinger, Janet Macomber, and Kristen Jacobson · 1939
Earlier work this paper cites.
The Construction of Reality in the Child
Jean Piaget · 1954
Earlier work this paper cites.
Object permanence in five-month-old infants
Renee Baillargeon, Elizabeth S Spelke, and Stanley Wasserman · 1985
Earlier work this paper cites.
Preferential-looking methods as tools for the study of cognition in infancy
Elizabeth S Spelke · 1985
Earlier work this paper cites.
Mind children: The future of robot and human intelligence
Hans Moravec · 1988
Earlier work this paper cites.
Object Permanence in Young Infants: Further Evidence
Renee Baillargeon and Julie DeVos · 1991
Earlier work this paper cites.
The development of young infants’ intuitions about support
Renée Baillargeon, Amy Needham, and Julie Devos · 1992
Earlier work this paper cites.
Infants’ sensitivity to effects of gravity on visible object motion
In Kyeong Kim and Elizabeth S Spelke · 1992
Earlier work this paper cites.
Physical reasoning in infancy
Renee Baillargeon · 1995
Earlier work this paper cites.
Spatiotemporal continuity, smoothness of motion and object identity in infancy
Elizabeth S. Spelke, Roberta Kestenbaum, Daniel J. Simons, and Debra Wein · 1995
Earlier work this paper cites.
Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects
Rajesh PN Rao and Dana H Ballard · 1999
Earlier work this paper cites.
Object individuation: Infants’ use of shape, size, pattern, and color
Teresa Wilcox · 1999
Earlier work this paper cites.
The origin of concepts
Susan Carey · 2000
Earlier work this paper cites.
Recognizing impossible object relations: intuitions about support in chimpanzees (pan troglodytes)
Trix Cacchione and Horst Krist · 2004
Earlier work this paper cites.
Priming infants to attend to color and pattern information in an individuation task
Teresa Wilcox and Catherine Chapa · 2004
Earlier work this paper cites.
Raising the level: orangutans use water as a tool
Natacha Mendes, Daniel Hanus, and Josep Call · 2007
Earlier work this paper cites.
Core knowledge
Elizabeth S. Spelke and Katherine D. Kinzler · 2007
Earlier work this paper cites.
Innate Ideas Revisited: For a Principle of Persistence in Infants’ Physical Reasoning
Renée Baillargeon · 2008
Earlier work this paper cites.
Rooks use stones to raise the water level to reach a floating worm
Christopher David Bird and Nathan John Emery · 2009
Earlier work this paper cites.
What laboratory research has told us about dolphin cognition
Louis M Herman · 2010
Cited alongside, same era.
New caledonian crows reason about hidden causal agents
Alex H Taylor, Rachael Miller, and Russell D Gray · 2012
Cited alongside, same era.
Core knowledge of object, number, and geometry: A comparative and neural approach
Giorgio Vallortigara · 2012
Cited alongside, same era.
Simulation as an engine of physical scene understanding
Peter W. Battaglia, Jessica B. Hamrick, and Joshua B. Tenenbaum · 2013
Cited alongside, same era.
Whatever next? predictive brains, situated agents, and the future of cognitive science
Andy Clark · 2013
Cited alongside, same era.
The predictive mind
Jakob Hohwy · 2013
Cited alongside, same era.
IntPhys 2019: A Benchmark for Visual Intuitive Physics Understanding
Ronan Riochet, Mario Ynocente Castro, Mathieu Bernard, Adam Lerer, Rob Fergus, Véronique Izard, and Emmanuel Dupoux · 2021
Later among the works it cites.
Roformer: enhanced transformer with rotary position embedding. corr abs/2104.09864 (2021)
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Later among the works it cites.
Saycam: A large, longitudinal audiovisual dataset recorded from the infant’s perspective
Jessica Sullivan, Michelle Mei, Andrew Perfors, Erica Wojcik, and Michael C Frank · 2021
Later among the works it cites.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
Yann LeCun · 2022
Later among the works it cites.
Intuitive physics learning in a deep-learning model inspired by developmental psychology
Luis S. Piloto, Ari Weinstein, Peter Battaglia, and Matthew Botvinick · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Is the top object adequately supported by the bottom object? young infants’ understanding of support relations
Renée Baillargeon and Stephanie Hanko-Summers · 2014
Cited alongside, same era.
Object permanence in marine mammals using the violation of expectation procedure
Rebecca Singer and Elizabeth Henderson · 2015
Cited alongside, same era.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Cited alongside, same era.
Learning physical intuition of block towers by example
Adam Lerer, Sam Gross, and Rob Fergus · 2016
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Cited alongside, same era.
Mind games: Game engines as an architecture for intuitive physics
Tomer D Ullman, Elizabeth Spelke, Peter Battaglia, and Joshua B Tenenbaum · 2017
Cited alongside, same era.
Benchmarking progress to infant-level physical reasoning in ai
Luca Weihs, Amanda Rose Yuile, Renée Baillargeon, Cynthia Fisher, Gary Marcus, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2022
Later among the works it cites.
Worldsense: A synthetic benchmark for grounded reasoning in large language models
Youssef Benchekroun, Megi Dervishi, Mark Ibrahim, Jean-Baptiste Gaya, Xavier Martinet, Grégoire Mialon, Thomas Scialom, Emmanuel Dupoux, Dieuwke Hupkes, and Pascal Vincent · 2023
Later among the works it cites.
Videomae v2: Scaling video masked autoencoders with dual masking
Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Later among the works it cites.
Videophy: Evaluating physical commonsense for video generation
Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang, and Aditya Grover · 2024
Later among the works it cites.
Revisiting feature prediction for learning visual representations from video
Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mido Assran, and Nicolas Ballas · 2024
Later among the works it cites.
Video generation models as world simulators, 2024
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh · 2024
Later among the works it cites.
Grasp: A novel benchmark for evaluating language grounding and situated physics understanding in multimodal language models
Serwan Jassim, Mario Holubar, Annika Richter, Cornelius Wolff, Xenia Ohmer, and Elia Bruni · 2024
Later among the works it cites.
How far is video generation from world model: A physical law perspective
Bingyi Kang, Yang Yue, Rui Lu, Zhijie Lin, Yang Zhao, Kaixin Wang, Gao Huang, and Jiashi Feng · 2024
Later among the works it cites.
Bria Long, Violet Xiang, Stefan Stojanov, Robert Z. Sparks, Zi Yin, Grace E. Keene, Alvin W. M. Tan, Steven Y. Feng, Chengxu Zhuang, Virginia A. Marchman, Daniel L. K. Yamins, and Michael C. Frank · 2024
Later among the works it cites.
Gpt-4 technical report, 2024
OpenAI · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin · 2024
Later among the works it cites.
Newborn chickens generate invariant object representations at the onset of visual object experience
Justin N Wood · 2024
Later among the works it cites.
Do generative video models learn physical principles from watching videos?
Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, and Robert Geirhos · 2025
Closest in time.