Fetching the paper…
Reading the bibliography…
We present IntPhys 2, a video benchmark designed to evaluate the intuitive physics understanding of deep learning models.
Huggingface’s transformers: State-of-the-art natural language processing, 2020
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 1910
Earlier work this paper cites.
The violation-of-expectation paradigm: A conceptual overview
F. Margoni, L. Surian, and R. Baillargeon · 1939
Earlier work this paper cites.
Origins of knowledge
E. S. Spelke, K. Breinlinger, J. Macomber, and K. Jacobson · 1939
Earlier work this paper cites.
The Construction of Reality in the Child
J. Piaget · 1954
Earlier work this paper cites.
Object permanence in five-month-old infants
R. Baillargeon, E. S. Spelke, and S. Wasserman · 1985
Earlier work this paper cites.
Preferential-looking methods as tools for the study of cognition in infancy
E. S. Spelke · 1985
Earlier work this paper cites.
Object permanence in young infants: Further evidence
R. Baillargeon and J. DeVos · 1991
Earlier work this paper cites.
The development of young infants’ intuitions about support
R. Baillargeon, A. Needham, and J. Devos · 1992
Earlier work this paper cites.
Spatiotemporal continuity, smoothness of motion and object identity in infancy
E. S. Spelke, R. Kestenbaum, D. J. Simons, and D. Wein · 1995
Earlier work this paper cites.
Spatiotemporal continuity, smoothness of motion and object identity in infancy
E. S. Spelke, R. Kestenbaum, D. J. Simons, and D. Wein · 1995
Earlier work this paper cites.
Object individuation: Infants’ use of shape, size, pattern, and color
T. Wilcox · 1999
Earlier work this paper cites.
Priming infants to attend to color and pattern information in an individuation task
T. Wilcox and C. Chapa · 2004
Earlier work this paper cites.
Simulation as an engine of physical scene understanding
P. W. Battaglia, J. B. Hamrick, and J. B. Tenenbaum · 2013
Earlier work this paper cites.
Is the top object adequately supported by the bottom object? young infants’ understanding of support relations
R. Baillargeon and S. Hanko-Summers · 2014
Earlier work this paper cites.
Carla: An open urban driving simulator
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun · 2017
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi · 2017
Earlier work this paper cites.
Unrealcv: Virtual worlds for computer vision
W. Qiu, Q. Zhou, C. Chen, and A. Yuille · 2017
Earlier work this paper cites.
N. Watters, A. Tacchetti, T. Weber, R. Pascanu, P. Battaglia, and D. Zoran · 2017
Earlier work this paper cites.
World models, 2018
D. Ha and J. Schmidhuber · 2018
Earlier work this paper cites.
Intphys: A framework and benchmark for visual intuitive physics reasoning
R. Riochet, M. Y. Castro, M. Bernard, A. Lerer, R. Fergus, V. Izard, and E. Dupoux · 2018
Cited alongside, same era.
Phyre: A new benchmark for physical reasoning
A. Bakhtin, L. van der Maaten, J. Johnson, L. Gustafson, and R. Girshick · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Cited alongside, same era.
Modeling expectation violation in intuitive physics with coarse probabilistic object representations
K. Smith, L. Mei, S. Yao, J. Wu, E. Spelke, J. Tenenbaum, and T. Ullman · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, J. Gao, and Y. Choi · 2020
Cited alongside, same era.
Worldsense: A synthetic benchmark for grounded reasoning in large language models
Y. Benchekroun, M. Dervishi, M. Ibrahim, J.-B. Gaya, X. Martinet, G. Mialon, T. Scialom, E. Dupoux, D. Hupkes, and P. Vincent · 2023
Later among the works it cites.
PUG: Photorealistic and semantically controllable synthetic data for representation learning
F. Bordes, S. Shekhar, M. Ibrahim, D. Bouchacourt, P. Vincent, and A. S. Morcos · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Physion++: Evaluating physical scene understanding that requires online inference of different physical properties
H.-Y. Tung, M. Ding, Z. Chen, D. Bear, C. Gan, J. B. Tenenbaum, D. L. Yamins, J. E. Fan, and K. A. Smith · 2023
Later among the works it cites.
Videomae v2: Scaling video masked autoencoders with dual masking
L. Wang, B. Huang, Z. Zhao, Z. Tong, Y. He, Y. Wang, Y. Wang, and Y. Qiao · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
Oops! predicting unintentional action in video
D. Epstein, B. Chen, and C. Vondrick · 2020
Cited alongside, same era.
Cater: A diagnostic dataset for compositional actions & temporal reasoning
R. Girdhar and D. Ramanan · 2020
Cited alongside, same era.
Visual grounding of learned physical models
Y. Li, T. Lin, K. Yi, D. Bear, D. Yamins, J. Wu, J. Tenenbaum, and A. Torralba · 2020
Cited alongside, same era.
Occlusion resistant learning of intuitive physics from videos
R. Riochet, J. Sivic, I. Laptev, and E. Dupoux · 2020
Cited alongside, same era.
Learning to simulate complex physics with graph networks
A. Sanchez-Gonzalez, J. Godwin, T. Pfaff, R. Ying, J. Leskovec, and P. W. Battaglia · 2020
Cited alongside, same era.
Clevrer: Collision events for video representation and reasoning
K. Yi*, C. Gan*, Y. Li, P. Kohli, J. Wu, A. Torralba, and J. B. Tenenbaum · 2020
Cited alongside, same era.
Videophy: Evaluating physical commonsense for video generation
H. Bansal, Z. Lin, T. Xie, Z. Zong, M. Yarom, Y. Bitton, C. Jiang, Y. Sun, K.-W. Chang, and A. Grover · 2024
Later among the works it cites.
Revisiting feature prediction for learning visual representations from video
A. Bardes, Q. Garrido, J. Ponce, X. Chen, M. Rabbat, Y. LeCun, M. Assran, and N. Ballas · 2024
Later among the works it cites.
An introduction to vision-language modeling, 2024
F. Bordes, R. Y. Pang, A. Ajay, A. C. Li, A. Bardes, S. Petryk, O. Mañas, Z. Lin, A. Mahmoud, B. Jayaraman, M. Ibrahim, M. Hall, Y. Xiong, J. Lebensold, C. Ross, S. Jayakumar, C. Guo, D. Bouchacourt, H. Al-Tahan, K. Padthe, V. Sharma, H. Xu, X. E. Tan, M. Richards, S. Lavoie, P. Astolfi, R. A. Hemmat, J. Chen, K. Tirumala, R. Assouel, M. Moayeri, A. Talattof, K. Chaudhuri, Z. Liu, X. Chen, Q. Garrido, K. Ullrich, A. Agrawal, K. Saenko, A. Celikyilmaz, and V. Chandra · 2024
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2024
Gemini Team · 2024
Later among the works it cites.
Grasp: A novel benchmark for evaluating language grounding and situated physics understanding in multimodal language models
S. Jassim, M. Holubar, A. Richter, C. Wolff, X. Ohmer, and E. Bruni · 2024
Later among the works it cites.
Babilong: Testing the limits of llms with long context reasoning-in-a-haystack, 2024
Y. Kuratov, A. Bulatov, P. Anokhin, I. Rodkin, D. Sorokin, A. Sorokin, and M. Burtsev · 2024
Later among the works it cites.
Cosmos world foundation model platform for physical ai
N. Agarwal, A. Ali, M. Bala, Y. Balaji, E. Barker, T. Cai, P. Chattopadhyay, Y. Chen, Y. Cui, Y. Ding, et al · 2025
Closest in time.
V-jepa 2: Self-supervised video models can understand, predict and plan
M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, M. Komeili, M. Muckley, A. Rizvi, C. Roberts, K. Sinha, A. Zholus, S. Arnaud, A. Gejji, A. Martin, F. R. Hogan, D. Dugas, P. Bojanowski, V. Khalidov, P. Labatut, F. Massa, M. Szafraniec, K. Krishnakumar, Y. Li, X. Ma, S. Chandar, F. Meier, Y. LeCun, M. Rabbat, and N. Ballas · 2025
Closest in time.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin · 2025
Closest in time.
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
Q. Garrido, N. Ballas, M. Assran, A. Bardes, L. Najman, M. Rabbat, E. Dupoux, and Y. LeCun · 2025
Closest in time.
D. X. Long, H. N. Ngoc, T. Sim, H. Dao, S. Joty, K. Kawaguchi, N. F. Chen, and M.-Y. Kan · 2025
Closest in time.
Do generative video models understand physical principles?
S. Motamed, L. Culp, K. Swersky, P. Jaini, and R. Geirhos · 2025
Closest in time.
Compositional 4d dynamic scenes understanding with physics priors for video question answering
X. Wang, W. Ma, A. Wang, S. Chen, A. Kortylewski, and A. Yuille · 2025
Closest in time.