Fetching the paper…
Reading the bibliography…
Physical reasoning remains a significant challenge for Vision-Language Models (VLMs).
Intuitive physics: the straight-down belief and its origin
Michael McCloskey, Allyson Washburn, and Linda Felch · 1983
Earlier work this paper cites.
Infants’ physical world
Renée Baillargeon · 2004
Earlier work this paper cites.
Pure reasoning in 12-month-old infants as probabilistic inference
Ernő Téglás, Edward Vul, Vittorio Girotto, Michel Gonzalez, Joshua B Tenenbaum, and Luca L Bonatti · 2011
Earlier work this paper cites.
Simulation as an engine of physical scene understanding
Peter W Battaglia, Jessica B Hamrick, and Joshua B Tenenbaum · 2013
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Galileo: Perceiving physical object properties by integrating a physics engine with deep learning
Jiajun Wu, Ilker Yildirim, Joseph J Lim, Bill Freeman, and Josh Tenenbaum · 2015
Earlier work this paper cites.
Interaction networks for learning about objects, relations and physics
Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al · 2016
Earlier work this paper cites.
Learning physical intuition of block towers by example
Adam Lerer, Sam Gross, and Rob Fergus · 2016
Earlier work this paper cites.
Intuitive physics: Current research and controversies
James R Kubricht, Keith J Holyoak, and Hongjing Lu · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Earlier work this paper cites.
Learning to see physics via visual de-animation
Jiajun Wu, Erika Lu, Pushmeet Kohli, Bill Freeman, and Josh Tenenbaum · 2017
Earlier work this paper cites.
ShapeStacks: Learning vision-based physical intuition for generalised object stacking
Oliver Groth, Fabian B Fuchs, Ingmar Posner, and Andrea Vedaldi · 2018
Earlier work this paper cites.
Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Disentangling physical dynamics from unknown factors for unsupervised video prediction
Vincent Le Guen and Nicolas Thome · 2020
Earlier work this paper cites.
Clevrer: Collision events for video representation and reasoning
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B Tenenbaum · 2020
Cited alongside, same era.
Grounding physical concepts of objects and events through dynamic visual reasoning
Z Chen, J Mao, J Wu, KKY Wong, JB Tenenbaum, and C Gan · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Cited alongside, same era.
Pip: Physical interaction prediction via mental simulation with span selection
Jiafei Duan, Samson Yu, Soujanya Poria, Bihan Wen, and Cheston Tan · 2022
Cited alongside, same era.
Exploring failure cases in multimodal reasoning about physical dynamics
Sadaf Ghaffari and Nikhil Krishnaswamy · 2024
Closest in time.
Phygrasp: generalizing robotic grasping with physics-informed large multimodal models
Dingkun Guo, Yuqi Xiang, Shuqi Zhao, Xinghao Zhu, Masayoshi Tomizuka, Mingyu Ding, and Wei Zhan · 2024
Closest in time.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Closest in time.
Think before you simulate: Symbolic reasoning to orchestrate neural computation for counterfactual question answering
Adam Ishay, Zhun Yang, Joohyung Lee, Ilgu Kang, and Dongjae Lim · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2022
Cited alongside, same era.
Mind’s eye: Grounded language model reasoning through simulation
Ruibo Liu, Jason Wei, Shixiang Shane Gu, Te-Yen Wu, Soroush Vosoughi, Claire Cui, Denny Zhou, and Andrew M Dai · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adri Garriga-Alonso, et al · 2023
Cited alongside, same era.
Zhiyuan Li, Heng Wang, Dongnan Liu, Chaoyi Zhang, Ao Ma, Jieting Long, and Weidong Cai · 2024
Closest in time.
Moka: Open-vocabulary robotic manipulation through mark-based visual prompting
Fangchen Liu, Kuan Fang, Pieter Abbeel, and Sergey Levine · 2024
Closest in time.
Zero-shot visual reasoning by vision-language models: Benchmarking and analysis
Aishik Nagar, Shantanu Jaiswal, and Cheston Tan · 2024
Closest in time.
Vision language models are blind
Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri, and Anh Totti Nguyen · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al · 2024
Closest in time.
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2024
Closest in time.
Contphy: Continuum physical concept learning and reasoning from videos
Zhicheng Zheng, Xin Yan, Zhenfang Chen, Jingzhou Wang, Qin Zhi Eddie Lim, Joshua B Tenenbaum, and Chuang Gan · 2024
Closest in time.
Physbench: Benchmarking and enhancing vision-language models for physical world understanding
Wei Chow, Jiageng Mao, Boyi Li, Daniel Seita, Vitor Campagnolo Guizilini, and Yue Wang · 2025
Closest in time.
Probing mechanical reasoning in large vision language models
Haoran Sun, Yijiang Li, Qingying Gao, Haiyun Lyu, Dezhi Luo, and Hokin Deng · 2025
Closest in time.
NL-eye: Abductive NLI for images
Mor Ventura, Michael Toker, Nitay Calderon, Zorik Gekhman, Yonatan Bitton, and Roi Reichart · 2025
Closest in time.
Maps: Advancing multi-modal reasoning in expert-level physical science
Erle Zhu, Yadi Liu, Zhe Zhang, Xujun Li, Xinjie Yu, Minlie Huang, Hongning Wang, et al · 2025
Closest in time.