Fetching the paper…
Reading the bibliography…
As general-purpose tools, Large Language Models (LLMs) must often reason about everyday physical environments.
Construct validity in psychological tests
Lee J Cronbach and Paul E Meehl · 1955
Earlier work this paper cites.
Grounded in the world: Developmental origins of the embodied mind
Esther Thelen · 2000
Earlier work this paper cites.
Folk Physics for Apes: The Chimpanzee’s theory of how the world works
D. J. Povinelli · 2003
Earlier work this paper cites.
The concept of validity
Denny Borsboom, Gideon J Mellenbergh, and Jaap Van Heerden · 2004
Earlier work this paper cites.
Mechanisms of theory formation in young children
Alison Gopnik and Laura Schulz · 2004
Earlier work this paper cites.
Representing causation
Phillip Wolff · 2007
Earlier work this paper cites.
Metaphors we live by
George Lakoff and Mark Johnson · 2008
Earlier work this paper cites.
Embodiment and the inner life: cognition and consciousness in the space of possible minds
Murray Shanahan · 2010
Earlier work this paper cites.
Intuitive physical reasoning about occluded objects by inexperienced chicks
Cinzia Chiandetti and Giorgio Vallortigara · 2011
Earlier work this paper cites.
How to grow a mind: Statistics, structure, and abstraction
Joshua B Tenenbaum, Charles Kemp, Thomas L Griffiths, and Noah D Goodman · 2011
Earlier work this paper cites.
Computational models of intuitive physics
Peter Battaglia, Tomer Ullman, Joshua Tenenbaum, Adam Sanborn, Kenneth Forbus, Tobias Gerstenberg, and David Lagnado · 2012
Earlier work this paper cites.
Evaluation in artificial intelligence: from task-oriented to ability-oriented measurement
José Hernández-Orallo · 2017
Earlier work this paper cites.
Intuitive physics: Current research and controversies
James R Kubricht, Keith J Holyoak, and Hongjing Lu · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Earlier work this paper cites.
Interpretable counting for visual question answering
Alexander Trott, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Unity: A general platform for intelligent agents
Arthur Juliani · 2018
Earlier work this paper cites.
Virtualhome: Simulating household activities via programs
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba · 2018
Earlier work this paper cites.
Different physical intuitions exist between tasks, not domains
Kevin A Smith, Peter W Battaglia, and Edward Vul · 2018
Earlier work this paper cites.
Modeling human intuitions about liquid flow with particle-based simulation
Christopher J Bates, Ilker Yildirim, Joshua B Tenenbaum, and Peter Battaglia · 2019
Earlier work this paper cites.
The Animal-AI environment: Training and testing animal-like artificial cognition
Benjamin Beyret, José Hernández-Orallo, Lucy Cheke, Marta Halina, Murray Shanahan, and Matthew Crosby · 2019
Earlier work this paper cites.
The Animal-AI Olympics
Matthew Crosby, Benjamin Beyret, and Marta Halina · 2019
Earlier work this paper cites.
Shane Storks, Qiaozi Gao, and Joyce Y Chai · 2019
Earlier work this paper cites.
Climbing towards nlu: On meaning, form, and understanding in the age of data
Emily M Bender and Alexander Koller · 2020
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al · 2020
Earlier work this paper cites.
The animal-ai testbed and competition
Matthew Crosby, Benjamin Beyret, Murray Shanahan, José Hernández-Orallo, Lucy Cheke, and Marta Halina · 2020
Cited alongside, same era.
Commonsense reasoning for natural language processing
Maarten Sap, Vered Shwartz, Antoine Bosselut, Yejin Choi, and Dan Roth · 2020
Cited alongside, same era.
Artificial intelligence and the common sense of animals
Murray Shanahan, Matthew Crosby, Benjamin Beyret, and Lucy Cheke · 2020
Cited alongside, same era.
Prost: Physical reasoning of objects through space and time
Stéphane Aroca-Ouellette, Cory Paik, Alessandro Roncone, and Katharina Kann · 2021
Cited alongside, same era.
Benchmarking the spectrum of agent capabilities
Danijar Hafner · 2021
Cited alongside, same era.
Cognitive architectures for language agents
Theodore R Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L Griffiths · 2023
Later among the works it cites.
Macgyver: Are large language models creative problem solvers?
Yufei Tian, Abhilasha Ravichander, Lianhui Qin, Ronan Le Bras, Raja Marjieh, Nanyun Peng, Yejin Choi, Thomas L Griffiths, and Faeze Brahman · 2023
Later among the works it cites.
Behind the magic, merlim: Multi-modal evaluation benchmark for large image-language models
Andrés Villa, Juan Carlos León Alcázar, Alvaro Soto, and Bernard Ghanem · 2023
Later among the works it cites.
Animal-ai 3: What’s new & why you should care
Konstantinos Voudouris, Ibrahim Alhas, Wout Schellaert, Matthew Crosby, Joel Holmes, John Burden, Niharika Chaubey, Niall Donnelly, Matishalin Patel, Marta Halina, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Melanie Mitchell · 2021
Cited alongside, same era.
Piglet: Language grounding through neuro-symbolic interaction in a 3d world
Rowan Zellers, Ari Holtzman, Matthew Peters, Roozbeh Mottaghi, Aniruddha Kembhavi, Ali Farhadi, and Yejin Choi · 2021
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al · 2022
Cited alongside, same era.
Not a number: Identifying instance features for capability-oriented evaluation
Ryan Burnell, John Burden, Danaja Rutar, Konstantinos Voudouris, Lucy Cheke, and José Hernández-Orallo · 2022
Cited alongside, same era.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar · 2022
Cited alongside, same era.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch · 2022
Cited alongside, same era.
Vima: General robot manipulation with multimodal prompts
Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan · 2022
Cited alongside, same era.
Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, et al · 2023
Later among the works it cites.
Evaluating ai evaluation: Perils and prospects
John Burden · 2024
Closest in time.
Have we built machines that think like people?
Luca M Schulze Buschoff, Elif Akata, Matthias Bethge, and Eric Schulz · 2024
Closest in time.
Chatgpt in action: Analyzing its use in software development
Arifa Islam Champa, Md Fazle Rabbi, Costain Nachuma, and Minhaz F Zibran · 2024
Closest in time.
S-agents: self-organizing agents in open-ended environment
Jiaqi Chen, Yuxian Jiang, Jiachen Lu, and Li Zhang · 2024
Closest in time.
Mathematical capabilities of chatgpt
Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Petersen, and Julius Berner · 2024
Closest in time.
The development of human causal learning and reasoning
Mariel K Goddu and Alison Gopnik · 2024
Closest in time.
A survey on large language model-based game agents
Sihao Hu, Tiansheng Huang, Fatih Ilhan, Selim Tekin, Gaowen Liu, Ramana Kompella, and Ling Liu · 2024
Closest in time.
Serwan Jassim, Mario Holubar, Annika Richter, Cornelius Wolff, Xenia Ohmer, and Elia Bruni · 2024
Closest in time.
Learning to localize objects improves spatial reasoning in visual-llms
Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S Ryoo, and Tsung-Yu Lin · 2024
Closest in time.
General interaction battery: Simple object navigation and affordances (gibsona)
Danaja Rutar, Lucy Gaia Cheke, José Hernández-Orallo, Alva Markelius, and Wout Schellaert · 2024
Closest in time.
Swarmbrain: Embodied agent for real-time strategy game starcraft ii via large language models
Xiao Shao, Weifu Jiang, Fei Zuo, and Mengqing Liu · 2024
Closest in time.
Regal: Refactoring programs to discover generalizable abstractions
Elias Stengel-Eskin, Archiki Prasad, and Mohit Bansal · 2024
Closest in time.
Investigating object permanence in deep reinforcement learning agents
Konstantinos Voudouris, Jason Darwin Liu, Natasza Siwinska, Wout Schellaert, and Lucy G Cheke · 2024
Closest in time.
Spring: Studying papers and reasoning to play games
Yue Wu, So Yeon Min, Shrimai Prabhumoye, Yonatan Bisk, Russ R Salakhutdinov, Amos Azaria, Tom M Mitchell, and Yuanzhi Li · 2024
Closest in time.
Language models meet world models: Embodied experiences enhance language models
Jiannan Xiang, Tianhua Tao, Yi Gu, Tianmin Shu, Zirui Wang, Zichao Yang, and Zhiting Hu · 2024
Closest in time.
Benchmark data contamination of large language models: A survey
Cheng Xu, Shuhao Guan, Derek Greene, M Kechadi, et al · 2024
Closest in time.
Adarefiner: Refining decisions of language models with adaptive feedback
Wanpeng Zhang and Zongqing Lu · 2024
Closest in time.
Hierarchical auto-organizing system for open-ended multi-agent navigation
Zhonghan Zhao, Kewei Chen, Dongxu Guo, Wenhao Chai, Tian Ye, Yanting Zhang, and Gaoang Wang · 2024
Closest in time.