Fetching the paper…
Reading the bibliography…
Large language models (LLMs) demonstrate impressive performance on a wide variety of tasks, but they often struggle with tasks that require multi-step reasoning or goal-directed planning.
Cognitive maps in rats and men
Edward C Tolman · 1948
Earlier work this paper cites.
The logic theory machine–a complex information processing system
Allen Newell and Herbert Simon · 1956
Earlier work this paper cites.
Report on a general problem solving program
Allen Newell, John C Shaw, and Herbert A Simon · 1959
Earlier work this paper cites.
The functional equivalence of problem solving skills
Herbert A Simon · 1975
Earlier work this paper cites.
Controlled and automatic human information processing: I. detection, search, and attention
Walter Schneider and Richard M Shiffrin · 1977
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test
Patricia A Carpenter, Marcel A Just, and Peter Shell · 1990
Earlier work this paper cites.
The empirical case for two systems of reasoning
Steven A Sloman · 1996
Earlier work this paper cites.
Cognitive planning in humans: neuropsychological, neuroanatomical and neuropharmacological perspectives
Adrian M Owen · 1997
Earlier work this paper cites.
Conflict monitoring versus selection-for-action in anterior cingulate cortex
Matthew Botvinick, Leigh E Nystrom, Kate Fissell, Cameron S Carter, and Jonathan D Cohen · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
An integrative theory of prefrontal cortex function
Earl K Miller and Jonathan D Cohen · 2001
Earlier work this paper cites.
Reward representations and reward-related learning in the human brain: insights from neuroimaging
John P O’doherty · 2004
Earlier work this paper cites.
Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control
Nathaniel D Daw, Yael Niv, and Peter Dayan · 2005
Earlier work this paper cites.
Determining the neural substrates of goal-directed learning in the human brain
Vivian V Valentin, Anthony Dickinson, and John P O’Doherty · 2007
Earlier work this paper cites.
Orbitofrontal cortex and its contribution to decision-making
Jonathan D Wallis · 2007
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
Expectancy-related changes in firing of dopamine neurons depend on orbitofrontal cortex
Yuji K Takahashi, Matthew R Roesch, Robert C Wilson, Kathy Toreson, Patricio O’donnell, Yael Niv, and Geoffrey Schoenbaum · 2011
Earlier work this paper cites.
Model-based reinforcement learning as cognitive search: neurocomputational theories
Nathaniel D Daw · 2012
Earlier work this paper cites.
Neural representations of events arise from temporal community structure
Anna C Schapiro, Timothy T Rogers, Natalia I Cordova, Nicholas B Turk-Browne, and Matthew M Botvinick · 2013
Earlier work this paper cites.
From conflict management to reward-based decision making: actors and critics in primate medial frontal cortex
Massimo Silvetti, William Alexander, Tom Verguts, and Joshua W Brown · 2014
Cited alongside, same era.
A map for social navigation in the human brain
Rita Morais Tavares, Avi Mendelsohn, Yael Grossman, Christian Hamilton Williams, Matthew Shapiro, Yaacov Trope, and Daniela Schiller · 2015
Cited alongside, same era.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Cited alongside, same era.
Human orbitofrontal cortex represents a cognitive map of state space
Nicolas W Schuck, Ming Bo Cai, Robert C Wilson, and Yael Niv · 2016
Cited alongside, same era.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Cited alongside, same era.
What is a cognitive map? organizing knowledge for flexible behavior
Timothy EJ Behrens, Timothy H Muller, James CR Whittington, Shirley Mark, Alon B Baram, Kimberly L Stachenfeld, and Zeb Kurth-Nelson · 2018
Planning in the brain
Marcelo G Mattar and Máté Lengyel · 2022
Later among the works it cites.
Plansformer: Generating symbolic plans using transformers
Vishal Pallagani, Bharath Muppasani, Keerthiram Murugesan, Francesca Rossi, Lior Horesh, Biplav Srivastava, Francesco Fabiano, and Andrea Loreggia · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Closest in time.
Grounding large language models in interactive environments with online reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Offline replay supports planning in human reinforcement learning
I Momennejad, A R Otto, N D Daw, and K A Norman · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Prefrontal cortex as a meta-reinforcement learning system
Jane X Wang, Zeb Kurth-Nelson, Dharshan Kumaran, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Demis Hassabis, and Matthew Botvinick · 2018
Cited alongside, same era.
Reinforcement learning, fast and slow
Matthew Botvinick, Sam Ritter, Jane X Wang, Zeb Kurth-Nelson, Charles Blundell, and Demis Hassabis · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Recurrent independent mechanisms
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf · 2019
Cited alongside, same era.
Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer · 2023
Closest in time.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch · 2023
Closest in time.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jian, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D Hwang, et al · 2023
Closest in time.
Hosein Hasanbeig, Hiteshi Sharma, Leo Betthauser, Felipe Vieira Frujeri, and Ida Momennejad · 2023
Closest in time.
Yilun Kong, Jingqing Ruan, Yihong Chen, Bin Zhang, Tianpeng Bao, Shiwei Shi, Guoqing Du, Xiaoru Hu, Hangyu Mao, Ziyue Li, et al · 2023
Closest in time.
Dissociating language and thought in large language models: a cognitive perspective
Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko · 2023
Closest in time.
Evaluating cognitive maps in large language models with cogeval: No emergent planning
Ida Momennejad, Hosein Hasanbeig, Felipe Vieira Frujeri, Hiteshi Sharma, Robert Osazuwa Ness, Nebojsa Jojic, Hamid Palangi, and Jonathan Larson · 2023
Closest in time.
Tptu: Task planning and tool usage of large language model-based ai agents
Jingqing Ruan, Yihong Chen, Bin Zhang, Zhiwei Xu, Tianpeng Bao, Guoqing Du, Shiwei Shi, Hangyu Mao, Xingyu Zeng, and Rui Zhao · 2023
Closest in time.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Noah Shinn, Beck Labash, and Ashwin Gopinath · 2023
Closest in time.
On the planning abilities of large language models–a critical investigation
Karthik Valmeekam, Matthew Marquez, Sarath Sreedharan, and Subbarao Kambhampati · 2023
Closest in time.
Zihao Wang, Shaofei Cai, Anji Liu, Xiaojian Ma, and Yitao Liang · 2023
Closest in time.
Emergent analogical reasoning in large language models
Taylor Webb, Keith J Holyoak, and Hongjing Lu · 2023
Closest in time.
Embodied task planning with large language models
Zhenyu Wu, Ziwei Wang, Xiuwei Xu, Jiwen Lu, and Haibin Yan · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Closest in time.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning
Yuexiang Zhai, Hao Bai, Zipeng Lin, Jiayi Pan, Shengbang Tong, Yifei Zhou, Alane Suhr, Saining Xie, Yann LeCun, Yi Ma, et al · 2024
Closest in time.
Archer: Training language model agents via hierarchical multi-turn rl
Yifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine, and Aviral Kumar · 2024
Closest in time.