Fetching the paper…
Reading the bibliography…
The ability to plan a course of action that achieves a desired state of affairs has long been considered a core competence of intelligent agents and has been an integral part of AI research since its inception.
The one-machine sequencing problem
Jacques Carlier · 1982
Earlier work this paper cites.
Sokoban is pspace-complete
Joseph Culberson · 1997
Earlier work this paper cites.
International planning competition
IPC · 1998
Earlier work this paper cites.
Pddl-the planning domain definition language
Drew McDermott, Malik Ghallab, Adele E. Howe, Craig A. Knoblock, Ashwin Ram, Manuela M. Veloso, Daniel S. Weld, and David E. Wilkins · 1998
Earlier work this paper cites.
Course of action generation for cyber security using classical planning
Mark S Boddy, Johnathan Gohde, Thomas Haigh, and Steven A Harp · 2005
Earlier work this paper cites.
The fast downward planning system
Malte Helmert · 2006
Earlier work this paper cites.
Engineering benchmarks for planning: the domains used in the deterministic part of ipc-4
Jörg Hoffmann, Stefan Edelkamp, Sylvie Thiébaux, Roman Englert, Frederico Liporace, and Sebastian Trüg · 2006
Earlier work this paper cites.
Automated Planning and Acting
Malik Ghallab, Dana S. Nau, and Paolo Traverso · 2016
Earlier work this paper cites.
GPT3-to-plan: Extracting plans from text using GPT-3
Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, et al · 2022
Earlier work this paper cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch · 2022
Earlier work this paper cites.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al · 2022
Earlier work this paper cites.
Large Language Models are Zero-Shot Reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Limitations of language models in arithmetic and symbolic induction
Jing Qian, Hong Wang, Zekun Li, Shiyang Li, and Xifeng Yan · 2022
Earlier work this paper cites.
Behavior: Benchmark for everyday household activities in virtual, interactive, and ecological environments
Sanjana Srivastava, Chengshu Li, Michael Lingelbach, Roberto Martín-Martín, Fei Xia, Kent Elliott Vainio, Zheng Lian, Cem Gokmen, Shyamal Buch, Karen Liu, et al · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
Llm+ p: Empowering large language models with optimal planning proficiency
Bo Liu, Yuqian Jiang, Xiaohan Zhang, Qiang Liu, Shiqi Zhang, Joydeep Biswas, and Peter Stone · 2023
Cited alongside, same era.
Ai agents that matter, 2024
Sayash Kapoor, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, and Arvind Narayanan · 2024
Closest in time.
Thought of search: Planning with language models through the lens of efficiency, 2024
Michael Katz, Harsha Kokel, Kavitha Srinivas, and Shirin Sohrabi · 2024
Closest in time.
Guides: Reasoning, 2024
OpenAI · 2024
Closest in time.
Introducing openai o1-preview, 2024
OpenAI · 2024
Closest in time.
Openai o1 system card
OpenAI · 2024
Closest in time.
Mathematical discoveries from program search with large language models
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Man Luo, Shrinidhi Kumbhar, Mihir Parmar, Neeraj Varshney, Pratyay Banerjee, Somak Aditya, Chitta Baral, et al · 2023
Cited alongside, same era.
Data contamination through the lens of time
Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, and Samuel Dooley · 2023
Cited alongside, same era.
On the planning abilities of large language models–a critical investigation
Karthik Valmeekam, Matthew Marquez, Sarath Sreedharan, and Subbarao Kambhampati · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao · 2023
Cited alongside, same era.
o1 gets it right almost always, 2024
Noam Brown · 2024
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras, et al · 2024
Cited alongside, same era.
Ban warnings fly as users dare to probe the “thoughts” of openai’s latest model
Benj Edwards · 2024
Cited alongside, same era.
Can large language models reason and plan?
Subbarao Kambhampati · 2024
Cited alongside, same era.
Kaya Stechly, Karthik Valmeekam, and Subbarao Kambhampati · 2024
Closest in time.
On the self-verification limitations of large language models on reasoning and planning tasks
Kaya Stechly, Karthik Valmeekam, and Subbarao Kambhampati · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong · 2024
Closest in time.
Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Karthik Valmeekam, Matthew Marquez, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati · 2024
Closest in time.
Travelplanner: A benchmark for real-world planning with language agents
Jian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, and Yu Su · 2024
Closest in time.
log. probs returned by openai’s api are *incredibly* unstable, 2023
Tan Zhi Xuan · 2024
Closest in time.
Natural plan: Benchmarking llms on natural language planning
Huaixiu Steven Zheng, Swaroop Mishra, Hugh Zhang, Xinyun Chen, Minmin Chen, Azade Nova, Le Hou, Heng-Tze Cheng, Quoc V Le, Ed H Chi, et al · 2024
Closest in time.