Fetching the paper…
Reading the bibliography…
LLMs have immense potential for generating plans, transforming an initial world state into a desired goal state.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Step-by-step: Separating planning from realization in neural data-to-text generation
Amit Moryossef, Yoav Goldberg, and Ido Dagan. 2019 · 1904
Earlier work this paper cites.
Elements of a theory of human problem solving
Allen Newell, John Calman Shaw, and Herbert A Simon. 1958 · 1958
Earlier work this paper cites.
Multiple-agent planning systems
Kurt Konolige and Nils J Nilsson. 1980 · 1980
Earlier work this paper cites.
Towards a foundation for evaluating ai planners
Nabil A Kartam and David E Wilkins. 1990 · 1990
Earlier work this paper cites.
Pddl-the planning domain definition language
Drew McDermott, Malik Ghallab, Adele E. Howe, Craig A. Knoblock, Ashwin Ram, Manuela M. Veloso, Daniel S. Weld, and David E. Wilkins. 1998 · 1998
Earlier work this paper cites.
Lpg: A planner based on local search for planning graphs with action costs
Alfonso Gerevini, Ivan Serina, et al. 2002 · 2002
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe. 2004 · 2004
Earlier work this paper cites.
Val: Automatic plan validation, continuous effects and mixed initiative planning using pddl
Richard Howey, Derek Long, and Maria Fox. 2004 · 2004
Earlier work this paper cites.
The fast downward planning system
Malte Helmert. 2006 · 2006
Earlier work this paper cites.
Peter A Jansen. 2020 · 2009
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2020b · 2010
Earlier work this paper cites.
A new representation and associated algorithms for generalized planning
Siddharth Srivastava, Neil Immerman, and Shlomo Zilberstein. 2011 · 2011
Earlier work this paper cites.
Width and inference based planners: Siw, bfs (f), and probe
Nir Lipovetzky, Miquel Ramirez, Christian Muise, and Hector Geffner. 2014 · 2014
Earlier work this paper cites.
Artificial intelligence: a modern approach
Stuart J Russell and Peter Norvig. 2016 · 2016
Earlier work this paper cites.
Cooperative multi-agent planning: A survey
Alejandro Torreno, Eva Onaindia, Antonín Komenda, and Michal Štolba. 2017 · 2017
Earlier work this paper cites.
Virtualhome: Simulating household activities via programs
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba. 2018 · 2018
Earlier work this paper cites.
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
The scip optimization suite 8.0
Ksenia Bestuzheva, Mathieu Besançon, Wei-Kun Chen, Antonia Chmiela, Tim Donkiewicz, Jasper van Doornmalen, Leon Eifler, Oliver Gaul, Gerald Gamrath, Ambros Gleixner, et al. 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Earlier work this paper cites.
Skill induction and planning with latent language
Pratyusha Sharma, Antonio Torralba, and Jacob Andreas. 2021 · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. 2022 · 2022
Earlier work this paper cites.
Pre-trained language models for interactive decision-making
Shuang Li, Xavier Puig, Chris Paxton, Yilun Du, Clinton Wang, Linxi Fan, Tao Chen, De-An Huang, Ekin Akyürek, Anima Anandkumar, et al. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Earlier work this paper cites.
Planning with large language models via corrective re-prompting
Shreyas Sundara Raman, Vanya Cohen, Eric Rosen, Ifrah Idrees, David Paulius, and Stefanie Tellex. 2022 · 2022
Earlier work this paper cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Abulhair Saparov and He He. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Earlier work this paper cites.
Pg3: Policy-guided planning for generalized policy generation
Ryan Yang, Tom Silver, Aidan Curtis, Tomas Lozano-Perez, and Leslie Pack Kaelbling. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Mohamed Aghzal, Erion Plaku, and Ziyu Yao. 2023 · 2023
Earlier work this paper cites.
Codeplan: Repository-level coding using llms and planning
Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade, Vageesh D C, Arun Iyer, Suresh Parthasarathy, Sriram Rajamani, B. Ashok, and Shashank Shet. 2023 · 2023
Earlier work this paper cites.
Learning to reason over scene graphs: a case study of finetuning gpt-2 into a robot language model for grounded task planning
Georgia Chalvatzaki, Ali Younes, Daljeet Nandha, An Thai Le, Leonardo FR Ribeiro, and Iryna Gurevych. 2023 · 2023
Earlier work this paper cites.
Gautier Dagan, Frank Keller, and Alex Lascarides. 2023 · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023 · 2023
Earlier work this paper cites.
Leveraging pre-trained large language models to construct and utilize world models for model-based task planning
Lin Guan, Karthik Valmeekam, Sarath Sreedharan, and Subbarao Kambhampati. 2023 · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. 2023 · 2023
Cited alongside, same era.
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023 · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2023 · 2023
Cited alongside, same era.
To the globe (ttg): Towards language-driven guaranteed travel planning
Da Ju, Song Jiang, Andrew Cohen, Aaron Foss, Sasha Mitts, Arman Zharmagambetov, Brandon Amos, Xian Li, Justine T Kao, Maryam Fazel-Zarandi, et al. 2024 · 2024
Later among the works it cites.
Can large language models reason and plan?
Subbarao Kambhampati. 2024 · 2024
Later among the works it cites.
Llms can’t plan, but can help planning in llm-modulo frameworks
Subbarao Kambhampati, Karthik Valmeekam, Lin Guan, Kaya Stechly, Mudit Verma, Siddhant Bhambri, Lucas Saldyt, and Anil Murthy. 2024 · 2024
Later among the works it cites.
Thought of search: Planning with language models through the lens of efficiency
Michael Katz, Harsha Kokel, Kavitha Srinivas, and Shirin Sohrabi. 2024 · 2024
Later among the works it cites.
Tree search for language model agents
Jing Yu Koh, Stephen McAleer, Daniel Fried, and Ruslan Salakhutdinov. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song, Hangyu Li, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. 2023 · 2023
Cited alongside, same era.
Multimodal procedural planning via dual text-image prompting
Yujie Lu, Pan Lu, Zhiyu Chen, Wanrong Zhu, Xin Eric Wang, and William Yang Wang. 2023 · 2023
Cited alongside, same era.
Bioplanner: automatic evaluation of llms on protocol planning in biology
Odhran O’Donoghue, Aleksandar Shtedritski, John Ginger, Ralph Abboud, Ali Essa Ghareeb, Justin Booth, and Samuel G Rodriques. 2023 · 2023
Cited alongside, same era.
Data-efficient learning of natural language to linear temporal logic translators for robot task specification
Jiayi Pan, Glen Chou, and Dmitry Berenson. 2023 · 2023
Cited alongside, same era.
Adapt: As-needed decomposition and planning with language models
Archiki Prasad, Alexander Koller, Mareike Hartmann, Peter Clark, Ashish Sabharwal, Mohit Bansal, and Tushar Khot. 2023 · 2023
Cited alongside, same era.
Robots that ask for help: Uncertainty alignment for large language model planners
Allen Z Ren, Anushri Dixit, Alexandra Bodrova, Sumeet Singh, Stephen Tu, Noah Brown, Peng Xu, Leila Takayama, Fei Xia, Jake Varley, et al. 2023 · 2023
Cited alongside, same era.
Tptu: Large language model-based ai agents for task planning and tool usage
Jingqing Ruan, Yihong Chen, Bin Zhang, Zhiwei Xu, Tianpeng Bao, Guoqing Du, Shiwei Shi, Hangyu Mao, Ziyue Li, Xingyu Zeng, and Rui Zhao. 2023 · 2023
Cited alongside, same era.
Progprompt: Generating situated robot task plans using large language models
Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Beyond a*: Better planning with transformers via search dynamics bootstrapping
Lucas Lehnert, Sainbayar Sukhbaatar, DiJia Su, Qinqing Zheng, Paul Mcvay, Michael Rabbat, and Yuandong Tian. 2024 · 2024
Later among the works it cites.
Graph-enhanced large language models in asynchronous plan reasoning
Fangru Lin, Emanuele La Malfa, Valentin Hofmann, Elle Michelle Yang, Anthony Cohn, and Janet B Pierrehumbert. 2024 · 2024
Later among the works it cites.
Multimodal large language models for inverse molecular design with retrosynthetic planning
Gang Liu, Michael Sun, Wojciech Matusik, Meng Jiang, and Jie Chen. 2024 · 2024
Later among the works it cites.
Worldapis: The world is worth how many apis? a thought experiment
Jiefu Ou, Arda Uzunoglu, Benjamin Van Durme, and Daniel Khashabi. 2024 · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024 · 2024
Later among the works it cites.
Cape: Corrective actions from precondition errors using large language models
Shreyas Sundara Raman, Vanya Cohen, Ifrah Idrees, Eric Rosen, Raymond Mooney, Stefanie Tellex, and David Paulius. 2024 · 2024
Later among the works it cites.
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024 · 2024
Later among the works it cites.
Generalized planning in pddl domains with pretrained large language models
Tom Silver, Soham Dan, Kavitha Srinivas, Joshua B Tenenbaum, Leslie Kaelbling, and Michael Katz. 2024 · 2024
Later among the works it cites.
Trial and error: Exploration-based trajectory optimization for llm agents
Yifan Song, Da Yin, Xiang Yue, Jie Huang, Sujian Li, and Bill Yuchen Lin. 2024 · 2024
Later among the works it cites.
Adaplanner: Adaptive planning from feedback with language models
Haotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai, and Chao Zhang. 2024 · 2024
Later among the works it cites.
Llms still can’t plan; can lrms? a preliminary evaluation of openai’s o1 on planbench
Karthik Valmeekam, Kaya Stechly, and Subbarao Kambhampati. 2024 · 2024
Later among the works it cites.
Renxi Wang, Haonan Li, Xudong Han, Yixuan Zhang, and Timothy Baldwin. 2024 · 2024
Later among the works it cites.
Hui Wei, Shenghua He, Tian Xia, Andy Wong, Jingyang Lin, and Mei Han. 2024 · 2024
Later among the works it cites.
Can graph learning improve planning in llm-based agents?
Xixi Wu, Yifei Shen, Caihua Shan, Kaitao Song, Siwei Wang, Bohang Zhang, Jiarui Feng, Hong Cheng, Wei Chen, Yun Xiong, et al. 2024 · 2024
Later among the works it cites.
Travelplanner: A benchmark for real-world planning with language agents
Jian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, and Yu Su. 2024 · 2024
Later among the works it cites.
Selfgoal: Your language agents already know how to achieve high-level goals
Ruihan Yang, Jiangjie Chen, Yikai Zhang, Siyu Yuan, Aili Chen, Kyle Richardson, Yanghua Xiao, and Deqing Yang. 2024 · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024 · 2024
Later among the works it cites.
Justice or prejudice? quantifying biases in llm-as-a-judge
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, et al. 2024 · 2024
Later among the works it cites.
Tasklama: probing the complex task understanding of language models
Quan Yuan, Mehran Kazemi, Xin Xu, Isaac Noble, Vaiva Imbrasaite, and Deepak Ramachandran. 2024 · 2024
Later among the works it cites.
Agentohana: Design unified data and training pipeline for effective agent learning
Jianguo Zhang, Tian Lan, Rithesh Murthy, Zhiwei Liu, Weiran Yao, Juntao Tan, Thai Hoang, Liangwei Yang, Yihao Feng, Zuxin Liu, et al. 2024 · 2024
Later among the works it cites.
Large language models as commonsense knowledge for large-scale task planning
Zirui Zhao, Wee Sun Lee, and David Hsu. 2024 · 2024
Later among the works it cites.
Natural plan: Benchmarking llms on natural language planning
Huaixiu Steven Zheng, Swaroop Mishra, Hugh Zhang, Xinyun Chen, Minmin Chen, Azade Nova, Le Hou, Heng-Tze Cheng, Quoc V Le, Ed H Chi, et al. 2024 · 2024
Later among the works it cites.
Planetarium: A rigorous benchmark for translating text to structured planning languages
Max Zuo, Francisco Piedrahita Velez, Xiaochen Li, Michael L Littman, and Stephen H Bach. 2024 · 2024
Later among the works it cites.
Autonomous agents from automatic reward modeling and planning
Zhenfang Chen, Delin Chen, Wenjun Sun, Rui abd Liu, and Chuang Gan. 2025 · 2025
Closest in time.
Robotouille: An asynchronous planning benchmark for llm agents
Gonzalo Gonzalez-Pumariega, Leong Su Yean, Neha Sunkara, and Sanjiban Choudhury. 2025 · 2025
Closest in time.
Planet: A collection of benchmarks for evaluating llms’ planning capabilities
Haoming Li, Zhaoliang Chen, Jonathan Zhang, and Fei Liu. 2025 · 2025
Closest in time.
Benchmarking prompt sensitivity in large language models
Amirhossein Razavi, Mina Soltangheis, Negar Arabzadeh, Sara Salamat, Morteza Zihayat, and Ebrahim Bagheri. 2025 · 2025
Closest in time.
Monte carlo planning with large language model for text-based games
Zijing Shi, Meng Fang, and Ling Chen. 2025 · 2025
Closest in time.
Marcus Tantakoun, Xiaodan Zhu, and Christian Muise. 2025 · 2025
Closest in time.
Isr-llm: Iterative self-refined large language model for long-horizon sequential task planning
Zhehua Zhou, Jiayang Song, Kunpeng Yao, Zhan Shu, and Lei Ma. 2024 · 2088
Closest in time.