Fetching the paper…
Reading the bibliography…
The planning ability of Large Language Models (LLMs) has garnered increasing attention in recent years due to their remarkable capacity for multi-step reasoning and their ability to generalize across a wide range of domains.
TextWorld: A Learning Environment for Text-based Games
Marc-Alexandre Côté, Ákos Kádár, , et al · 2018
Earlier work this paper cites.
The Book of Why: The New Science of Cause and Effect
Judea Pearl and Dana Mackenzie · 2018
Earlier work this paper cites.
VirtualHome: Simulating Household Activities via Programs
Xavier Puig, Kevin Ra, et al · 2018
Earlier work this paper cites.
BabyAI: First steps towards grounded language learning with a human in the loop
Maxime Chevalier-Boisvert, Dzmitry Bahdanau, et al · 2019
Earlier work this paper cites.
Interactive fiction games: A colossal adventure
Matthew Hausknecht, Prithviraj Ammanabrolu, et al · 2020
Earlier work this paper cites.
A Benchmark for Systematic Generalization in Grounded Language Understanding
Laura Ruis, Jacob Andreas, et al · 2020
Earlier work this paper cites.
ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, et al · 2021
Earlier work this paper cites.
A comprehensive review on resolving ambiguities in natural language processing
Apurwa Yadav, Aarshil Patel, et al · 2021
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, et al · 2022
Earlier work this paper cites.
A path towards autonomous machine intelligence
Yann LeCun · 2022
Earlier work this paper cites.
ProgPrompt: Generating Situated Robot Task Plans using Large Language Models
Ishika Singh, Valts Blukis, et al · 2022
Earlier work this paper cites.
Emergent Abilities of Large Language Models, 2022
Jason Wei, Yi Tay, et al · 2022
Earlier work this paper cites.
Chain of Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, et al · 2022
Earlier work this paper cites.
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Shunyu Yao, Howard Chen, et al · 2022
Earlier work this paper cites.
Vision-Language Models as a Source of Rewards
Harris Chan, Volodymyr Mnih, et al · 2023
Earlier work this paper cites.
NL2TL: Transforming natural languages to temporal logics using large language models
Yongchao Chen, Rujul Gandhi, Yang Zhang, and Chuchu Fan · 2023
Earlier work this paper cites.
Task and motion planning with large language models for object rearrangement
Yan Ding, Xiaohan Zhang, et al · 2023
Earlier work this paper cites.
Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task Planning
Lin Guan, Karthik Valmeekam, et al · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, et al · 2023
Earlier work this paper cites.
Reward Design with Language Models
Minae Kwon, Sang Michael Xie, et al · 2023
Earlier work this paper cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, et al · 2023
Earlier work this paper cites.
API-bank: A comprehensive benchmark for tool-augmented LLMs
Minghao Li, Yingxiu Zhao, et al · 2023
Earlier work this paper cites.
LLM+P: Empowering Large Language Models with Optimal Planning Proficiency, 2023
Bo Liu, Yuqian Jiang, et al · 2023
Earlier work this paper cites.
Self-Refine: Iterative Refinement with Self-Feedback
Aman Madaan, Niket Tandon, et al · 2023
Earlier work this paper cites.
Reflexion: language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, et al · 2023
Earlier work this paper cites.
AdaPlanner: Adaptive Planning from Feedback with Language Models
Haotian Sun, Yuchen Zhuang, et al · 2023
Earlier work this paper cites.
PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change
Karthik Valmeekam, Matthew Marquez, et al · 2023
Earlier work this paper cites.
Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models
Lei Wang, Wanyu Xu, et al · 2023
Earlier work this paper cites.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Xuezhi Wang, Jason Wei, et al · 2023
Earlier work this paper cites.
Translating Natural Language to Planning Goals with Large-Language Models, 2023
Yaqi Xie, Chen Yu, et al · 2023
Earlier work this paper cites.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Shunyu Yao, Dian Yu, et al · 2023
Earlier work this paper cites.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Shunyu Yao, Dian Yu, et al · 2023
Earlier work this paper cites.
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao, Jeffrey Zhao, et al · 2023
Earlier work this paper cites.
Language to Rewards for Robotic Skill Synthesis
Wenhao Yu, Nimrod Gileadi, et al · 2023
Cited alongside, same era.
Grounding Classical Task Planners via Vision-Language Models, 2023
Xiaohan Zhang, Yan Ding, et al · 2023
Cited alongside, same era.
Large Language Models as Commonsense Knowledge for Large-Scale Task Planning
Zirui Zhao, Wee Sun Lee, et al · 2023
Cited alongside, same era.
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Denny Zhou, Nathanael Schärli, et al · 2023
Cited alongside, same era.
Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning
Mohamed Aghzal, Erion Plaku, and Ziyu Yao · 2024
Cited alongside, same era.
Look Further Ahead: Testing the Limits of GPT-4 in Path Planning
Mohamed Aghzal, Erion Plaku, and Ziyu Yao · 2024
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models, 2024
Iman Mirzadeh, Keivan Alizadeh, et al · 2024
Later among the works it cites.
AmbigNLG: Addressing task ambiguity in instruction for NLG
Ayana Niwa and Hayate Iso · 2024
Later among the works it cites.
Language-guided Manipulator Motion Planning with Bounded Task Space
Thies Oelerich, Christian Hartl-Nesic, et al · 2024
Later among the works it cites.
Large Language Models as Planning Domain Generators
James Oswald, Kavitha Srinivas, et al · 2024
Later among the works it cites.
Robust Agents Learn Causal World Models
Jonathan Richens and Tom Everitt · 2024
Later among the works it cites.
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Juan Rocamonde, Victoriano Montesinos, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Evaluating vision-language models as evaluators in path planning
Mohamed Aghzal, Xiang Yue, Erion Plaku, and Ziyu Yao · 2024
Cited alongside, same era.
Large language models for mathematical reasoning: Progresses and challenges
Janice Ahn, Rishu Verma, et al · 2024
Cited alongside, same era.
Graph of Thoughts: Solving Elaborate Problems with Large Language Models
Maciej Besta, Nils Blach, et al · 2024
Cited alongside, same era.
Exploring and benchmarking the planning capabilities of large language models, 2024
Bernd Bohnet, Azade Nova, et al · 2024
Cited alongside, same era.
Can we rely on LLM agents to draft long-horizon plans? Let’s take TravelPlanner as an example
Yanan Chen, Ali Pesaranghader, et al · 2024
Cited alongside, same era.
AutoTAMP: Autoregressive task and motion planning with LLMs as translators and checkers
Yongchao Chen, Jacob Arkin, et al · 2024
Cited alongside, same era.
Later among the works it cites.
TaskBench: Benchmarking Large Language Models for Task Automation, 2024
Yongliang Shen, Kaitao Song, et al · 2024
Later among the works it cites.
Generating consistent pddl domains with large language models, 2024
Pavel Smirnov, Frank Joublin, et al · 2024
Later among the works it cites.
VLM-Social-Nav: Socially Aware Robot Navigation through Scoring using Vision-Language Models
Daeun Song, Jing Liang, et al · 2024
Later among the works it cites.
Chain of Thoughtlessness? An Analysis of CoT in Planning
Kaya Stechly, Karthik Valmeekam, et al · 2024
Later among the works it cites.
On the self-verification limitations of large language models on reasoning and planning tasks, 2024
Kaya Stechly, Karthik Valmeekam, et al · 2024
Later among the works it cites.
A comprehensive survey of hallucination mitigation techniques in large language models, 2024
S. M Towhidul Islam Tonmoy, S M Mehedi Zaman, et al · 2024
Later among the works it cites.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, et al · 2024
Later among the works it cites.
Evaluating the World Model Implicit in a Generative Model
Keyon Vafa, Justin Y Chen, et al · 2024
Later among the works it cites.
LLMs Still Can’t Plan; Can LRMs? A Preliminary Evaluation of OpenAI’s o1 on PlanBench
Karthik Valmeekam, Kaya Stechly, et al · 2024
Later among the works it cites.
On the Brittle Foundations of ReAct Prompting for Agentic Large Language Models, 2024
Mudit Verma, Siddhant Bhambri, et al · 2024
Later among the works it cites.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, et al · 2024
Later among the works it cites.
LLM3: Large language model-based task and motion planning with motion failure reasoning
Shu Wang, Muzhi Han, et al · 2024
Later among the works it cites.
MINT: Evaluating LLMs in multi-turn interaction with tools and language feedback
Xingyao Wang, Zihan Wang, et al · 2024
Later among the works it cites.
TravelPlanner: A Benchmark for Real-World Planning with Language Agents
Jian Xie, Kai Zhang, et al · 2024
Later among the works it cites.
OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Tianbao Xie, Danyang Zhang, et al · 2024
Later among the works it cites.
Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
Tianbao Xie, Siheng Zhao, et al · 2024
Later among the works it cites.
Guiding Long-Horizon Task and Motion Planning with Vision Language Models
Zhutian Yang, Caelan Garrett, et al · 2024
Later among the works it cites.
From task structures to world models: what do LLMs know?
Ilker Yildirim and L.A. Paul · 2024
Later among the works it cites.
PDDLEGO: Iterative planning in textual environments
Li Zhang, Peter Jansen, Tianyi Zhang, Peter Clark, Chris Callison-Burch, and Niket Tandon · 2024
Later among the works it cites.
Policy Improvement using Language Feedback Models
Victor Zhong, Dipendra Misra, et al · 2024
Later among the works it cites.
WebArena: A Realistic Web Environment for Building Autonomous Agents
Shuyan Zhou, Frank F. Xu, et al · 2024
Later among the works it cites.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, 2025
DeepSeek-AI · 2025
Closest in time.
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
Gonzalo Gonzalez-Pumariega, Leong Su Yean, et al · 2025
Closest in time.
Planning anything with rigor: General-purpose zero-shot planning with LLM-based formalized programming
Yilun Hao, Yang Zhang, et al · 2025
Closest in time.
System 1.x: Learning to balance fast and slow planning with language models
Swarnadeep Saha, Archiki Prasad, et al · 2025
Closest in time.
DOTS: Learning to reason dynamically in LLMs via optimal reasoning trajectories search
Murong Yue, Wenlin Yao, Haitao Mi, Dian Yu, Ziyu Yao, and Dong Yu · 2025
Closest in time.