Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have shown remarkable performance in various natural language tasks, but they often struggle with planning problems that require structured reasoning.
International planning competition, 1998
IPC · 1998
Earlier work this paper cites.
The 1998 ai planning systems competition
D. McDermott · 2000
Earlier work this paper cites.
Val: Automatic plan validation, continuous effects and mixed initiative planning using pddl
R. Howey, D. Long, and M. Fox · 2004
Earlier work this paper cites.
The fast downward planning system
M. Helmert · 2006
Earlier work this paper cites.
Textworld: A learning environment for text-based games
M.-A. Côté, A. Kádár, X. Yuan, B. Kybartas, T. Barnes, E. Fine, J. Moore, R. Y. Tao, M. Hausknecht, L. E. Asri, M. Adada, W. Tay, and A. Trischler · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Deepsym: Deep symbol generation and rule learning for planning from unsupervised robot interaction
A. Ahmetoglu, M. Y. Seker, J. Piater, E. Oztop, and E. Ugur · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Earlier work this paper cites.
PDDL generators
J. Seipp, Á. Torralba, and J. Hoffmann · 2022
Earlier work this paper cites.
PDDL planning with pretrained large language models
T. Silver, V. Hariprasad, R. S. Shuttleworth, N. Kumar, T. Lozano-Pérez, and L. P. Kaelbling · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Cited alongside, same era.
G. Dagan, F. Keller, and A. Lascarides · 2023
Cited alongside, same era.
Leveraging pre-trained large language models to construct and utilize world models for model-based task planning
Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
K. Valmeekam, M. Marquez, A. Olmo, S. Sreedharan, and S. Kambhampati · 2023
Later among the works it cites.
Translating natural language to planning goals with large-language models
Y. Xie, C. Yu, T. Zhu, J. Bai, Z. Gong, and H. Soh · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. R. Narasimhan · 2023
Later among the works it cites.
Coder reviewer reranking for code generation
T. Zhang, T. Yu, T. Hashimoto, M. Lewis, W.-t. Yih, D. Fried, and S. Wang · 2023
Later among the works it cites.
Teaching large language models to self-debug
X. Chen, M. Lin, N. Schärli, and D. Zhou · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Guan, K. Valmeekam, S. Sreedharan, and S. Kambhampati · 2023
Cited alongside, same era.
Llm+p: Empowering large language models with optimal planning proficiency
B. Liu, Y. Jiang, X. Zhang, Q. Liu, S. Zhang, J. Biswas, and P. Stone · 2023
Cited alongside, same era.
Faithful chain-of-thought reasoning
Q. Lyu, S. Havaldar, A. Stein, L. Zhang, D. Rao, E. Wong, M. Apidianaki, and C. Callison-Burch · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark · 2023
Cited alongside, same era.
Lever: Learning to verify language-to-code generation with execution
A. Ni, S. Iyer, D. Radev, V. Stoyanov, W.-t. Yih, S. I. Wang, and X. V. Lin · 2023
Cited alongside, same era.
Predicate invention for bilevel planning
T. Silver, R. Chitnis, N. Kumar, W. McClinton, T. Lozano-Pérez, L. Kaelbling, and J. B. Tenenbaum · 2023
Cited alongside, same era.
Autoplanbench:: Automatically generating benchmarks for llm planners from pddl
K. Stein and A. Koller · 2023
Cited alongside, same era.
N. Dziri, X. Lu, M. Sclar, X. L. Li, L. Jiang, B. Y. Lin, S. Welleck, P. West, C. Bhagavatula, R. Le Bras, et al · 2024
Closest in time.
Interpret: Interactive predicate learning from language feedback for generalizable task planning
M. Han, Y. Zhu, S.-C. Zhu, Y. N. Wu, and Y. Zhu · 2024
Closest in time.
Large language models as planning domain generators
J. Oswald, K. Srinivas, H. Kokel, J. Lee, M. Katz, and S. Sohrabi · 2024
Closest in time.
S. Sagar, A. Taparia, and R. Senanayake · 2024
Closest in time.
Generalized planning in PDDL domains with pretrained large language models
T. Silver, S. Dan, K. Srinivas, J. Tenenbaum, L. Kaelbling, and M. Katz · 2024
Closest in time.