Fetching the paper…
Reading the bibliography…
Recent agent frameworks and inference-time algorithms often struggle with complex planning problems due to limitations in verifying generated plans or reasoning and varying complexity of instances within a single task.
Leveraging pre-trained large language models to construct and utilize world models for model-based task planning
L. Guan, K. Valmeekam, S. Sreedharan, and S. Kambhampati · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model
S. Hao, Y. Gu, H. Ma, J. Hong, Z. Wang, D. Wang, and Z. Hu · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao · 2023
Earlier work this paper cites.
Explicit planning helps language models in logical reasoning
H. Zhao, K. Wang, M. Yu, and H. Mei · 2023
Earlier work this paper cites.
Exploring and benchmarking the planning capabilities of large language models
B. Bohnet, A. Nova, A. T. Parisi, K. Swersky, K. Goshvadi, H. Dai, D. Schuurmans, N. Fiedel, and H. Sedghi · 2024
Earlier work this paper cites.
Large language monkeys: Scaling inference compute with repeated sampling
B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. Ré, and A. Mirhoseini · 2024
Earlier work this paper cites.
RePrompt: Planning by Automatic Prompt Engineering for Large Language Models Agents
W. Chen, S. Koenig, and B. Dilkina · 2024
Earlier work this paper cites.
UCB algorithms for multi-armed bandits: Precise regret and adaptive inference
Q. Han, K. Khamaru, and C.-H. Zhang · 2024
Earlier work this paper cites.
OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems
C. He, R. Luo, Y. Bai, S. Hu, Z. Thai, J. Shen, J. Hu, X. Han, Y. Huang, Y. Zhang, J. Liu, L. Qi, Z. Liu, and M. Sun · 2024
Earlier work this paper cites.
Large language models cannot self-correct reasoning yet
J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song, and D. Zhou · 2024
Earlier work this paper cites.
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al · 2024
Cited alongside, same era.
Learning planning-based reasoning by trajectories collection and process reward synthesizing
F. Jiao, C. Qin, Z. Liu, N. F. Chen, and S. Joty · 2024
Cited alongside, same era.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2024
Cited alongside, same era.
Tool-Planner: Dynamic Solution Tree Planning for Large Language Model with Tool Clustering
Y. Liu, X. Peng, Y. Zhang, J. Cao, X. Zhang, S. Cheng, X. Wang, J. Yin, and T. Du · 2024
Cited alongside, same era.
DocFinQA: A long-context financial reasoning dataset
V. Reddy, R. Koncel-Kedziorski, V. D. Lai, M. Krumdick, C. Lovering, and C. Tanner · 2024
Cited alongside, same era.
Chain-of-experts: When LLMs meet complex operations research problems
Z. Xiao, D. Zhang, Y. Wu, L. Xu, Y. J. Wang, X. Han, X. Fu, T. Zhong, J. Zeng, M. Song, and G. Chen · 2024
Later among the works it cites.
A human-like reasoning framework for multi-phases planning task with large language models
C. Xie and D. Zou · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan · 2024
Later among the works it cites.
D. Zhang, X. Huang, D. Zhou, Y. Li, and W. Ouyang · 2024
Later among the works it cites.
Natural plan: Benchmarking llms on natural language planning
H. S. Zheng, S. Mishra, H. Zhang, X. Chen, M. Chen, A. Nova, L. Hou, H.-T. Cheng, Q. V. Le, E. H. Chi, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GPQA: A graduate-level google-proof q&a benchmark
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman · 2024
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
G. Team, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang, et al · 2024
Cited alongside, same era.
Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
K. Valmeekam, M. Marquez, A. Olmo, S. Sreedharan, and S. Kambhampati · 2024
Cited alongside, same era.
From decoding to meta-generation: Inference-time algorithms for large language models
S. Welleck, A. Bertsch, M. Finlayson, H. Schoelkopf, A. Xie, G. Neubig, I. Kulikov, and Z. Harchaoui · 2024
Cited alongside, same era.
Os-copilot: Towards generalist computer agents with self-improvement
Z. Wu, C. Han, Z. Ding, Z. Weng, Z. Liu, S. Yao, T. Yu, and L. Kong · 2024
Cited alongside, same era.
A survey on large language model based autonomous agents
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, et al
Cited in the paper.
Promptagent: Strategic planning with language models enables expert-level prompt optimization
X. Wang, C. Li, Z. Wang, F. Bai, H. Luo, J. Zhang, N. Jojic, E. Xing, and Z. Hu
Cited in the paper.
Later among the works it cites.
Knowagent: Knowledge-augmented planning for llm-based agents
Y. Zhu, S. Qiao, Y. Ou, S. Deng, N. Zhang, S. Lyu, Y. Shen, L. Liang, J. Gu, and H. Chen · 2024
Later among the works it cites.
K.-H. Lee, I. Fischer, Y.-H. Wu, D. Marwood, S. Baluja, D. Schuurmans, and X. Chen · 2025
Closest in time.
Scaling test-time compute optimally can be more effective than scaling LLM parameters
C. V. Snell, J. Lee, K. Xu, and A. Kumar · 2025
Closest in time.
Planning in natural language improves LLM search for code generation
E. Z. Wang, F. Cassano, C. Wu, Y. Bai, W. Song, V. Nath, Z. Han, S. M. Hendryx, S. Yue, and H. Zhang · 2025
Closest in time.