Fetching the paper…
Reading the bibliography…
While Large Language Models (LLMs) have shown remarkable advancements in reasoning and tool use, they often fail to generate optimal, grounded solutions under complex constraints.
Strips: A new approach to the application of theorem proving to problem solving
R. E. Fikes and N. J. Nilsson · 1971
Earlier work this paper cites.
Consistency in networks of relations
A. K. Mackworth · 1977
Earlier work this paper cites.
Belief maintenance in dynamic constraint networks
R. Dechter and A. Dechter · 1988
Earlier work this paper cites.
1990-Dynamic Constraint Satisfaction Problems
S. Mittal and B. Falkenhainer · 1990
Earlier work this paper cites.
Trains-95: Towards a mixed-initiative planning assistant
G. Ferguson, J. Allen, and B. Miller · 1996
Earlier work this paper cites.
Suggestion strategies for constraint-based matchmaker agents
E. C. Freuder and R. J. Wallace · 1998
Earlier work this paper cites.
Mixed-initiative interaction
M. A. Hearst · 1999
Earlier work this paper cites.
Bidirectional reasoning in decision making by constraint satisfaction
K. J. Holyoak and D. Simon · 1999
Earlier work this paper cites.
A knowledge-based approach to planning with incomplete information and sensing
R. P. A. Petrick and F. Bacchus · 2002
Earlier work this paper cites.
Constraint Processing
R. Dechter · 2003
Earlier work this paper cites.
Pddl2.1: An extension to PDDL for expressing temporal planning domains
M. Fox and D. Long · 2003
Earlier work this paper cites.
VAL: Automatic plan validation, continuous effects and mixed initiative planning using PDDL
R. Howey, D. Long, and M. Fox · 2004
Earlier work this paper cites.
Extending the knowledge-based approach to planning with incomplete information and sensing
R. P. A. Petrick and F. Bacchus · 2004
Earlier work this paper cites.
Acquiring both constraint and solution preferences in interactive constraint systems
F. Rossi and A. Sperduti · 2004
Earlier work this paper cites.
Query-driven constraint acquisition
C. Bessiere, R. Coletta, B. O’Sullivan, and M. Paulin · 2007
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Earlier work this paper cites.
CRoW: Benchmarking commonsense reasoning in real-world tasks
M. Ismayilzada, D. Paul, S. Montariol, M. Geva, and A. Bosselut · 2023
Cited alongside, same era.
Commonsense reasoning and explainable artificial intelligence using large language models
S. Krause and F. Stolzenburg · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning
N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao · 2023
Cited alongside, same era.
On the planning abilities of large language models: A critical investigation
K. Valmeekam, M. Marquez, A. Olmo, S. Sreedharan, and S. Kambhampati · 2023
Cited alongside, same era.
ReAct: Synergizing Reasoning and Acting in Language Models
System Card: Claude Opus 4 & Claude Sonnet 4
Anthropic · 2025
Closest in time.
SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling, Feb. 2025
J. Chen, J. Ren, X. Chen, C. Yang, R. Sun, and S. Ö. Arık · 2025
Closest in time.
G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, et al · 2025
Closest in time.
Mind2web 2: Evaluating agentic search with agent-as-a-judge
B. Gou, Z. Huang, Y. Ning, Y. Gu, M. Lin, W. Qi, A. Kopanev, B. Yu, B. J. Gutiérrez, Y. Shu, et al · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2023
Cited alongside, same era.
Large language models as commonsense knowledge for large-scale task planning
Z. Zhao, W. S. Lee, and D. Hsu · 2023
Cited alongside, same era.
TravelAgent: An AI Assistant for Personalized Travel Planning, Sept. 2024
A. Chen, X. Ge, Z. Fu, Y. Xiao, and J. Chen · 2024
Cited alongside, same era.
Malade: Orchestration of llm-powered agents with retrieval augmented generation for pharmacovigilance
J. Choi, N. Palumbo, P. Chalasani, M. M. Engelhard, S. Jha, A. Kumar, and D. Page · 2024
Cited alongside, same era.
CRITIC: Large language models can self-correct with tool-interactive critiquing
Z. Gou, Z. Shao, Y. Gong, yelong shen, Y. Yang, N. Duan, and W. Chen · 2024
Cited alongside, same era.
Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning, May 2024
A. Gundawar, M. Verma, L. Guan, K. Valmeekam, S. Bhambri, and S. Kambhampati · 2024
Cited alongside, same era.
Position: LLMs can’t plan, but can help planning in LLM-modulo frameworks
S. Kambhampati, K. Valmeekam, L. Guan, M. Verma, K. Stechly, S. Bhambri, L. P. Saldyt, and A. B. Murthy · 2024
Cited alongside, same era.
B. Jiang, Z. Hao, Y.-M. Cho, B. Li, Y. Yuan, S. Chen, L. Ungar, C. J. Taylor, and D. Roth · 2025
Closest in time.
Evolving Deeper LLM Thinking, Jan. 2025
K.-H. Lee, I. Fischer, Y.-H. Wu, D. Marwood, S. Baluja, D. Schuurmans, and X. Chen · 2025
Closest in time.
Hello again! llm-powered personalized agent for long-term dialogue
H. Li, C. Yang, A. Zhang, Y. Deng, X. Wang, and T.-S. Chua · 2025
Closest in time.
Z. Lu, W. Lu, Y. Tao, Y. Dai, Z. Chen, H. Zhuang, C. Chen, H. Peng, and Z. Zeng · 2025
Closest in time.
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents, June 2025
J. Oh, E. Kim, and A. Oh · 2025
Closest in time.
PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving, Feb. 2025
M. Parmar, X. Liu, P. Goyal, Y. Chen, L. Le, S. Mishra, H. Mobahi, J. Gu, Z. Wang, H. Nakhost, C. Baral, C.-Y. Lee, T. Pfister, and H. Palangi · 2025
Closest in time.
Agentic reasoning: A streamlined framework for enhancing LLM reasoning with agentic tools
J. Wu, J. Zhu, Y. Liu, M. Xu, and Y. Jin · 2025
Closest in time.
GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning, May 2025
Z. Xiang, L. Zheng, Y. Li, J. Hong, Q. Li, H. Xie, J. Zhang, Z. Xiong, C. Xie, C. Yang, D. Song, and B. Li · 2025
Closest in time.
Revealing the Barriers of Language Agents in Planning
J. Xie, K. Zhang, J. Chen, S. Yuan, K. Zhang, Y. Zhang, L. Li, and Y. Xiao · 2025
Closest in time.
EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms, Mar. 2025
S. Yuan, K. Song, J. Chen, X. Tan, D. Li, and D. Yang · 2025
Closest in time.
Planning with Multi-Constraints via Collaborative Language Agents
C. Zhang, X. D. Goh, D. Li, H. Zhang, and Y. Liu · 2025
Closest in time.