Fetching the paper…
Reading the bibliography…
This benchmark suite provides a comprehensive evaluation framework for assessing both individual LLMs and multi-agent systems in Real-world planning and scheduling scenarios.
Benchmarks for shop scheduling problems
Ebru Demirkol, Sanjay Mehta, and Reha Uzsoy. 1998 · 1998
Earlier work this paper cites.
The Process Planning Competition 2020: An Overview. In Proceedings of the International Conference on Automated Planning and Scheduling (ICAPS) . ICAPS, Virtual
A. Smith and B. Johnson. 2020 · 2020
Earlier work this paper cites.
The Dynamic Planning Competition 2022: Challenges and Results. In Proceedings of the International Conference on Automated Planning and Scheduling (ICAPS) . ICAPS, Virtual
J. Doe and R. Roe. 2022 · 2022
Earlier work this paper cites.
Timebench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Haotian Wang, Ming Liu, and Bing Qin. 2023 · 2023
Earlier work this paper cites.
Flexible job shop scheduling problem under Industry 5.0: A survey on human reintegration, environmental consideration and resilience improvement
Candice Destouet, Houda Tlahig, Belgacem Bettayeb, and Bélahcène Mazari. 2023 · 2023
Earlier work this paper cites.
DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, and more. 2023 · 2023
Earlier work this paper cites.
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023 · 2023
Earlier work this paper cites.
Automated Negotiation Agents Competition 2023: Benchmarking Negotiation Strategies. In Proceedings of the International Conference on Autonomous Agents and Multiagent Systems (AAMAS) . IFAAMAS, Virtual
E. Miller and F. Davis. 2023 · 2023
Cited alongside, same era.
Taskbench: Benchmarking large language models for task automation
Yongliang Shen et al · 2023
Cited alongside, same era.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, et al · 2023
Cited alongside, same era.
Karthik Valmeekam, Matthew Marquez, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati. 2023 · 2023
Cited alongside, same era.
CrewAI Framework
Joao Moura. 2024 · 2024
Later among the works it cites.
The 2023 International Planning Competition
Ayal Taitler, Ron Alford, Joan Espasa, Gregor Behnke, Daniel Fišer, Michael Gimelfarb, Florian Pommerening, Scott Sanner, Enrico Scala, Dominik Schreiber, Javier Segovia-Aguas, and Jendrik Seipp. 2024 · 2024
Later among the works it cites.
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. In COLM 2024
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, and Chi Wang. 2024 · 2024
Later among the works it cites.
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
Edward Y. Chang and Longling Geng. 2025b · 2025
Closest in time.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI, Daya Guo, Dejian Yang, and more. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
XAgent: An Autonomous Agent for Complex Task Solving
XAgent Team. 2023 · 2023
Cited alongside, same era.
Claude Technical Report
Anthropic. 2024 · 2024
Cited alongside, same era.
LangGraph: Building Structured Applications with LLMs
LangChain AI. 2024 · 2024
Cited alongside, same era.
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning. In arXiv:2505.12501
Edward Y. Chang and Longling Geng. 2025a
Cited in the paper.
Hello GPT-4o
OpenAI. 2024 · 2025
Closest in time.