Fetching the paper…
Reading the bibliography…
This paper introduces SagaLLM, a structured multi-agent architecture designed to address four foundational limitations of current LLM-based planning systems: unreliable self-validation, context loss, lack of transactional safeguards, and insufficient inter-agent coordination.
On formally undecidable propositions of Principia Mathematica and related systems i
Kurt Gödel · 1967
Earlier work this paper cites.
The transaction concept: virtues and limitations
Jim Gray · 1981
Earlier work this paper cites.
Commitments and conventions: The foundation of coordination in multi-agent systems
Nicholas R Jennings · 1993
Earlier work this paper cites.
Distributed problem solving and planning , page 121–164
Edmund H. Durfee · 1999
Earlier work this paper cites.
Transactional Information Systems: Theory, Algorithms, and the Practice of Concurrency Control and Recovery
Gerhard Weikum and Gottfried Vossen · 2001
Earlier work this paper cites.
Distributed sensor networks: A multiagent perspective
Victor Lesser, Charles L Ortiz Jr, and Milind Tambe · 2004
Earlier work this paper cites.
Yawl: yet another workflow language
Wil M.P. van der Aalst and Arthur H. M. ter Hofstede · 2005
Earlier work this paper cites.
Base: An acid alternative
Dan Pritchett · 2008
Earlier work this paper cites.
An Introduction to MultiAgent Systems
Michael Wooldridge · 2009
Earlier work this paper cites.
Teamwork in multi-agent systems - a formal approach
Barbara Dunin-Keplicz and Rineke Verbrugge · 2010
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, and more · 2017
Earlier work this paper cites.
Microservices Patterns: With examples in Java
Chris Richardson · 2018
Earlier work this paper cites.
Text and patterns: For effective chain of thought, it takes two to tango, 2022
Aman Madaan and Amir Yazdanbakhsh · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate, 2023
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch · 2023
Earlier work this paper cites.
Large language models as commonsense knowledge for large-scale task planning
Zirui Zhao, Wee Sun Lee, and David Hsu · 2023
Cited alongside, same era.
Claude Technical Report, 2024
Anthropic · 2024
Cited alongside, same era.
PLASMA: Making small language models better procedural knowledge models for (counterfactual) planning
Faeze Brahman, Chandra Bhagavatula, Valentina Pyatkin, and Yejin Choi · 2024
Cited alongside, same era.
EVINCE: Optimizing Adversarial LLM Dialogues via Conditional Statistics and Information Theory
Edward Y Chang · 2024
Cited alongside, same era.
Multi-LLM Agent Collaborative Intelligence: The Path to Artificial General Intelligence
Edward Y. Chang · 2024
Cited alongside, same era.
A survey on evaluation of large language models
Failure modes of llms for causal reasoning on narratives, 2024
Khurram Yamin, Shantanu Gupta, Gaurav R. Ghosal, Zachary C. Lipton, and Bryan Wilder · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shinnosuke Yao, Dong Yu, Jianfeng Zhao, Izhak Shafran, Thomas Griffiths, Yanjun Cao, and Karthik Narasimhan · 2024
Later among the works it cites.
Aflow: Automating agentic workflow generation, 2024
Jiayi Zhang, Jinyu Xiang, Zhaoyang Yu, Fengwei Teng, Xionghui Chen, Jiaqi Chen, Mingchen Zhuge, Xin Cheng, Sirui Hong, Jinlin Wang, Bingnan Zheng, Bang Liu, Yuyu Luo, and Chenglin Wu · 2024
Later among the works it cites.
https://aws.amazon.com/step-functions/ , 2023
AWS Step Functions · 2025
Closest in time.
https://azure.microsoft.com/en-us/products/logic-apps/ , 2023
Azure Logic Apps · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, and more · 2024
Cited alongside, same era.
Agentscope: A flexible yet robust multi-agent platform, 2024
Dawei Gao, Zitao Li, Xuchen Pan, Weirui Kuang, Zhijian Ma, Bingchen Qian, Fei Wei, Wenhao Zhang, Yuexiang Xie, Daoyuan Chen, and more · 2024
Cited alongside, same era.
Found in the middle: Calibrating positional attention bias improves long context utilization, 2024
Cheng-Yu Hsieh, Yung-Sung Chuang, Chun-Liang Li, Zifeng Wang, Long T. Le, and more · 2024
Cited alongside, same era.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou · 2024
Cited alongside, same era.
Self-[in]correct: Llms struggle with refining self-generated responses
Dongwei Jiang, Jingyu Zhang, Orion Weller, Nathaniel Weir, Benjamin Van Durme, and Daniel Khashabi · 2024
Cited alongside, same era.
Langgraph: Building structured applications with llms
LangChain AI · 2024
Cited alongside, same era.
Lost in the middle: How language models use long contexts
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang · 2024
Cited alongside, same era.
Why do multi-agent llm systems fail?, 2025
Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, and Ion Stoica · 2025
Closest in time.
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning, 2025
Edward Y. Chang and Longling Geng · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, Daya Guo, Dejian Yang, and more · 2025
Closest in time.
Source code for sagallm paper experiments
Longling Geng · 2025
Closest in time.
Realm-bench: A real-world planning benchmark for llms and multi-agent systems, 2025
Longling Geng and Edward Y. Chang · 2025
Closest in time.
Nolima: Long-context evaluation beyond literal matching, 2025
Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt, Trung Bui, Ryan A. Rossi, Seunghyun Yoon, and Hinrich Schütze · 2025
Closest in time.
Hello GPT-4o, 2024
OpenAI · 2025
Closest in time.
Plangenllms: A modern survey of llm planning capabilities, 2025
Hui Wei, Zihao Zhang, Shenghua He, Tian Xia, Shijia Pan, and Fei Liu · 2025
Closest in time.
A survey of large language models (updated 2025), 2025
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, and more · 2025
Closest in time.