Fetching the paper…
Reading the bibliography…
Despite enthusiasm for Multi-Agent LLM Systems (MAS), their performance gains on popular benchmarks are often minimal.
The Discovery of Grounded Theory: Strategies for Qualitative Research
Barney G. Glaser and Anselm L. Strauss · 1967
Earlier work this paper cites.
Normal Accidents: Living with High-Risk Technologies
Charles Perrow · 1984
Earlier work this paper cites.
New challenges in organizational research: High reliability organizations
Karlene H. Roberts · 1989
Earlier work this paper cites.
Reliable organizations: Present research and future directions
Gene I Rochlin · 1996
Earlier work this paper cites.
Uncertainty, action, and interaction: In pursuit of mixed-initiative computing
Eric Horvitz · 1999
Earlier work this paper cites.
Theoretical sampling and category development in grounded theory
Claire B Draucker, Donna S Martsolf, Ratchneewan Ross, and Thomas B Rusk · 2007
Earlier work this paper cites.
Open coding
Shahedul Huq Khandkar · 2009
Earlier work this paper cites.
Learning attentional communication for multi-agent cooperation
Jiechuan Jiang and Zongqing Lu · 2018
Earlier work this paper cites.
Learning when to communicate at scale in multiagent cooperative and competitive tasks
Amanpreet Singh, Tushar Jain, and Sainbayar Sukhbaatar · 2018
Earlier work this paper cites.
Question answering over knowledge bases by leveraging semantic parsing and neuro-symbolic reasoning
Pavan Kapanipathi, Ibrahim Abdelaziz, Srinivas Ravishankar, Salim Roukos, Alexander Gray, Ramon Astudillo, Maria Chang, Cristina Cornelio, Saswati Dana, Achille Fokoue, et al · 2020
Earlier work this paper cites.
Multi-agent graph-attention communication and teaming
Yaru Niu, Rohan R Paleja, and Matthew C Gombolay · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
The surprising effectiveness of ppo in cooperative multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu · 2022
Earlier work this paper cites.
Gorilla: Large language model connected with massive apis, 2023
Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez · 2023
Earlier work this paper cites.
Chatdev: Communicative agents for software development
Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun · 2023
Earlier work this paper cites.
Roco: Dialectic multi-robot collaboration with large language models, 2023
Zhao Mandi, Shreeya Jain, and Shuran Song · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate, 2023
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Earlier work this paper cites.
Dspy: Compiling declarative language model calls into self-improving pipelines, 2023
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts · 2023
Earlier work this paper cites.
Trustworthy llms: A survey and guideline for evaluating large language models’ alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li · 2023
Earlier work this paper cites.
An empirical study of code generation errors made by large language models
Song Da, Zijie Zhou, Zhijie Wang, Yuheng Huang, Shengmai Chen, Bonan Kou, Lei Ma, and Tianyi Zhang · 2023
Earlier work this paper cites.
Metagpt: Meta programming for multi-agent collaborative framework
Sirui Hong, Xiawu Zheng, Jonathan Chen, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, et al · 2023
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang · 2023
Cited alongside, same era.
Multi-agent collaboration: Harnessing the power of intelligent llm agents
Yashar Talebirad and Amirhossein Nadiri · 2023
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu · 2023
Cited alongside, same era.
Specifications: The missing link to making the development of llm systems an engineering discipline
Ion Stoica, Matei Zaharia, Joseph Gonzalez, Ken Goldberg, Hao Zhang, Anastasios Angelopoulos, Shishir G Patil, Lingjiao Chen, Wei-Lin Chiang, and Jared Q Davis · 2024
Later among the works it cites.
Challenges in human-agent communication
Gagan Bansal, Jennifer Wortman Vaughan, Saleema Amershi, Eric Horvitz, Adam Fourney, Hussein Mozannar, Victor Dibia, and Daniel S. Weld · 2024
Later among the works it cites.
Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues
Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, and Wanli Ouyang · 2024
Later among the works it cites.
Assessing and verifying task utility in llm-powered applications, 2024
Negar Arabzadeh, Siqing Huo, Nikhil Mehta, Qinqyun Wu, Chi Wang, Ahmed Awadallah, Charles L. A. Clarke, and Julia Kiseleva · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Large language models are better reasoners with self-verification
Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun, Kang Liu, and Jun Zhao · 2023
Cited alongside, same era.
Towards reasoning in large language models via multi-agent peer review collaboration
Zhenran Xu, Senbao Shi, Baotian Hu, Jindi Yu, Dongfang Li, Min Zhang, and Yuxiang Wu · 2023
Cited alongside, same era.
Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, et al · 2023
Cited alongside, same era.
Servicenow: From startup to world’s most innovative company
Barnali Chakraborty and Debapratim Purkayastha · 2023
Cited alongside, same era.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et al · 2023
Cited alongside, same era.
Memgpt: Towards llms as operating systems, 2024
Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez · 2024
Cited alongside, same era.
The virtual lab: Ai agents design new sars-cov-2 nanobodies with experimental validation
Kyle Swanson, Wesley Wu, Nash L. Bulaong, John E. Pak, and James Zou · 2024
Cited alongside, same era.
Magentic-one: A generalist multi-agent system for solving complex tasks
Adam Fourney, Gagan Bansal, Hussein Mozannar, Cheng Tan, Eduardo Salinas, Friederike Niedtner, Grace Proebsting, Griffin Bassman, Jack Gerrits, Jacob Alber, et al · 2024
Cited alongside, same era.
Huy Nhat Phan, Tien N Nguyen, Phong X Nguyen, and Nghi DQ Bui · 2024
Later among the works it cites.
Appworld: A controllable world of apps and people for benchmarking interactive coding agents
Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, and Niranjan Balasubramanian · 2024
Later among the works it cites.
Chatdev: Communicative agents for software development
Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, et al · 2024
Later among the works it cites.
Langgraph, 2024
LangChain · 2024
Later among the works it cites.
Improving llm reasoning with multi-agent tree-of-thought validator agent
Fatemeh Haji, Mazal Bethany, Maryam Tabar, Jason Chiang, Anthony Rios, and Peyman Najafirad · 2024
Later among the works it cites.
Inference scaling f laws: The limits of llm resampling with imperfect verifiers
Benedikt Stroebl, Sayash Kapoor, and Arvind Narayanan · 2024
Later among the works it cites.
Testgeneval: A real world unit test generation and test completion benchmark
Kush Jain, Gabriel Synnaeve, and Baptiste Rozière · 2024
Later among the works it cites.
Qwen2. 5-coder technical report
Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Kai Dang, et al · 2024
Later among the works it cites.
Towards an ai co-scientist, 2025
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, Khaled Saab, Dan Popovici, Jacob Blum, Fan Zhang, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Pushmeet Kohli, Yossi Matias, Andrew Carroll, Kavita Kulkarni, Nenad Tomasev, Yuan Guan, Vikram Dhillon, Eeshit Dhaval Vaishnav, Byron Lee, Tiago R D Costa, José R Penadés, Gary Peltz, Yunhan Xu, Annalisa Pawlosky, Alan Karthikesalingam, and Vivek Natarajan · 2025
Closest in time.
Openmanus: An open-source framework for building general ai agents
Xinbin Liang, Jinyu Xiang, Zhaoyang Yu, Jiayi Zhang, and Sirui Hong · 2025
Closest in time.
Multi-agent risks from advanced ai, 2025
Lewis Hammond, Alan Chan, Jesse Clifton, Jason Hoelscher-Obermaier, Akbir Khan, Euan McLean, Chandler Smith, Wolfram Barfuss, Jakob Foerster, Tomáš Gavenčiak, The Anh Han, Edward Hughes, Vojtěch Kovařík, Jan Kulveit, Joel Z. Leibo, Caspar Oesterheld, Christian Schroeder de Witt, Nisarg Shah, Michael Wellman, Paolo Bova, Theodor Cimpeanu, Carson Ezell, Quentin Feuillade-Montixi, Matija Franklin, Esben Kran, Igor Krawczuk, Max Lamparth, Niklas Lauffer, Alexander Meinke, Sumeet Motwani, Anka Reuel, Vincent Conitzer, Michael Dennis, Iason Gabriel, Adam Gleave, Gillian Hadfield, Nika Haghtalab, Atoosa Kasirzadeh, Sébastien Krier, Kate Larson, Joel Lehman, David C. Parkes, Georgios Piliouras, and Iyad Rahwan · 2025
Closest in time.
Interactive debugging and steering of multi-agent ai systems
Will Epperson, Gagan Bansal, Victor Dibia, Adam Fourney, Jack Gerrits, Erkang (Eric) Zhu, and Saleema Amershi · 2025
Closest in time.
Shaokun Zhang, Ming Yin, Jieyu Zhang, Jiale Liu, Zhiguang Han, Jingyang Zhang, Beibin Li, Chi Wang, Huazheng Wang, Yiran Chen, and Qingyun Wu · 2025
Closest in time.
A2a: A new era of agent interoperability, April 2025
Rao Surapaneni, Miku Jha, Michael Vakoc, and Todd Segal · 2025
Closest in time.
Saaket Agashe, Yue Fan, Anthony Reyna, and Xin Eric Wang · 2025
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen · 2095
Closest in time.