Fetching the paper…
Reading the bibliography…
While multi-agent debate has been proposed as a promising strategy for improving AI reasoning ability, we find that debate can sometimes be harmful rather than helpful.
Beyond the running tally: Partisan bias in political perceptions
Larry M. Bartels · 2002
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2009
Earlier work this paper cites.
Partisan bias in factual beliefs about politics
John G. Bullock, Alan S. Gerber, Seth J. Hill, and Gregory A. Huber · 2015
Earlier work this paper cites.
You cannot be serious: The impact of accuracy incentives on partisan bias in reports of economic perceptions
Markus Prior, Gaurav Sood, and Kabir Khanna · 2015
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge, 2019
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate, 2023
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch · 2023
Earlier work this paper cites.
Improving language model negotiation with self-play and in-context learning from ai feedback
Yao Fu, Hao Peng, Tushar Khot, and Mirella Lapata · 2023
Earlier work this paper cites.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou · 2023
Earlier work this paper cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Cited alongside, same era.
Camel: Communicative agents for "mind" exploration of large language model society, 2023
Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem · 2023
Cited alongside, same era.
Encouraging divergent thinking in large language models through multi-agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2023
Cited alongside, same era.
A survey of reinforcement learning from human feedback, 2024
Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier · 2024
Later among the works it cites.
On scalable oversight with weak llms judging strong llms, 2024
Zachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen, Samuel Albanie, Jannis Bulian, Rishabh Agarwal, David Lindner, Yunhao Tang, Noah D. Goodman, and Rohin Shah · 2024
Later among the works it cites.
Debating with more persuasive llms leads to more truthful answers, 2024
Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, and Ethan Perez · 2024
Later among the works it cites.
Gpt-4o mini: Advancing cost-efficient intelligence
OpenAI · 2024
Later among the works it cites.
Mixture-of-agents enhances large language model capabilities
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Julian Michael, Salsabila Mahdi, David Rein, Jackson Petty, Julien Dirani, Vishakh Padmakumar, and Samuel R. Bowman · 2023
Cited alongside, same era.
Towards understanding sycophancy in language models, 2023
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez · 2023
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models, 2023
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation, 2023
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang · 2023
Cited alongside, same era.
Alfonso Amayuelas, Xianjun Yang, Antonis Antoniades, Wenyue Hua, Liangming Pan, and William Wang · 2024
Cited alongside, same era.
Multi-llm debate: Framework, principals, and interventions
Andrew Estornell and Yang Liu · 2024
Cited alongside, same era.
The llama 3 herd of models, 2024
Aaron Grattafiori et al · 2024
Cited alongside, same era.
When can llms actually correct their own mistakes? a critical survey of self-correction of llms
Ryo Kamoi, Yusen Zhang, Nan Zhang, Jiawei Han, and Rui Zhang · 2024
Cited alongside, same era.
Junlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang, and James Zou · 2024
Later among the works it cites.
Mahak Agarwal and Divyam Khanna · 2025
Closest in time.
Acc-collab: An actor-critic approach to multi-agent llm collaboration, 2025
Andrew Estornell, Jean-Francois Ton, Yuanshun Yao, and Yang Liu · 2025
Closest in time.
Enhancing llm reasoning with multi-path collaborative reactive and reflection agents, 2025
Chengbo He, Bochao Zou, Xin Li, Jiansheng Chen, Junliang Xing, and Huimin Ma · 2025
Closest in time.
Multiagent finetuning: Self improvement with diverse reasoning chains
Vighnesh Subramaniam, Yilun Du, Joshua B Tenenbaum, Antonio Torralba, Shuang Li, and Igor Mordatch · 2025
Closest in time.
Multi-agent collaboration mechanisms: A survey of llms
Khanh-Tung Tran, Dung Dao, Minh-Duong Nguyen, Quoc-Viet Pham, Barry O’Sullivan, and Hoang D Nguyen · 2025
Closest in time.
Hanqing Yang, Jingdi Chen, Marie Siew, Tania Lorido-Botran, and Carlee Joe-Wong · 2025
Closest in time.
Peacemaker or troublemaker: How sycophancy shapes multi-agent debate, 2025
Binwei Yao, Chao Shang, Wanyu Du, Jianfeng He, Ruixue Lian, Yi Zhang, Hang Su, Sandesh Swamy, and Yanjun Qi · 2025
Closest in time.