Fetching the paper…
Reading the bibliography…
Multi-agent debate (MAD) has gained significant attention as a promising line of research to improve the factual accuracy and reasoning capabilities of large language models (LLMs).
Learning to solve arithmetic word problems with verb categorization
Hosseini, M. J · 2014
Earlier work this paper cites.
Parsing algebraic word problems into equations
Koncel-Kedziorski, R · 2015
Earlier work this paper cites.
Solving general arithmetic word problems
Roy, S · 2016
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Ling, W · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P · 2018
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Talmor, A · 2019
Earlier work this paper cites.
The box is in the pen: Evaluating commonsense reasoning in neural machine translation
He, J · 2020
Earlier work this paper cites.
Unsupervised evaluation of interactive dialog with dialogpt
Mehri, S · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K · 2021
Earlier work this paper cites.
Are nlp models really able to solve simple math word problems?
Patel, A · 2021
Earlier work this paper cites.
Language models are multilingual chain-of-thought reasoners
Shi, F · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S · 2022
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Chiang, C.-H · 2023
Cited alongside, same era.
Improving factuality and reasoning in language models through multiagent debate
Du, Y · 2023
Cited alongside, same era.
Topical-chat: Towards knowledge-grounded open-domain conversations
Comm: Collaborative multi-agent, multi-reasoning-path prompting for complex problem solving
Chen, P · 2024
Later among the works it cites.
The llama 3 herd of models
Dubey, A · 2024
Later among the works it cites.
Position: Why we must rethink empirical research in machine learning
Herrmann, M · 2024
Later among the works it cites.
Debating with more persuasive llms leads to more truthful answers
Khan, A · 2024
Later among the works it cites.
Improving multi-agent debate with sparse communication topology
Li, Y · 2024
Later among the works it cites.
Encouraging divergent thinking in large language models through multi-agent debate
Liang, T · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gopalakrishnan, K · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Madaan, A · 2023
Cited alongside, same era.
Let models speak ciphers: Multiagent debate through embeddings
Pham, C · 2023
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A · 2023
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Wang, X · 2023
Cited alongside, same era.
Examining inter-consistency of large language models collaboration: An in-depth analysis via debate
Xiong, K · 2023
Cited alongside, same era.
Exchange-of-thought: Enhancing large language model capabilities through cross-model communication
Yin, Z · 2023
Cited alongside, same era.
Liu, A · 2024
Later among the works it cites.
Scaling large-language-model-based multi-agent collaboration
Qian, C · 2024
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N · 2024
Later among the works it cites.
Should we be going mad? a look at multi-agent debate strategies for llms
Smit, A. P · 2024
Later among the works it cites.
Exploring collaboration mechanisms for LLM agents: A social psychology view
Zhang, J · 2024
Later among the works it cites.
AGIEval: A human-centric benchmark for evaluating foundation models
Zhong, W · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D · 2025
Closest in time.