Fetching the paper…
Reading the bibliography…
Multi-agent AI systems powered by large language models (LLMs) are increasingly applied to solve complex tasks.
The neural basis of economic decision-making in the ultimatum game
Sanfey, A. G., Rilling, J. K., Aronson, J. A., Nystrom, L. E., and Cohen, J. D · 2003
Earlier work this paper cites.
Counterfactual multi-agent policy gradients
Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Earlier work this paper cites.
Decoupling strategy and generation in negotiation dialogues
He, H., Chen, D., Balakrishnan, A., and Liang, P · 2018
Earlier work this paper cites.
Irving, G., Christiano, P., and Amodei, D · 2018
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., and Lu, X · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Chen, W., Ma, X., Wang, X., and Cohen, W. W · 2022
Earlier work this paper cites.
Large language models can self-improve
Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J · 2022
Earlier work this paper cites.
Chatgpt: Optimizing language models for dialogue
Schulman, J., Zoph, B., Kim, C., Hilton, J., Menick, J., Weng, J., Uribe, J. F. C., Fedus, L., Metz, L., Pokorny, M., et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Generating sequences by learning to self-correct
Welleck, S., Lu, X., West, P., Brahman, F., Shen, T., Khashabi, D., and Choi, Y · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., Mu, J., and Goodman, N · 2022
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chan, C.-M., Chen, W., Su, Y., Yu, J., Xue, W., Zhang, S., Fu, J., and Liu, Z · 2023
Earlier work this paper cites.
Theoremqa: A theorem-driven question answering dataset
Chen, W., Yin, M., Ku, M., Lu, P., Wan, Y., Ma, X., Xu, J., Wang, X., and Xia, T · 2023
Earlier work this paper cites.
Emergent cooperation and strategy adaptation in multi-agent systems: An extended coevolutionary theory with llms
de Zarzà, I., de Curtò, J., Roig, G., Manzoni, P., and Calafate, C. T · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
Hao, S., Gu, Y., Ma, H., Hong, J. J., Wang, Z., Wang, D. Z., and Hu, Z · 2023
Cited alongside, same era.
Encouraging divergent thinking in large language models through multi-agent debate
Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Shi, S., and Tu, Z · 2023
Cited alongside, same era.
Gpt-4 technical report. arxiv 2303.08774
OpenAI, R · 2023
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et al · 2024
Later among the works it cites.
Iterative reasoning preference optimization
Pang, R. Y., Yuan, W., Cho, K., He, H., Sukhbaatar, S., and Weston, J · 2024
Later among the works it cites.
Regenesis: Llms can grow into reasoning generalists via self-improvement
Peng, X., Xia, C., Yang, X., Xiong, C., Wu, C.-S., and Xing, C · 2024
Later among the works it cites.
Self-refinement of language models from external proxy metrics feedback
Ramji, K., Lee, Y.-S., Astudillo, R. F., Sultan, M. A., Naseem, T., Munawar, A., Florian, R., and Roukos, S · 2024
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Zhang, S., Zhu, E., Li, B., Jiang, L., Zhang, X., and Wang, C · 2023
Cited alongside, same era.
Teaching language models to self-improve through interactive demonstrations
Yu, X., Peng, B., Galley, M., Gao, J., and Yu, Z · 2023
Cited alongside, same era.
Graph of thoughts: Solving elaborate problems with large language models
Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Podstawski, M., Gianinazzi, L., Gajda, J., Lehmann, T., Niewiadomski, H., Nyczyk, P., et al · 2024
Cited alongside, same era.
How well can llms negotiate? negotiationarena platform and analysis
Bianchi, F., Chia, P. J., Yuksekgonul, M., Tagliabue, J., Jurafsky, D., and Zou, J · 2024
Cited alongside, same era.
Scalable multi-robot collaboration with large language models: Centralized or decentralized systems?
Chen, Y., Arkin, J., Zhang, Y., Roy, N., and Fan, C · 2024
Cited alongside, same era.
Combating adversarial attacks with multi-agent debate
Chern, S., Fan, Z., and Liu, A · 2024
Cited alongside, same era.
Large language model based multi-agents: A survey of progress and challenges
Guo, T., Chen, X., Wang, Y., Chang, R., Pei, S., Chawla, N. V., Wiest, O., and Zhang, X · 2024
Cited alongside, same era.
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2024
Later among the works it cites.
Should we be going mad? a look at multi-agent debate strategies for llms
Smit, A. P., Grinsztajn, N., Duckworth, P., Barrett, T. D., and Pretorius, A · 2024
Later among the works it cites.
Llm-based multi-agent reinforcement learning: Current and future directions
Sun, C., Huang, S., and Pompili, D · 2024
Later among the works it cites.
The virtual lab: Ai agents design new sars-cov-2 nanobodies with experimental validation
Swanson, K., Wu, W., Bulaong, N. L., Pak, J. E., and Zou, J · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Later among the works it cites.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Later among the works it cites.
Textgrad: Automatic" differentiation" via text
Yuksekgonul, M., Bianchi, F., Boen, J., Liu, S., Huang, Z., Guestrin, C., and Zou, J · 2024
Later among the works it cites.
Small language models need strong verifiers to self-correct reasoning
Zhang, Y., Khalifa, M., Logeswaran, L., Kim, J., Lee, M., Lee, H., and Wang, L · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
Multiagent finetuning: Self improvement with diverse reasoning chains
Subramaniam, V., Du, Y., Tenenbaum, J. B., Torralba, A., Li, S., and Mordatch, I · 2025
Closest in time.