Fetching the paper…
Reading the bibliography…
Large language model-based multi-agent systems have shown great abilities across various tasks due to the collaboration of expert agents, each focusing on a specific domain.
Hierarchical structure and search in complex organizations
Mihm, J., Loch, C. H., Wilkinson, D., and Huberman, B. A · 2010
Earlier work this paper cites.
The resilient organization
Boin, A. and Van Eeten, M. J · 2013
Earlier work this paper cites.
Team resilience: How teams flourish under pressure
Alliger, G. M., Cerasoli, C. P., Tannenbaum, S. I., and Vessey, W. B · 2015
Earlier work this paper cites.
Communication and the optimality of hierarchy in organizations
Yang, H. and Zhang, L · 2019
Earlier work this paper cites.
Workplace team resilience: A systematic review and conceptual development
Hartwig, A., Clarke, S., Johnson, S., and Willis, S · 2020
Earlier work this paper cites.
The box is in the pen: Evaluating commonsense reasoning in neural machine translation
He, J., Wang, T., Xiong, D., and Liu, Q · 2020
Earlier work this paper cites.
Bleurt: Learning robust metrics for text generation
Sellam, T., Das, D., and Parikh, A · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Learning compact metrics for mt
Pu, A., Chung, H. W., Parikh, A., Gehrmann, S., and Sellam, T · 2021
Earlier work this paper cites.
Reliability testing for natural language processing systems
Tan, S., Joty, S., Baxter, K., Taeihagh, A., Bennett, G. A., and Kan, M.-Y · 2021
Earlier work this paper cites.
How flat can it get? from better at flatter to the promise of the decentralized, boundaryless organization
Alexy, O · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Earlier work this paper cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Pal, A., Umapathi, L. K., and Sankarasubbu, M · 2022
Earlier work this paper cites.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Earlier work this paper cites.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Guha, N., Nyarko, J., Ho, D., Ré, C., Chilton, A., Chohlas-Wood, A., Peters, A., Waldon, B., Rockmore, D., Zambrano, D., et al · 2023
Earlier work this paper cites.
Is chatgpt a good translator? a preliminary study
Jiao, W., Wang, W., Huang, J.-t., Wang, X., and Tu, Z · 2023
Cited alongside, same era.
Camel: Communicative agents for” mind” exploration of large language model society
Li, G., Hammoud, H., Itani, H., Khizbullin, D., and Ghanem, B · 2023
Cited alongside, same era.
Leveraging word guessing games to assess the intelligence of large language models
Liang, T., He, Z., Huang, J.-t., Wang, W., Jiao, W., Wang, R., Yang, Y., Tu, Z., Shi, S., and Wang, X · 2023
Cited alongside, same era.
Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation
Liu, J., Xia, C. S., Wang, Y., and Zhang, L · 2023
Cited alongside, same era.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
Agent hospital: A simulacrum of hospital with evolvable medical agents
Li, J., Wang, S., Zhang, M., Li, W., Lai, Y., Kang, X., Ma, W., and Liu, Y · 2024
Closest in time.
Encouraging divergent thinking in large language models through multi-agent debate
Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Tu, Z., and Shi, S · 2024
Closest in time.
Interintent: Investigating social intelligence of llms via intention understanding in an interactive game context
Liu, Z., Anand, A., Zhou, P., Huang, J.-t., and Zhao, J · 2024
Closest in time.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J · 2024
Closest in time.
Chatdev: Communicative agents for software development
Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., Yang, C., Chen, W., Su, Y., Cong, X., et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Evil geniuses: Delving into the safety of llm-based agents
Tian, Y., Yang, X., Zhang, J., Dong, Y., and Su, H · 2023
Cited alongside, same era.
Agents: An open-source framework for autonomous language agents
Zhou, W., Jiang, Y. E., Li, L., Wu, J., Wang, T., Qiu, S., Zhang, J., Chen, J., Wu, R., Wang, S., et al · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Multiagent collaboration attack: Investigating adversarial attacks in large language model collaborations via debate
Amayuelas, A., Yang, X., Antoniades, A., Hua, W., Pan, L., and Wang, W · 2024
Cited alongside, same era.
Multi-agent large language models for conversational task-solving
Becker, J · 2024
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chan, C.-M., Chen, W., Su, Y., Yu, J., Xue, W., Zhang, S., Fu, J., and Liu, Z · 2024
Cited alongside, same era.
Self-collaboration code generation via chatgpt
Dong, Y., Jiang, X., Jin, Z., and Li, G · 2024
Cited alongside, same era.
Unleashing the emergent cognitive synergy in large language models: A task-solving agent through multi-persona self-collaboration
Wang, Z., Mao, S., Wu, W., Ge, T., Wei, F., and Ji, H · 2024
Closest in time.
Oasis: Open agents social interaction simulations on one million agents
Yang, Z., Zhang, Z., Zheng, Z., Jiang, Y., Gan, Z., Wang, Z., Ling, Z., Chen, J., Ma, M., Dong, B., et al · 2024
Closest in time.
Infecting llm agents via generalizable adversarial attack
Yu, W., Hu, K., Pang, T., Du, C., Lin, M., and Fredrikson, M · 2024
Closest in time.
Psysafe: A comprehensive framework for psychological-based attack, defense, and evaluation of multi-agent system safety
Zhang, Z., Zhang, Y., Li, L., Gao, H., Wang, L., Lu, H., Zhao, F., Qiao, Y., and Shao, J · 2024
Closest in time.
Is this the real life? is this just fantasy? the misleading success of simulating social interactions with llms
Zhou, X., Su, Z., Eisape, T., Kim, H., and Sap, M · 2024
Closest in time.
Gptswarm: Language agents as optimizable graphs
Zhuge, M., Wang, W., Kirsch, L., Faccio, F., Khizbullin, D., and Schmidhuber, J · 2024
Closest in time.
Competing large language models in multi-agent gaming environments
Huang, J.-t., Li, E. J., Lam, M. H., Liang, T., Wang, W., Yuan, Y., Jiao, W., Wang, X., Tu, Z., and Lyu, M. R · 2025
Closest in time.
Mao, J., Meng, F., Duan, Y., Yu, M., Jia, X., Fang, J., Liang, Y., Wang, K., and Wen, Q · 2025
Closest in time.
Multi-agent collaboration mechanisms: A survey of llms
Tran, K.-T., Dao, D., Nguyen, M.-D., Pham, Q.-V., O’Sullivan, B., and Nguyen, H. D · 2025
Closest in time.
A survey on trustworthy llm agents: Threats and countermeasures
Yu, M., Meng, F., Zhou, X., Wang, S., Mao, J., Pang, L., Chen, T., Wang, K., Li, X., Zhang, Y., et al · 2025
Closest in time.
Corba: Contagious recursive blocking attacks on multi-agent systems based on large language models
Zhou, Z., Li, Z., Zhang, J., Zhang, Y., Wang, K., Liu, Y., and Guo, Q · 2025
Closest in time.