Fetching the paper…
Reading the bibliography…
Leveraging multiple large language model (LLM) agents has shown to be a promising approach for tackling complex tasks, while the effective design of multiple agents for a particular application remains an art.
Evolutionary psychology: Controversies, questions, prospects, and limitations
Confer, J. C., Easton, J. A., Fleischman, D. S., Goetz, C. D., Lewis, D. M., Perilloux, C., and Buss, D. M · 2010
Earlier work this paper cites.
An experimental study of team size and performance on a complex task
Mao, A., Mason, W., Suri, S., and Watts, D. J · 2016
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Network neuroscience and the adapted mind: Rethinking the role of network theories in evolutionary psychology
Elimari, N. and Lafargue, G · 2020
Earlier work this paper cites.
Deep learning for source code modeling and generation: Models, applications, and challenges
Le, T. H., Chen, H., and Babar, M. A · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
A survey on in-context learning
Dong, Q., Li, L., Dai, D., Zheng, C., Wu, Z., Chang, B., Sun, X., Xu, J., and Sui, Z · 2022
Earlier work this paper cites.
Large language models are reasoning teachers
Ho, N., Schmid, L., and Yun, S.-Y · 2022
Earlier work this paper cites.
Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Valmeekam, K., Olmo, A., Sreedharan, S., and Kambhampati, S · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2022
Earlier work this paper cites.
Github — babyagi
BabyAGI · 2023
Earlier work this paper cites.
Bharadhwaj, H., Vakil, J., Sharma, M., Gupta, A., Tulsiani, S., and Kumar, V · 2023
Earlier work this paper cites.
Large language models as tool makers
Cai, T., Wang, X., Ma, T., Chen, X., and Zhou, D · 2023
Earlier work this paper cites.
Autoagents: A framework for automatic agent generation
Chen, G., Dong, S., Shu, Y., Zhang, G., Sesay, J., Karlsson, B. F., Fu, J., and Shi, Y · 2023
Earlier work this paper cites.
Why can gpt learn in-context? language models secretly perform gradient descent as meta-optimizers
Dai, D., Sun, Y., Dong, L., Hao, Y., Ma, S., Sui, Z., and Wei, F · 2023
Earlier work this paper cites.
Educhat: A large-scale language model-based chatbot system for intelligent education
Dan, Y., Lei, Z., Gu, Y., Li, Y., Yin, J., Lin, J., Ye, L., Tie, Z., Zhou, Y., Wang, Y., et al · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I · 2023
Cited alongside, same era.
Bridging the gap: A survey on integrating (human) feedback for natural language generation
Fernandes, P., Madaan, A., Liu, E., Farinhas, A., Martins, P. H., Bertsch, A., de Souza, J. G., Zhou, S., Wu, T., Neubig, G., et al · 2023
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., and Wang, H · 2023
Cited alongside, same era.
Metagpt: Meta programming for multi-agent collaborative framework
Hong, S., Zheng, X., Chen, J., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., et al · 2023
Cited alongside, same era.
Metatool benchmark for large language models: Deciding whether to use tools and which to use
Github — autogenbench
AutoGenBench · 2024
Closest in time.
Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors
Chen, W., Su, Y., Zuo, J., Yang, C., Yuan, C., Chan, C.-M., Yu, H., Lu, Y., Hung, Y.-H., Qian, C., Qin, Y., Cong, X., Xie, R., Liu, Z., Sun, M., and Zhou, J · 2024
Closest in time.
Multimodal web navigation with instruction-finetuned foundation models
Furuta, H., Lee, K.-H., Nachum, O., Matsuo, Y., Faust, A., Gu, S. S., and Gur, I · 2024
Closest in time.
Github — autogen: Gaia orchestrator
GAIA_Orchestrator · 2024
Closest in time.
Data interpreter: An llm agent for data science
Hong, S., Lin, Y., Liu, B., Wu, B., Li, D., Chen, J., Zhang, J., Wang, J., Zhang, L., Zhuge, M., et al · 2024
Closest in time.
m&m’s: A benchmark to evaluate tool-use for multi-step multi-modal tasks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huang, Y., Shi, J., Li, Y., Fan, C., Wu, S., Zhang, Q., Liu, Y., Zhou, P., Wan, Y., Gong, N. Z., et al · 2023
Cited alongside, same era.
Encouraging divergent thinking in large language models through multi-agent debate
Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Tu, Z., and Shi, S · 2023
Cited alongside, same era.
Gaia: a benchmark for general ai assistants
Mialon, G., Fourrier, C., Swift, C., Wolf, T., LeCun, Y., and Scialom, T · 2023
Cited alongside, same era.
Learning deductive reasoning from synthetic corpus based on formal logic
Morishita, T., Morio, G., Yamaguchi, A., and Sogawa, Y · 2023
Cited alongside, same era.
Art: Automatic multi-step reasoning and tool-use for large language models
Paranjape, B., Lundberg, S., Singh, S., Hajishirzi, H., Zettlemoyer, L., and Ribeiro, M. T · 2023
Cited alongside, same era.
In-context retrieval-augmented language models
Ram, O., Levine, Y., Dalmedigos, I., Muhlgay, D., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2023
Cited alongside, same era.
Branch-solve-merge improves large language model evaluation and generation
Saha, S., Levy, O., Celikyilmaz, A., Bansal, M., Weston, J., and Li, X · 2023
Cited alongside, same era.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Song, C. H., Wu, J., Washington, C., Sadler, B. M., Chao, W.-L., and Su, Y · 2023
Cited alongside, same era.
Ma, Z., Huang, W., Zhang, J., Gupta, T., and Krishna, R · 2024
Closest in time.
GAIA: a benchmark for general AI assistants
Mialon, G., Fourrier, C., Wolf, T., LeCun, Y., and Scialom, T · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2024
Closest in time.
Shi, W., Xu, R., Zhuang, Y., Yu, Y., Zhang, J., Wu, H., Zhu, Y., Ho, J., Yang, C., and Wang, M. D · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2024
Closest in time.
Adaplanner: Adaptive planning from feedback with language models
Sun, H., Zhuang, Y., Kong, L., Dai, B., and Zhang, C · 2024
Closest in time.
Tdag: A multi-agent framework based on dynamic task decomposition and agent generation
Wang, Y., Wu, Z., Yao, J., and Su, J · 2024
Closest in time.
Os-copilot: Towards generalist computer agents with self-improvement
Wu, Z., Han, C., Ding, Z., Weng, Z., Liu, Z., Yao, S., Yu, T., and Kong, L · 2024
Closest in time.
Travelplanner: A benchmark for real-world planning with language agents
Xie, J., Zhang, K., Chen, J., Zhu, T., Lou, R., Tian, Y., Xiao, Y., and Su, Y · 2024
Closest in time.
Benchmarking benchmark leakage in large language models
Xu, R., Wang, Z., Fan, R.-Z., and Liu, P · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Closest in time.
Gpt-4v(ision) is a generalist web agent, if grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y · 2024
Closest in time.