Fetching the paper…
Reading the bibliography…
Modern AI agents, driven by advances in large foundation models, promise to enhance our productivity and transform our lives by augmenting our knowledge and capabilities.
Intelligent agents: theory and practice
M. Wooldridge and N. R. Jennings · 1995
Earlier work this paper cites.
Stuart russell and peter norvig, artificial intelligence: A modern approach
N. J. Nilsson · 1996
Earlier work this paper cites.
Applications of intelligent agents
N. R. Jennings and M. Wooldridge · 1998
Earlier work this paper cites.
Implementing agent teams in dynamic multiagent environments
M. Tambe · 1998
Earlier work this paper cites.
The evolution of sharedplans
B. J. Grosz and S. Kraus · 1999
Earlier work this paper cites.
Multiagent systems: A survey from a machine learning perspective
P. Stone and M. Veloso · 2000
Earlier work this paper cites.
Adjustable autonomy in real-world multi-agent environments
P. Scerri, D. V. Pynadath, and M. Tambe · 2001
Earlier work this paper cites.
An introduction to multiagent systems
B. Messing · 2002
Earlier work this paper cites.
World of bits: An open-domain platform for web-based agents
T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang · 2017
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Earlier work this paper cites.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, et al · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Earlier work this paper cites.
Github — babyagi
BabyAGI · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web, 2023
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su · 2023
Earlier work this paper cites.
Aligning offline metrics and human judgments of value for code generation models
V. Dibia, A. Fourney, G. Bansal, F. Poursabzi-Sangdeh, H. Liu, and S. Amershi · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch · 2023
Earlier work this paper cites.
Metagpt: Meta programming for multi-agent collaborative framework
S. Hong, X. Zheng, J. Chen, Y. Cheng, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, et al · 2023
Earlier work this paper cites.
Camel: Communicative agents for ”mind” exploration of large scale language model society, 2023
G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem · 2023
Earlier work this paper cites.
Encouraging divergent thinking in large language models through multi-agent debate, 2023
T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, Z. Tu, and S. Shi · 2023
Earlier work this paper cites.
Gaia: a benchmark for general ai assistants, 2023
G. Mialon, C. Fourrier, C. Swift, T. Wolf, Y. LeCun, and T. Scialom · 2023
Earlier work this paper cites.
Gaia: benchmark for general ai assistants
G. Mialon, C. Fourrier, C. Swift, T. Wolf, Y. LeCun, and T. Scialom · 2023
Earlier work this paper cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Earlier work this paper cites.
Art: Automatic multi-step reasoning and tool-use for large language models
B. Paranjape, S. Lundberg, S. Singh, H. Hajishirzi, L. Zettlemoyer, and M. T. Ribeiro · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior, 2023
J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein · 2023
Earlier work this paper cites.
Tool learning with foundation models, 2023
Y. Qin, S. Hu, Y. Lin, W. Chen, N. Ding, G. Cui, Z. Zeng, Y. Huang, C. Xiao, C. Han, Y. R. Fung, Y. Su, H. Wang, C. Qian, R. Tian, K. Zhu, S. Liang, X. Shen, B. Xu, Z. Zhang, Y. Ye, B. Li, Z. Tang, J. Yi, Y. Zhu, Z. Dai, L. Yan, X. Cong, Y. Lu, W. Zhao, Y. Huang, J. Yan, X. Han, X. Sun, D. Li, J. Phang, C. Yang, T. Wu, H. Ji, Z. Liu, and M. Sun · 2023
Earlier work this paper cites.
Toolllm: Facilitating large language models to master 16000+ real-world apis, 2023
Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, S. Zhao, R. Tian, R. Xie, J. Zhou, M. Gerstein, D. Li, Z. Liu, and M. Sun · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools, 2023
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Earlier work this paper cites.
Multi-agent collaboration: Harnessing the power of intelligent llm agents, 2023
Y. Talebirad and A. Nadiri · 2023
Cited alongside, same era.
A survey on large language model based autonomous agents
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, et al · 2023
Cited alongside, same era.
The rise and potential of large language model based agents: A survey, 2023
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, R. Zheng, X. Fan, X. Wang, L. Xiong, Y. Zhou, W. Wang, C. Jiang, Y. Zou, X. Liu, Z. Yin, S. Dou, R. Weng, W. Cheng, Q. Zhang, W. Qin, Y. Zheng, X. Qiu, X. Huang, and T. Gui · 2023
Cited alongside, same era.
Auto-gpt for online decision making: Benchmarks and additional opinions, 2023
H. Yang, S. Yue, and Y. He · 2023
Cited alongside, same era.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao · 2023
The landscape of emerging ai agent architectures for reasoning, planning, and tool calling: A survey
T. Masterman, S. Besen, M. Sawtell, and A. Chao · 2024
Closest in time.
Autonomous evaluation and refinement of digital agents, 2024
J. Pan, Y. Zhang, N. Tomlin, Y. Zhou, S. Levine, and A. Suhr · 2024
Closest in time.
Webcanvas: Benchmarking web agents in online environments, 2024
Y. Pan, D. Kong, S. Zhou, C. Cui, Y. Leng, B. Jiang, H. Liu, Y. Shang, S. Zhou, T. Wu, and Z. Wu · 2024
Closest in time.
REFINER: Reasoning feedback on intermediate representations
D. Paul, M. Ismayilzada, M. Peyrard, B. Borges, A. Bosselut, R. West, and B. Faltings · 2024
Closest in time.
Agent q: Advanced reasoning and learning for autonomous ai agents, 2024
P. Putta, E. Mills, N. Garg, S. Motwani, C. Finn, D. Garg, and R. Rafailov · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Webshop: Towards scalable real-world web interaction with grounded language agents, 2023
S. Yao, H. Chen, J. Yang, and K. Narasimhan · 2023
Cited alongside, same era.
Tree of thoughts: Deliberate problem solving with large language models, 2023
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2023
Cited alongside, same era.
Agenttuning: Enabling generalized agent abilities for llms, 2023
A. Zeng, M. Liu, R. Lu, B. Wang, X. Liu, Y. Dong, and J. Tang · 2023
Cited alongside, same era.
Agent-e: From autonomous web navigation to foundational design principles in agentic systems, 2024
T. Abuelsaad, D. Akkil, P. Dey, A. Jagmohan, A. Vempaty, and R. Kokku · 2024
Cited alongside, same era.
Windows agent arena: Evaluating multi-modal os agents at scale, 2024
R. Bonatti, D. Zhao, F. Bonacci, D. Dupont, S. Abdali, Y. Li, Y. Lu, J. Wagle, K. Koishida, A. Bucker, L. Jang, and Z. Hui · 2024
Cited alongside, same era.
Spider2-v: How far are multimodal agents from automating data science and engineering workflows?, 2024
R. Cao, F. Lei, H. Wu, J. Chen, Y. Fu, H. Gao, X. Xiong, H. Zhang, Y. Mao, W. Hu, T. Xie, H. Xu, D. Zhang, S. Wang, R. Sun, P. Yin, C. Xiong, A. Ni, Q. Liu, V. Zhong, L. Chen, K. Yu, and T. Yu · 2024
Cited alongside, same era.
Red Cell Partners · 2024
Closest in time.
Great, now write an article about that: The crescendo multi-turn llm jailbreak attack
M. Russinovich, A. Salem, and R. Eldan · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao · 2024
Closest in time.
Step: Stacked llm policies for web actions, 2024
P. Sodhi, S. R. K. Branavan, Y. Artzi, and R. McDonald · 2024
Closest in time.
Trial and error: Exploration-based trajectory optimization for llm agents, 2024
Y. Song, D. Yin, X. Yue, J. Huang, S. Li, and B. Y. Lin · 2024
Closest in time.
Opendevin: An open platform for ai software developers as generalist agents, 2024
X. Wang, B. Li, Y. Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y. Song, B. Li, J. Singh, H. H. Tran, F. Li, R. Ma, M. Zheng, B. Qian, Y. Shao, N. Muennighoff, Y. Zhang, B. Hui, J. Lin, R. Brennan, H. Peng, H. Ji, and G. Neubig · 2024
Closest in time.
Sibyl: Simple yet effective agent framework for complex real-world reasoning, 2024
Y. Wang, T. Shen, L. Liu, and J. Xie · 2024
Closest in time.
Agent workflow memory, 2024
Z. Z. Wang, J. Mao, D. Fried, and G. Neubig · 2024
Closest in time.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, A. H. Awadallah, R. W. White, D. Burger, and C. Wang · 2024
Closest in time.
Os-copilot: Towards generalist computer agents with self-improvement, 2024
Z. Wu, C. Han, Z. Ding, Z. Weng, Z. Liu, S. Yao, T. Yu, and L. Kong · 2024
Closest in time.
Agentless: Demystifying llm-based software engineering agents
C. S. Xia, Y. Deng, S. Dunn, and L. Zhang · 2024
Closest in time.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024
T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, Y. Liu, Y. Xu, S. Zhou, S. Savarese, C. Xiong, V. Zhong, and T. Yu · 2024
Closest in time.
Swe-agent: Agent-computer interfaces enable automated software engineering
J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press · 2024
Closest in time.
Assistantbench: Can web agents solve realistic and time-consuming tasks?, 2024
O. Yoran, S. J. Amouyal, C. Malaviya, B. Bogin, O. Press, and J. Berant · 2024
Closest in time.
Ufo: A ui-focused agent for windows os interaction, 2024
C. Zhang, L. Li, S. He, X. Zhang, B. Qiao, S. Qin, M. Ma, Y. Kang, Q. Lin, S. Rajmohan, D. Zhang, and Q. Zhang · 2024
Closest in time.
Cognitive kernel: An open-source agent system towards generalist autopilots, 2024
H. Zhang, X. Pan, H. Wang, K. Ma, W. Yu, and D. Yu · 2024
Closest in time.
Webpilot: A versatile and autonomous multi-agent system for web task execution with strategic exploration, 2024
Y. Zhang, Z. Ma, Y. Ma, Z. Han, Y. Wu, and V. Tresp · 2024
Closest in time.
Autocoderover: Autonomous program improvement
Y. Zhang, H. Ruan, Z. Fan, and A. Roychoudhury · 2024
Closest in time.
You only look at screens: Multimodal chain-of-action agents, 2024
Z. Zhang and A. Zhang · 2024
Closest in time.
Z. J. Zhang, E. Schoop, J. Nichols, A. Mahajan, and A. Swearngin · 2024
Closest in time.
Webarena: A realistic web environment for building autonomous agents, 2024
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, U. Alon, and G. Neubig · 2024
Closest in time.