Fetching the paper…
Reading the bibliography…
This survey paper examines the recent advancements in AI agent implementations, with a focus on their ability to achieve complex goals that require enhanced reasoning, planning, and tool execution capabilities.
“Measuring Massive Multitask Language Understanding” arXiv:2009.03300 [cs]
Dan Hendrycks et al · 2009
Earlier work this paper cites.
“Training Verifiers to Solve Math Word Problems” arXiv:2110.14168 [cs]
Karl Cobbe et al · 2021
Earlier work this paper cites.
Mor Geva et al · 2021
Earlier work this paper cites.
Weize Chen et al · 2023
Earlier work this paper cites.
“MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework”, 2023
Sirui Hong et al · 2023
Earlier work this paper cites.
“SWE-bench: Can Language Models Resolve Real-World GitHub Issues?” arXiv:2310.06770 [cs]
Carlos. Jimenez et al · 2023
Earlier work this paper cites.
Fangyu Lei et al · 2023
Earlier work this paper cites.
“AgentBench: Evaluating LLMs as Agents” arXiv:2308.03688 [cs]
Xiao Liu et al · 2023
Earlier work this paper cites.
“Dynamic LLM-Agent Network: An LLM-agent Collaboration Framework with Agent Team Optimization”, 2023
Zijun Liu et al · 2023
Earlier work this paper cites.
“AI Deception: A Survey of Examples, Risks, and Potential Solutions” arXiv:2308.14752 [cs]
Peter. Park et al · 2023
Earlier work this paper cites.
“Personality Traits in Large Language Models”, 2023
Greg Serapio-García et al · 2023
Earlier work this paper cites.
“Reflexion: Language Agents with Verbal Reinforcement Learning” arXiv:2303.11366 [cs]
Noah Shinn et al · 2023
Earlier work this paper cites.
“Chain-of-Thought Prompting Elicits Reasoning in Large Language Models” arXiv:2201.11903 [cs]
Jason Wei et al · 2023
Earlier work this paper cites.
“The Rise and Potential of Large Language Model Based Agents: A Survey”, 2023
Zhiheng Xi et al · 2023
Cited alongside, same era.
“ReAct: Synergizing Reasoning and Acting in Language Models” arXiv:2210.03629 [cs]
Shunyu Yao et al · 2023
Cited alongside, same era.
“Tree of Thoughts: Deliberate Problem Solving with Large Language Models” arXiv:2305.10601 [cs]
Shunyu Yao et al · 2023
Cited alongside, same era.
“How Language Model Hallucinations Can Snowball” arXiv:2305.13534 [cs]
Muru Zhang et al · 2023
Cited alongside, same era.
Na Liu et al · 2024
Closest in time.
“yoheinakajima/babyagi” original-date: 2023-04-03T00:40:27Z, 2024
Yohei Nakajima · 2024
Closest in time.
“Learning to Use Tools via Cooperative and Interactive Agents” arXiv:2403.03031 [cs]
Zhengliang Shi et al · 2024
Closest in time.
“Systematic Biases in LLM Simulations of Debates” arXiv:2402.04049 [cs]
Amir Taubenfeld, Yaniv Dover, Roi Reichart and Ariel Goldstein · 2024
Closest in time.
“Evil Geniuses: Delving into the Safety of LLM-based Agents” arXiv:2311.11855 [cs]
Yu Tian et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andy Zhou et al · 2023
Cited alongside, same era.
Timo Birr, Christoph Pohl, Abdelrahman Younes and Tamim Asfour · 2024
Cited alongside, same era.
“Large Language Model-based Human-Agent Collaboration for Complex Task Solving”, 2024
Xueyang Feng et al · 2024
Cited alongside, same era.
“Bias and Fairness in Large Language Models: A Survey” arXiv:2309.00770 [cs]
Isabel. Gallegos et al · 2024
Cited alongside, same era.
“Efficient Tool Use with Chain-of-Abstraction Reasoning” arXiv:2401.17464 [cs]
Silin Gao et al · 2024
Cited alongside, same era.
Shahriar Golchin and Mihai Surdeanu · 2024
Cited alongside, same era.
“Embodied LLM Agents Learn to Cooperate in Organized Teams”, 2024
Xudong Guo et al · 2024
Cited alongside, same era.
“Understanding the planning of LLM agents: A survey”, 2024
Xu Huang et al · 2024
Cited alongside, same era.
Closest in time.
“Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?” arXiv:2402.18272 [cs]
Qineng Wang et al · 2024
Closest in time.
“Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation” arXiv:2402.11443 [cs]
Siyuan Wang et al · 2024
Closest in time.
Zhenhailong Wang et al · 2024
Closest in time.
“SmartPlay: A Benchmark for LLMs as Intelligent Agents” arXiv:2310.01557 [cs]
Yue Wu, Xuan Tang, Tom. Mitchell and Yuanzhi Li · 2024
Closest in time.
“(InThe)WildChat: 570K ChatGPT Interaction Logs In The Wild”
Wenting Zhao et al · 2024
Closest in time.
“DyVal 2: Dynamic Evaluation of Large Language Models by Meta Probing Agents” arXiv:2402.14865 [cs]
Kaijie Zhu et al · 2024
Closest in time.
“DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks” arXiv:2309.17167 [cs]
Kaijie Zhu et al · 2024
Closest in time.