Fetching the paper…
Reading the bibliography…
Artificial Intelligence is moving from models that only generate text to Agentic AI, where systems behave as autonomous entities that can perceive, reason, plan, and act.
N. R. Jennings, “On agent-based software engineering,” Artificial Intelligence
2000
Earlier work this paper cites.
L. Busoniu, R. Babuska, & B. De Schutter, “A comprehensive survey of multiagent reinforcement learning,” IEEE Transactions on Systems, Man, and Cybernetics
2008
Earlier work this paper cites.
Z. Yang et al. , “HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering,” EMNLP
2018
Earlier work this paper cites.
M. Shridhar et al. , “ALFWorld: Aligning Text and Embodied Environments for Interactive Learning,” International Conference on Learning Representations (ICLR)
2020
Earlier work this paper cites.
D. Hendrycks et al. , “Measuring Massive Multitask Language Understanding,” International Conference on Learning Representations (ICLR)
2021
Earlier work this paper cites.
K. Cobbe et al. , “Training Verifiers to Solve Math Word Problems,” arXiv preprint arXiv:2110.14168
2021
Earlier work this paper cites.
J. Wei et al. , “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,” Advances in Neural Information Processing Systems (NeurIPS)
2022
Earlier work this paper cites.
M. Ahn et al. , “Do As I Can, Not As I Say: Grounding Language in Robotic Affordances,” Conference on Robot Learning (CoRL)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
G. Mialon et al. , “Augmented Language Models: a Survey,” Transactions on Machine Learning Research
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
G. Mialon et al. , “GAIA: a benchmark for general AI assistants,” arXiv preprint arXiv:2311.12983
2023
Earlier work this paper cites.
N. Shinn et al. , “Reflexion: Language Agents with Verbal Reinforcement Learning,” Advances in Neural Information Processing Systems (NeurIPS)
2023
Earlier work this paper cites.
T. Schick et al. , “Toolformer: Language Models Can Teach Themselves to Use Tools,” Advances in Neural Information Processing Systems (NeurIPS)
2023
Earlier work this paper cites.
S. G. Patil, T. Zhang, X. Wang, & J. E. Gonzalez, “Gorilla: Large Language Model Connected with Massive APIs,” Advances in Neural Information Processing Systems (NeurIPS)
2023
Earlier work this paper cites.
J. S. Park et al. , “Generative Agents: Interactive Simulacra of Human Behavior,” Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST)
2023
Earlier work this paper cites.
A. Madaan et al. , “Self-Refine: Iterative Refinement with Self-Feedback,” Advances in Neural Information Processing Systems (NeurIPS)
2023
Earlier work this paper cites.
A. Zhao et al. , “Expel: LLM Agents are Experiential Learners,” arXiv preprint arXiv:2308.10144
2023
Earlier work this paper cites.
G. Li et al. , “CAMEL: Communicative Agents for ’Mind’ Exploration,” Advances in Neural Information Processing Systems (NeurIPS)
2023
Earlier work this paper cites.
C. Xu et al. , “MemGPT: Towards LLMs as operating systems,” arXiv preprint arXiv:2310.08560
2023
Earlier work this paper cites.
X. Deng et al. , “Mind2Web: Towards a Generalist Agent for the Web,” Advances in Neural Information Processing Systems (NeurIPS)
2023
Earlier work this paper cites.
S. Zhou et al. , “WebArena: A Realistic Web Environment for Autonomous Agents,” Advances in Neural Information Processing Systems (NeurIPS)
2023
Earlier work this paper cites.
J. Liang et al. , “Code as Policies: Language Model Programs for Embodied Control,” International Conference on Robotics and Automation (ICRA)
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
P. Manakul, A. Liusie, & M. J. F. Gales, “SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection,” Conference on Empirical Methods in Natural Language Processing (EMNLP)
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Wang et al. , “Executable Code Actions Elicit Better LLM Agents,” International Conference on Machine Learning (ICML)
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
S. Hong et al. , “MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework,” International Conference on Learning Representations (ICLR)
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
L. Wang et al. , “A Survey on Large Language Model based Autonomous Agents,” Frontiers of Computer Science
2024
Earlier work this paper cites.
A. Li et al. , “Agent-oriented planning in multi-agent systems,” arXiv preprint arXiv:2410.02189
2024
Earlier work this paper cites.
A. Kumar, “Building Autonomous AI Agents based AI Infrastructure,” International Journal of Computer Trends and Technology
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
S. Yao et al. , “Tree of Thoughts: Deliberate Problem Solving with Large Language Models,” Advances in Neural Information Processing Systems (NeurIPS)
2024
Earlier work this paper cites.
OpenAI, “Learning to reason with LLMs,” 2024. [Online]. Available: https://openai.com/index/learning-to-reason-with-llms/
2024
Earlier work this paper cites.
OpenAI, “o1 system card,” 2024. [Online]. Available: https://openai.com/index/openai-o1-system-card/
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
LangChain AI, “LangGraph: building language agents as graphs,” 2024. [Online]. Available: https://github.com/langchain-ai/langgraph
2024
Earlier work this paper cites.
OpenAI, “Swarm,” 2024. [Online]. Available: https://github.com/openai/swarm
2024
Earlier work this paper cites.
Anthropic, “Introducing the Model Context Protocol,” 2024. [Online]. Available: https://www.anthropic.com/news/model-context-protocol
2024
Earlier work this paper cites.
Anthropic, “Model Context Protocol specification,” 2024. [Online]. Available: https://modelcontextprotocol.io/specification
2024
Earlier work this paper cites.
Anthropic, “Developing a computer use model,” 2024. [Online]. Available: https://www.anthropic.com/news/developing-computer-use
2024
Earlier work this paper cites.
OpenAI, “Introducing SWE bench Verified,” 2024. [Online]. Available: https://openai.com/index/introducing-swe-bench-verified/
2024
Earlier work this paper cites.
SWE bench, “SWE bench leaderboards,” 2024. [Online]. Available: https://www.swebench.com/
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Later among the works it cites.
Z. Xi et al. , “The Rise and Potential of Large Language Model Based Agents: A Survey,” Science China Information Sciences
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
AI CERTs Team, “RE-Bench: Economic Efficiency Factors in Agent Evaluation,” AI CERTs Technical Report
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
G. Wang et al. , “Voyager: An Open-Ended Embodied Agent with Large Language Models,” Transactions on Machine Learning Research (TMLR)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
C. Wang et al. , “AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation Framework,” Conference on Language Modeling (COLM)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Later among the works it cites.
B. Cottier et al. , “LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks,” Epoch AI
2025
Later among the works it cites.
T. Zhang et al. , “When Hallucination Costs Millions: Benchmarking AI Agents in High-Stakes Adversarial Financial Markets,” arXiv preprint
2025
Later among the works it cites.
Scale AI Research, “Smoothing Out LLM Variance for Reliable Enterprise Evals,” Scale AI Technical Blog
2025
Later among the works it cites.
M. Wornow et al. , “Top of the CLASS: Benchmarking LLM Agents on Real-World Enterprise Tasks,” ICLR 2025 Workshop on Trustworthy LLMs
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
D. Gosmar & D. A. Dahl, “Hallucination Mitigation with Agentic AI NLP-Based Frameworks,” Available at SSRN 5086241
2025
Later among the works it cites.
2025
Later among the works it cites.
OpenAI, “GPT-5.1 Codex Max system card,” 2025. [Online]. Available: https://openai.com/index/gpt-5-1-codex-max-system-card/
2025
Later among the works it cites.
Google, “Introducing Gemini 3: Our most capable model family yet,” 2025. [Online]. Available: https://blog.google/technology/google-deepmind/gemini-3/
2025
Later among the works it cites.
Anthropic, “Claude Haiku 4.5 system card,” 2025. [Online]. Available: https://assets.anthropic.com/m/37cf170ec9d01f5e/original/Claude-Haiku-4-5-System-Card.pdf
2025
Later among the works it cites.
XLANG Lab, “Introducing OSWorld Verified,” 2025. [Online]. Available: https://xlang.ai/blog/osworld-verified
2025
Later among the works it cites.
X. Xin et al. , “CoAct 1: Computer using agents with coding as actions,” 2025. [Online]. Available: https://linxins.net/coact/
2025
Later among the works it cites.
2025
Later among the works it cites.
DeepMind Team, “Genie 3: A New Frontier for World Models,” Technical Report
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
Y. Chen et al. , “LLM-Powered SQL Agents for BI & Data Analytics,” arXiv preprint arXiv:2408.06259
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
Google DeepMind Team, “Gemini Robotics 1.5: AI Agents into the Physical World,” Technical Report
2025
Later among the works it cites.
2025
Later among the works it cites.
D. Zhou et al. , “Autonomous Agents for Scientific Discovery,” arXiv preprint arXiv:2410.13567
2025
Later among the works it cites.
T. So et al. , “Scientific Discoveries by LLM Agents,” Nature
2025
Later among the works it cites.
2025
Later among the works it cites.
S. Wang et al. , “A Survey of LLM-based Agents in Medicine,” arXiv preprint arXiv:2404.11585
2025
Later among the works it cites.
Nature Digital Medicine et al. , “Healthcare Agent: Eliciting the Power of LLMs,” Nature Digital Medicine
2025
Later among the works it cites.
K. Yuan et al. , “Agentic Large Language Models for Healthcare,” arXiv preprint arXiv:2410.15890
2025
Later among the works it cites.
ACM et al. , “MindGuard: Autonomous LLM Agent for Mental Health Using Mobile Sensor Data,” ACM Proceedings
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
ACM et al. , “Proactive Conversational AI: A Comprehensive Survey,” arXiv preprint arXiv:2405.13987
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
2025
Later among the works it cites.
Y. Zhang et al. , “Mitigating Spatial Hallucination in LLMs,” arXiv preprint arXiv:2410.13567
2025
Later among the works it cites.
D. M. Anisuzzaman et al. , “Fine-tuning large language models for specialized use cases,” Mayo Clinic Proceedings: Digital Health
2025
Later among the works it cites.
OpenAI, “OpenAI Practical Guide to Building Agents,” 2025
2025
Later among the works it cites.
J. Yang et al. , “Magma: A Foundation Model for Multimodal AI Agents,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2025
Later among the works it cites.
2025
Later among the works it cites.