Fetching the paper…
Reading the bibliography…
Leading AI developers and startups are increasingly deploying agentic AI systems that can plan and execute complex tasks with limited human involvement.
Behavior, purpose and teleology
Rosenblueth, A., Wiener, N., and Bigelow, J · 1943
Earlier work this paper cites.
An Introduction to Cybernetics
Ashby, W. R · 1956
Earlier work this paper cites.
Cybernetics: Or Control and Communication in the Animal and the Machine
Wiener, N · 1961
Earlier work this paper cites.
The intentional stance
Dennett, D. C · 1989
Earlier work this paper cites.
Designing autonomous agents: Theory and practice from biology to engineering and back
Maes, P · 1990
Earlier work this paper cites.
Modeling rational agents within a bdi-architecture
Rao, A. S. and Georgeff, M. P · 1991
Earlier work this paper cites.
Modeling adaptive autonomous agents
Maes, P · 1993
Earlier work this paper cites.
Artificial life meets entertainment: lifelike autonomous agents
Maes, P · 1995
Earlier work this paper cites.
Intelligent agents: Theory and practice
Wooldridge, M. and Jennings, N. R · 1995
Earlier work this paper cites.
Is it an agent, or just a program?: A taxonomy for autonomous agents
Franklin, S. and Graesser, A · 1996
Earlier work this paper cites.
On agent-based software engineering
Jennings, N. R · 2000
Earlier work this paper cites.
Cosmetic compliance and the failure of negotiated governance
Krawiec, K. D · 2003
Earlier work this paper cites.
Scrutiny, norms, and selective disclosure: A global study of greenwashing
Marquis, C., Toffel, M. W., and Zhou, Y · 2016
Earlier work this paper cites.
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., and Crawford, K · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Model cards for model reporting
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., and Gebru, T · 2019
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Russell, S. and Norvig, P · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Earlier work this paper cites.
Preventing repeated real world ai failures by cataloging incidents: The ai incident database
McGregor, S · 2021
Earlier work this paper cites.
Why ai is harder than we think
Mitchell, M · 2021
Earlier work this paper cites.
Reward reports for reinforcement learning
Gilbert, T. K., Lambert, N., Dean, S., Zick, T., and Snoswell, A · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Taxonomy of risks posed by language models
Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., et al · 2022
Earlier work this paper cites.
Autonomous chemical research with large language models
Boiko, D. A., MacKnight, R., Kline, B., and Gomes, G · 2023
Earlier work this paper cites.
Harms from increasingly agentic algorithmic systems
Chan, A., Salganik, R., Markelius, A., Pang, C., Rajkumar, N., Krasheninnikov, D., Langosco, L., He, Z., Duan, Y., Carroll, M., et al · 2023
Earlier work this paper cites.
Mind2web: towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2023
Earlier work this paper cites.
What if gpt4 became autonomous: The auto-gpt project and use cases
Fırat, M. and Kuleli, S · 2023
Earlier work this paper cites.
A real-world webagent with planning, long context understanding, and program synthesis
Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., and Faust, A · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K · 2023
Earlier work this paper cites.
Discovering agents
Kenton, Z., Kumar, R., Farquhar, S., Richens, J., MacDermott, M., and Everitt, T · 2023
Cited alongside, same era.
The data provenance initiative: A large scale audit of dataset licensing & attribution in ai
Longpre, S., Mahari, R., Chen, A., Obeng-Marnu, N., Sileo, D., Brannon, W., Muennighoff, N., Khazam, N., Kabbara, J., Perisetla, K., et al · 2023
Cited alongside, same era.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., et al · 2023
Cited alongside, same era.
Identifying the risks of lm agents with an lm-emulated sandbox
Ruan, Y., Dong, H., Wang, A., Pitis, S., Zhou, Y., Ba, J., Dubois, Y., Maddison, C. J., and Hashimoto, T · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
Refusal-trained llms are easily jailbroken as browser agents
Kumar, P., Lau, E., Vijayakumar, S., Trinh, T., Team, S. R., Chang, E., Robinson, V., Hendryx, S., Zhou, S., Fredrikson, M., et al · 2024
Later among the works it cites.
Frontier ai ethics: Anticipating and evaluating the societal impacts of generative agents
Lazar, S · 2024
Later among the works it cites.
The ai scientist: Towards fully automated open-ended scientific discovery
Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., and Ha, D · 2024
Later among the works it cites.
Should users trust advanced ai assistants? justified trust as a function of competence and alignment
Manzini, A., Keeling, G., Marchal, N., McKee, K. R., Rieser, V., and Gabriel, I · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Cited alongside, same era.
Practices for governing agentic ai systems
Shavit, Y., Agarwal, S., Brundage, M., Adler, S., O’Keefe, C., Campbell, R., Lee, T., Mishkin, P., Eloundou, T., Hickey, A., et al · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., and Yao, S · 2023
Cited alongside, same era.
Evaluating the social impact of generative ai systems in systems and society
Solaiman, I., Talat, Z., Agnew, W., Ahmad, L., Baker, D., Blodgett, S. L., Chen, C., Daumé III, H., Dodge, J., Duan, I., et al · 2023
Cited alongside, same era.
Cognitive architectures for language agents
Sumers, T. R., Yao, S., Narasimhan, K., and Griffiths, T. L · 2023
Cited alongside, same era.
Pearl: Prompting large language models to plan and execute actions over long documents
Sun, S., Liu, Y., Wang, S., Zhu, C., and Iyyer, M · 2023
Cited alongside, same era.
The rise and potential of large language model based agents: A survey
Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., et al · 2023
Cited alongside, same era.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Cited alongside, same era.
McKernon, E., Glasser, G., Cheng, D., and Hadfield, G · 2024
Later among the works it cites.
Evaluating frontier models for dangerous capabilities
Phuong, M., Aitchison, M., Catt, E., Cogan, S., Kaskasoli, A., Krakovna, V., Lindner, D., Rahtz, M., Assael, Y., Hodkinson, S., et al · 2024
Later among the works it cites.
Falcon-ui: Understanding gui before following user instructions
Shen, H., Liu, C., Li, G., Wang, X., Zhou, Y., Ma, C., and Ji, X · 2024
Later among the works it cites.
Siegel, Z. S., Kapoor, S., Nagdir, N., Stroebl, B., and Narayanan, A · 2024
Later among the works it cites.
Slattery, P., Saeri, A. K., Grundy, E. A., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., and Thompson, N · 2024
Later among the works it cites.
Language agents: Foundations, prospects, and risks
Su, Y., Yang, D., Yao, S., and Yu, T · 2024
Later among the works it cites.
A survey on large language model based autonomous agents
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., et al · 2024
Later among the works it cites.
Re-bench: Evaluating frontier ai r&d capabilities of language model agents against human experts
Wijk, H., Lin, T., Becker, J., Jawhar, S., Parikh, N., Broadley, T., Chan, L., Chen, M., Clymer, J., Dhyani, J., et al · 2024
Later among the works it cites.
Improving governance outcomes through ai documentation: Bridging theory and practice
Winecoff, A. A. and Bogen, M · 2024
Later among the works it cites.
Os-copilot: Towards generalist computer agents with self-improvement
Wu, Z., Han, C., Ding, Z., Weng, Z., Liu, Z., Yao, S., Yu, T., and Kong, L · 2024
Later among the works it cites.
Theagentcompany: benchmarking llm agents on consequential real world tasks
Xu, F. F., Song, Y., Li, B., Tang, Y., Jain, K., Bao, M., Wang, Z. Z., Zhou, X., Guo, Z., Cao, M., et al · 2024
Later among the works it cites.
Language Agents: From Next-Token Prediction to Digital Automation
Yao, S · 2024
Later among the works it cites.
Assistantbench: Can web agents solve realistic and time-consuming tasks?
Yoran, O., Amouyal, S. J., Malaviya, C., Bogin, B., Press, O., and Berant, J · 2024
Later among the works it cites.
The shift from models to compound ai systems
Zaharia, M., Khattab, O., Chen, L., Davis, J. Q., Miller, H., Potts, C., Zou, J., Carbin, M., Frankle, J., Rao, N., and Ghodsi, A · 2024
Later among the works it cites.
Honeycomb: A flexible llm-based agent system for materials science
Zhang, H., Song, Y., Hou, Z., Miret, S., and Liu, B · 2024
Later among the works it cites.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al · 2024
Later among the works it cites.
Globant code fixer agent: Whitepaper, November 2024
Bel, M. A., Ríos, J. L., Carrasco, R. A. L., Michelini, J., Milano, G., Milano, G., Pérez, M., and Pasquero, G · 2025
Closest in time.
International ai safety report, 2025
Bengio, Y., Mindermann, S., Privitera, D., Besiroglu, T., Bommasani, R., Casper, S., Choi, Y., Fox, P., Garfinkel, B., Goldfarb, D., Heidari, H., Ho, A., Kapoor, S., Khalatbari, L., Longpre, S., Manning, S., Mavroudis, V., Mazeika, M., Michael, J., Newman, J., Ng, K. Y., Okolo, C. T., Raji, D., Sastry, G., Seger, E., Skeadas, T., South, T., Strubell, E., Tramèr, F., Velasco, L., Wheeler, N., Acemoglu, D., Adekanmbi, O., Dalrymple, D., Dietterich, T. G., Felten, E. W., Fung, P., Gourinchas, P.-O., Heintz, F., Hinton, G., Jennings, N., Krause, A., Leavy, S., Liang, P., Ludermir, T., Marda, V., Margetts, H., McDermid, J., Munga, J., Narayanan, A., Nelson, A., Neppel, C., Oh, A., Ramchurn, G., Russell, S., Schaake, M., Schölkopf, B., Song, D., Soto, A., Tiedrich, L., Varoquaux, G., Yao, A., Zhang, Y.-Q., Albalawi, F., Alserkal, M., Ajala, O., Avrin, G., Busch, C., de Leon Ferreira de Carvalho, A. C. P., Fox, B., Gill, A. S., Hatip, A. H., Heikkilä, J., Jolly, G., Katzir, Z., Kitano, H., Krüger, A., Johnson, C., Khan, S. M., Lee, K. M., Ligot, D. V., Molchanovskyi, O., Monti, A., Mwamanzi, N., Nemer, M., Oliver, N., Portillo, J. R. L., Ravindran, B., Rivera, R. P., Riza, H., Rugege, C., Seoighe, C., Sheehan, J., Sheikh, H., Wong, D., and Zeng, Y · 2025
Closest in time.
Infrastructure for ai agents, 2025
Chan, A., Wei, K., Huang, S., Rajkumar, N., Perrier, E., Lazar, S., Hadfield, G. K., and Anderljung, M · 2025
Closest in time.
Kolt, N · 2025
Closest in time.
Introducing openai o1-preview, September 2024
OpenAI · 2025
Closest in time.
Sager, P. J., Meyer, B., Yan, P., von Wartburg-Kottler, R., Etaiwi, L., Enayati, A., Nobel, G., Abdulkadir, A., Grewe, B. F., and Stadelmann, T · 2025
Closest in time.
Hal: A holistic agent leaderboard for centralized and reproducible agent evaluation
Stroebl, B., Kapoor, S., and Narayanan, A · 2025
Closest in time.
Technical blog: Strengthening ai agent hijacking evaluations, January 2025
U.S. AI Safety Institute · 2025
Closest in time.