Fetching the paper…
Reading the bibliography…
Recent advancements in Web AI agents have demonstrated remarkable capabilities in addressing complex web navigation tasks.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2022
Earlier work this paper cites.
Cognitive architectures for language agents
Theodore R Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L Griffiths · 2023
Earlier work this paper cites.
Language agents as hackers: Evaluating cybersecurity skills with capture the flag
John Yang, Akshara Prabhakar, Shunyu Yao, Kexin Pei, and Karthik R Narasimhan · 2023
Earlier work this paper cites.
Webarena: A realistic web environment for building autonomous agents
Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, et al · 2023
Earlier work this paper cites.
Autodan: interpretable gradient-based adversarial attacks on large language models
Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun · 2023
Earlier work this paper cites.
Agentharm: A benchmark for measuring harmfulness of llm agents
Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, Zico Kolter, Matt Fredrikson, et al · 2024
Earlier work this paper cites.
AI agents with formal security guarantees
Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Martin Vechev · 2024
Earlier work this paper cites.
Agentdojo: A dynamic environment to evaluate attacks and defenses for llm agents
Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr · 2024
Earlier work this paper cites.
WorkArena: How capable are web agents at solving common knowledge work tasks?
Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, David Vazquez, Nicolas Chapados, and Alexandre Lacoste · 2024
Earlier work this paper cites.
Llm agents can autonomously exploit one-day vulnerabilities
Richard Fang, Rohan Bindu, Akul Gupta, and Daniel Kang · 2024
Earlier work this paper cites.
Navigating the digital world as humans do: Universal visual grounding for gui agents
Boyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su · 2024
Cited alongside, same era.
Yifeng He, Ethan Wang, Yuyang Rong, Zifei Cheng, and Hao Chen · 2024
Cited alongside, same era.
Openwebagent: An open toolkit to enable web agents on large language models
Iat Long Iong, Xiao Liu, Yuxuan Chen, Hanyu Lai, Shuntian Yao, Pengbo Shen, Hao Yu, Yuxiao Dong, and Jie Tang · 2024
Cited alongside, same era.
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks, 2024
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried · 2024
Cited alongside, same era.
Refusal-trained llms are easily jailbroken as browser agents
Agent q: Advanced reasoning and learning for autonomous ai agents
Pranav Putta, Edmund Mills, Naman Garg, Sumeet Motwani, Chelsea Finn, Divyansh Garg, and Rafael Rafailov · 2024
Later among the works it cites.
Naviqate: Functionality-guided web application navigation
Mobina Shahbandeh, Parsa Alian, Noor Nashid, and Ali Mesbah · 2024
Later among the works it cites.
”do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2024
Later among the works it cites.
Learning to use tools via cooperative and interactive agents, 2024
Zhengliang Shi, Shen Gao, Xiuyi Chen, Yue Feng, Lingyong Yan, Haibo Shi, Dawei Yin, Pengjie Ren, Suzan Verberne, and Zhaochun Ren · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Priyanshu Kumar, Elaine Lau, Saranya Vijayakumar, Tu Trinh, Scale Red Team, Elaine Chang, Vaughn Robinson, Sean Hendryx, Shuyan Zhou, Matt Fredrikson, et al · 2024
Cited alongside, same era.
Eia: Environmental injection attack on generalist web agents for privacy leakage
Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang, Jiawei Zhang, Chaowei Xiao, Yuan Tian, Bo Li, and Huan Sun · 2024
Cited alongside, same era.
Breaking react agents: Foot-in-the-door attack will get you in, 2024
Itay Nakash, George Kour, Guy Uziel, and Ateret Anaby-Tavor · 2024
Cited alongside, same era.
On the prospects of incorporating large language models (llms) in automated planning and scheduling (aps)
Vishal Pallagani, Bharath Chandra Muppasani, Kaushik Roy, Francesco Fabiano, Andrea Loreggia, Keerthiram Murugesan, Biplav Srivastava, Francesca Rossi, Lior Horesh, and Amit Sheth · 2024
Cited alongside, same era.
Training software engineering agents and verifiers with swe-gym, 2024
Jiayi Pan, Xingyao Wang, Graham Neubig, Navdeep Jaitly, Heng Ji, Alane Suhr, and Yizhe Zhang · 2024
Cited alongside, same era.
Accessibility tree - mdn web docs glossary: Definitions of web-related terms — mdn
Mozilla
Cited in the paper.
Scribeagent: Towards specialized web agents using production-scale workflow data, 2024a
Junhong Shen, Atishay Jain, Zedian Xiao, Ishan Amlekar, Mouad Hadji, Aaron Podolny, and Ameet Talwalkar
Cited in the paper.
OpenHands: An Open Platform for AI Software Developers as Generalist Agents, 2024a
Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, Robert Brennan, Hao Peng, Heng Ji, and Graham Neubig
Cited in the paper.
Fangzhou Wu, Ethan Cecchetti, and Chaowei Xiao · 2024
Later among the works it cites.
τ \tau -bench: A benchmark for tool-agent-user interaction in real-world domains, 2024
Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan · 2024
Later among the works it cites.
Agent-as-a-judge: Evaluate agents with agents, 2024
Mingchen Zhuge, Changsheng Zhao, Dylan Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, and Jürgen Schmidhuber · 2024
Later among the works it cites.
Commercial llm agents are already vulnerable to simple yet dangerous attacks
Ang Li, Yin Zhou, Vethavikashini Chithrra Raghuram, Tom Goldstein, and Micah Goldblum · 2025
Closest in time.
Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments, 2025
Hongjin Su, Ruoxi Sun, Jinsung Yoon, Pengcheng Yin, Tao Yu, and Sercan Ö. Arık · 2025
Closest in time.