Fetching the paper…
Reading the bibliography…
Recent advancements in Large Language Models (LLMs) and multimodal counterparts have spurred significant interest in developing web agents -- AI systems capable of autonomously navigating and completing tasks within web environments.
Latent retrieval for weakly supervised open domain question answering, 2019
K. Lee, M.-W. Chang, and K. Toutanova · 1906
Earlier work this paper cites.
Realm: Retrieval-augmented language model pre-training, 2020
K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang · 2002
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering, 2020
V. Karpukhin, B. Oğuz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W. tau Yih · 2004
Earlier work this paper cites.
Colbert: Efficient and effective passage search via contextualized late interaction over bert, 2020
O. Khattab and M. Zaharia · 2004
Earlier work this paper cites.
Internet Security Glossary, Version 2
R. W. Shirey · 2007
Earlier work this paper cites.
Towards an iterative reinforcement approach for simultaneous document summarization and keyword extraction
X. Wan, J. Yang, and J. Xiao · 2007
Earlier work this paper cites.
Evolutionary timeline summarization: a balanced optimization framework via iterative substitution
R. Yan, X. Wan, J. Otterbacher, L. Kong, X. Li, and Y. Zhang · 2011
Earlier work this paper cites.
Json-rpc 2.0 specification, Jan. 2013
W. G. JSON-RPC · 2013
Earlier work this paper cites.
World of bits: An open-domain platform for web-based agents
T. Shi, A. Karpathy, L. J. Fan, J. Z. Hernández, and P. Liang · 2017
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
E. Z. Liu, K. Guu, P. Pasupat, T. Shi, and P. Liang · 2018
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback, 2022
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, X. Jiang, K. Cobbe, T. Eloundou, G. Krueger, K. Button, M. Knight, B. Chess, and J. Schulman · 2022
Earlier work this paper cites.
Introducing claude, Mar. 2023
Anthropic · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su · 2023
Earlier work this paper cites.
Multimodal web navigation with instruction-finetuned foundation models
H. Furuta, K.-H. Lee, O. Nachum, Y. Matsuo, A. Faust, S. S. Gu, and I. Gur · 2023
Earlier work this paper cites.
A real-world webagent with planning, long context understanding, and program synthesis
I. Gur, H. Furuta, A. Huang, M. Safdari, Y. Matsuo, D. Eck, and A. Faust · 2023
Earlier work this paper cites.
Agentbench: Evaluating llms as agents, 2023
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, S. Zhang, X. Deng, A. Zeng, Z. Du, C. Zhang, S. Shen, T. Zhang, Y. Su, H. Sun, M. Huang, Y. Dong, and J. Tang · 2023
Earlier work this paper cites.
Tool learning with foundation models, 2023
Y. Qin, S. Hu, Y. Lin, W. Chen, N. Ding, G. Cui, Z. Zeng, Y. Huang, C. Xiao, C. Han, Y. R. Fung, Y. Su, H. Wang, C. Qian, R. Tian, K. Zhu, S. Liang, X. Shen, B. Xu, Z. Zhang, Y. Ye, B. Li, Z. Tang, J. Yi, Y. Zhu, Z. Dai, L. Yan, X. Cong, Y. Lu, W. Zhao, Y. Huang, J. Yan, X. Han, X. Sun, D. Li, J. Phang, C. Yang, T. Wu, H. Ji, Z. Liu, and M. Sun · 2023
Earlier work this paper cites.
Androidinthewild: A large-scale dataset for android device control
C. Rawles, A. Li, D. Rodriguez, O. Riva, and T. Lillicrap · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools, 2023
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Earlier work this paper cites.
From pixels to ui actions: Learning to follow instructions via graphical user interfaces
P. Shaw, M. Joshi, J. Cohan, J. Berant, P. Pasupat, H. Hu, U. Khandelwal, K. Lee, and K. N. Toutanova · 2023
Earlier work this paper cites.
Reflexion: Language agents with verbal reinforcement learning, 2023
N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao · 2023
Earlier work this paper cites.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao · 2023
Cited alongside, same era.
Summit: Iterative text summarization via chatgpt
H. Zhang, X. Liu, and J. Zhang · 2023
Cited alongside, same era.
Introducing the model context protocol, Nov. 2024
Anthropic · 2024
Cited alongside, same era.
Llm2vec: Large language models are secretly powerful text encoders
P. BehnamGhader, V. Adlakha, M. Mosbach, D. Bahdanau, N. Chapados, and S. Reddy · 2024
Cited alongside, same era.
Fine-tuning web agents: It works, but it’s trickier than you think
M. Caccia, M. Thakkar, L. Boisvert, T. L. S. de Chezelles, A. Piché, N. Chapados, A. Drouin, M. Gasse, and A. Lacoste · 2024
Cited alongside, same era.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, et al · 2024
Later among the works it cites.
Swe-agent: Agent-computer interfaces enable automated software engineering
J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press · 2024
Later among the works it cites.
Gpt-4v (ision) is a generalist web agent, if grounded
B. Zheng, B. Gou, J. Kil, H. Sun, and Y. Su · 2024
Later among the works it cites.
WebArena: A Realistic Web Environment for Building Autonomous Agents, Apr. 2024
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, U. Alon, and G. Neubig · 2024
Later among the works it cites.
Agentharm: A benchmark for measuring harmfulness of llm agents, 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Workarena: How capable are web agents at solving common knowledge work tasks?
A. Drouin, M. Gasse, M. Caccia, I. H. Laradji, M. Del Verme, T. Marty, L. Boisvert, M. Thakkar, Q. Cappart, D. Vazquez, et al · 2024
Cited alongside, same era.
How culture shapes what people want from ai
X. Ge, C. Xu, D. Misaki, H. R. Markus, and J. L. Tsai · 2024
Cited alongside, same era.
Webvoyager: Building an end-to-end web agent with large multimodal models
H. He, W. Yao, K. Ma, W. Yu, Y. Dai, H. Zhang, Z. Lan, and D. Yu · 2024
Cited alongside, same era.
Videowebarena: Evaluating long context multimodal agents with video understanding web tasks
L. Jang, Y. Li, D. Zhao, C. Ding, J. Lin, P. P. Liang, R. Bonatti, and K. Koishida · 2024
Cited alongside, same era.
Autowebglm: A large language model-based web navigating agent, 2024
H. Lai, X. Liu, I. L. Iong, S. Yao, Y. Chen, P. Shen, H. Yu, H. Zhang, X. Zhang, Y. Dong, and J. Tang · 2024
Cited alongside, same era.
St-webagentbench: A benchmark for evaluating safety and trustworthiness in web agents, 2024
I. Levy, B. Wiesel, S. Marreed, A. Oved, A. Yaeli, and S. Shlomov · 2024
Cited alongside, same era.
On the effects of data scale on ui control agents
W. Li, W. Bishop, A. Li, C. Rawles, F. Campbell-Ajala, D. Tyamagundlu, and O. Riva · 2024
Cited alongside, same era.
M. Andriushchenko, A. Souly, M. Dziemian, D. Duenas, M. Lin, J. Wang, D. Hendrycks, A. Zou, Z. Kolter, M. Fredrikson, E. Winsor, J. Wynne, Y. Gal, and X. Davies · 2025
Closest in time.
Large language models empowered personalized web agents
H. Cai, Y. Li, W. Wang, F. Zhu, X. Shen, W. Li, and T.-S. Chua · 2025
Closest in time.
The browsergym ecosystem for web agent research, 2025
T. L. S. D. Chezelles, M. Gasse, A. Drouin, M. Caccia, L. Boisvert, M. Thakkar, T. Marty, R. Assouel, S. O. Shayegan, L. K. Jang, X. H. Lù, O. Yoran, D. Kong, F. F. Xu, S. Reddy, Q. Cappart, G. Neubig, R. Salakhutdinov, N. Chapados, and A. Lacoste · 2025
Closest in time.
Function calling with the gemini api, May 2025
Google · 2025
Closest in time.
Navigating the digital world as humans do: Universal visual grounding for gui agents, 2025
B. Gou, R. Wang, B. Zheng, Y. Xie, C. Chang, Y. Shu, H. Sun, and Y. Su · 2025
Closest in time.
Is your llm secretly a world model of the internet? model-based planning for web agents, 2025
Y. Gu, K. Zhang, Y. Ning, B. Zheng, B. Gou, T. Xue, C. Chang, S. Srivastava, Y. Xie, P. Qi, H. Sun, and Y. Su · 2025
Closest in time.
Redteamcua: Realistic adversarial testing of computer-use agents in hybrid web-os environments, 2025
Z. Liao, J. Jones, L. Jiang, E. Fosler-Lussier, Y. Su, Z. Lin, and H. Sun · 2025
Closest in time.
Agentrewardbench: Evaluating automatic evaluations of web agent trajectories, 2025
X. H. Lù, A. Kazemnejad, N. Meade, A. Patel, D. Shin, A. Zambrano, K. Stańczak, P. Shaw, C. J. Pal, and S. Reddy · 2025
Closest in time.
Nnetnav: Unsupervised learning of browser agents through environment interaction in the wild, 2025
S. Murty, H. Zhu, D. Bahdanau, and C. D. Manning · 2025
Closest in time.
Introducing deep research, Feb. 2025
OpenAI · 2025
Closest in time.
Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning, 2025
Z. Qi, X. Liu, I. L. Iong, H. Lai, X. Sun, W. Zhao, Y. Yang, X. Yang, J. Sun, S. Yao, T. Zhang, W. Xu, J. Tang, and Y. Dong · 2025
Closest in time.
Ui-tars: Pioneering automated gui interaction with native agents
Y. Qin, Y. Ye, J. Fang, H. Wang, S. Liang, S. Tian, J. Zhang, J. Li, Y. Li, S. Huang, et al · 2025
Closest in time.
Beyond browsing: Api-based web agents, 2025
Y. Song, F. Xu, S. Zhou, and G. Neubig · 2025
Closest in time.
Towards internet-scale training for agents
B. Trabucco, G. A. Sigurdsson, R. Piramuthu, and R. Salakhutdinov · 2025
Closest in time.
Safearena: Evaluating the safety of autonomous web agents, 2025
A. D. Tur, N. Meade, X. H. Lù, A. Zambrano, A. Patel, E. Durmus, S. Gella, K. Stańczak, and S. Reddy · 2025
Closest in time.
An illusion of progress? assessing the current state of web agents, 2025
T. Xue, W. Qi, T. Shi, C. H. Song, B. Gou, D. Song, H. Sun, and Y. Su · 2025
Closest in time.
Skillweaver: Web agents can self-improve by discovering and honing skills
B. Zheng, M. Y. Fatemi, X. Jin, Z. Z. Wang, A. Gandhi, Y. Song, Y. Gu, J. Srinivasa, G. Liu, G. Neubig, and Y. Su · 2025
Closest in time.