Fetching the paper…
Reading the bibliography…
LLM-based agents are becoming increasingly proficient at solving web-based tasks.
A coefficient of agreement for nominal scales
Cohen, J · 1960
Earlier work this paper cites.
Language models are few-shot learners, July 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
World of bits: An open-domain platform for web-based agents
Shi, T., Karpathy, A., Fan, L., Hernandez, J., and Liang, P · 2017
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
Liu, E. Z., Guu, K., Pasupat, P., and Liang, P · 2018
Earlier work this paper cites.
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., Jones, A., Bowman, S., Chen, A., Conerly, T., DasSarma, N., Drain, D., Elhage, N., El-Showk, S., Fort, S., Hatfield-Dodds, Z., Henighan, T., Hernandez, D., Hume, T., Jacobson, J., Johnston, S., Kravec, S., Olsson, C., Ringer, S., Tran-Johnson, E., Amodei, D., Brown, T., Joseph, N., McCandlish, S., Olah, C., Kaplan, J., and Clark, J · 2022
Earlier work this paper cites.
Mind2Web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2023
Earlier work this paper cites.
Safe, secure, and trustworthy development and use of artificial intelligence
Executive Office of the President · 2023
Earlier work this paper cites.
Llama Guard: LLM-based input-output safeguard for human-AI conversations, December 2023
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., and Khabsa, M · 2023
Earlier work this paper cites.
Language models can solve computer tasks
Kim, G., Baldi, P., and McAleer, S · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
From pixels to UI actions: Learning to follow instructions via graphical user interfaces
Shaw, P., Joshi, M., Cohan, J., Berant, J., Pasupat, P., Hu, H., Khandelwal, U., Lee, K., and Toutanova, K · 2023
Earlier work this paper cites.
LLaMA: Open and efficient foundation language models, February 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Earlier work this paper cites.
Set-of-mark prompting unleashes extraordinary visual grounding in GPT-4V, 2023
Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J · 2023
Earlier work this paper cites.
Universal and transferable adversarial attacks on aligned language models, December 2023
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Earlier work this paper cites.
AgentHarm: A benchmark for measuring harmfulness of LLM agents, October 2024
Andriushchenko, M., Souly, A., Dziemian, M., Duenas, D., Lin, M., Wang, J., Hendrycks, D., Zou, A., Kolter, Z., Fredrikson, M., Winsor, E., Wynne, J., Gal, Y., and Davies, X · 2024
Earlier work this paper cites.
Many-shot jailbreaking, April 2024
Anil, C., Durmus, E., Sharma, M., Benton, J., Kundu, S., Batson, J., Rimsky, N., Tong, M., Mu, J., and Ford, D · 2024
Earlier work this paper cites.
Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku, October 2024
Anthropic · 2024
Cited alongside, same era.
The BrowserGym ecosystem for web agent research, December 2024
Chezelles, T. L. S. D., Gasse, M., Drouin, A., Caccia, M., Boisvert, L., Thakkar, M., Marty, T., Assouel, R., Shayegan, S. O., Jang, L. K., Lù, X. H., Yoran, O., Kong, D., Xu, F. F., Reddy, S., Cappart, Q., Neubig, G., Salakhutdinov, R., Chapados, N., and Lacoste, A · 2024
Cited alongside, same era.
Debenedetti, E., Zhang, J., Balunović, M., Beurer-Kellner, L., Fischer, M., and Tramèr, F · 2024
Cited alongside, same era.
WorkArena: How capable are web agents at solving common knowledge work tasks?
Drouin, A., Gasse, M., Caccia, M., Laradji, I. H., Verme, M. D., Marty, T., Vazquez, D., Chapados, N., and Lacoste, A · 2024
Cited alongside, same era.
OpenAI Team et al · 2024
Later among the works it cites.
WebCanvas: Benchmarking web agents in online environments
Pan, Y., Kong, D., Zhou, S., Cui, C., Leng, Y., Jiang, B., Liu, H., Shang, Y., Zhou, S., Wu, T., and Wu, Z · 2024
Later among the works it cites.
Identifying the risks of LM agents with an LM-emulated sandbox, May 2024
Ruan, Y., Dong, H., Wang, A., Pitis, S., Zhou, Y., Ba, J., Dubois, Y., Maddison, C. J., and Hashimoto, T · 2024
Later among the works it cites.
Russinovich, M., Salem, A., and Eldan, R · 2024
Later among the works it cites.
Step: Stacked LLM policies for web actions
Sodhi, P., Branavan, S., Artzi, Y., and McDonald, R · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multimodal web navigation with instruction-finetuned foundation models
Furuta, H., Lee, K.-H., Nachum, O., Matsuo, Y., Faust, A., Gu, S. S., and Gur, I · 2024
Cited alongside, same era.
Gemma: Open models based on Gemini research and technology, March 2024
Gemma Team et al · 2024
Cited alongside, same era.
OLMo: Accelerating the science of language models, February 2024
Groeneveld, D., Beltagy, I., Walsh, P., Bhagia, A., Kinney, R., Tafjord, O., Jha, A. H., Ivison, H., Magnusson, I., Wang, Y., Arora, S., Atkinson, D., Authur, R., Chandu, K. R., Cohan, A., Dumas, J., Elazar, Y., Gu, Y., Hessel, J., Khot, T., Merrill, W., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Pyatkin, V., Ravichander, A., Schwenk, D., Shah, S., Smith, W., Strubell, E., Subramani, N., Wortsman, M., Dasigi, P., Lambert, N., Richardson, K., Zettlemoyer, L., Dodge, J., Lo, K., Soldaini, L., Smith, N. A., and Hajishirzi, H · 2024
Cited alongside, same era.
Han, S., Rao, K., Ettinger, A., Jiang, L., Lin, B. Y., Lambert, N., Choi, Y., and Dziri, N · 2024
Cited alongside, same era.
SWE-bench: Can language models resolve real-world Github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R · 2024
Cited alongside, same era.
VisualWebArena: Evaluating multimodal agents on realistic visual web tasks, June 2024
Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M. C., Huang, P.-Y., Neubig, G., Zhou, S., Salakhutdinov, R., and Fried, D · 2024
Cited alongside, same era.
ST-WebAgentBench: A benchmark for evaluating safety and trustworthiness in web agents, October 2024
Levy, I., Wiesel, B., Marreed, S., Oved, A., Yaeli, A., and Shlomov, S · 2024
Cited alongside, same era.
EIA: Environmental injection attack on generalist web agents for privacy leakage, October 2024
Liao, Z., Mo, L., Xu, C., Kang, M., Zhang, J., Xiao, C., Tian, Y., Li, B., and Sun, H · 2024
Cited alongside, same era.
Multi-turn context jailbreak attack on large language models from first principles, August 2024
Sun, X., Zhang, D., Yang, D., Zou, Q., and Li, H · 2024
Later among the works it cites.
WebLINX: Real-world website navigation with multi-turn dialogue
ù, X. H., Kasner, Z., and Reddy, S · 2024
Later among the works it cites.
Bypassing the safety training of open-source LLMs with priming attacks
Vega, J., Chaudhary, I., Xu, C., and Singh, G · 2024
Later among the works it cites.
Hidden in plain sight: Exploring chat history tampering in interactive language models, 2024
Wei, C., Zhao, Y., Gong, Y., Chen, K., Xiang, L., and Zhu, S · 2024
Later among the works it cites.
InjecAgent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Zhan, Q., Liang, Z., Ying, Z., and Kang, D · 2024
Later among the works it cites.
WebOlympus: An open platform for web agents on live websites
Zheng, B., Gou, B., Salisbury, S., Du, Z., Sun, H., and Su, Y · 2024
Later among the works it cites.
WebArena: A realistic web environment for building autonomous agents, April 2024
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., Alon, U., and Neubig, G · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers, July 2024
Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., Wang, R., Kolter, Z., Fredrikson, M., and Hendrycks, D · 2024
Later among the works it cites.
Computer-Using Agent: Introducing a universal interface for AI to interact with the digital world, 2025
OpenAI · 2025
Closest in time.
The dark side of function calling: Pathways to jailbreaking large language models
Wu, Z., Gao, H., He, J., and Wang, P · 2025
Closest in time.