Fetching the paper…
Reading the bibliography…
We study the use of large language model-based agents for interacting with software via web browsers.
OpenAI gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
Liu, E. Z., Guu, K., Pasupat, P., Shi, T., and Liang, P · 2018
Earlier work this paper cites.
Mapping natural language instructions to mobile ui action sequences
Li, Y., He, J., Zhou, X., Zhang, Y., and Baldridge, J · 2020
Earlier work this paper cites.
Knowledge 2020: “The digital workflow revolution has just begun”
Maas, M · 2020
Earlier work this paper cites.
WebGPT: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J · 2021
Earlier work this paper cites.
Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles
SAE · 2021
Earlier work this paper cites.
A data-driven approach for learning to control computers
Humphreys, P. C., Raposo, D., Pohlen, T., Thornton, G., Chhaparia, R., Muldal, A., Abramson, J., Georgiev, P., Santoro, A., and Lillicrap, T · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., ichter, b., Xia, F., Chi, E., Le, Q. V., and Zhou, D · 2022
Earlier work this paper cites.
WebShop: Towards scalable real-world web interaction with grounded language agents
Yao, S., Chen, H., Yang, J., and Narasimhan, K · 2022
Earlier work this paper cites.
The unsolved challenges of LLMs in open-ended web tasks: A case study
Assouel, R., Marty, T., Caccia, M., Laradji, I., Drouin, A., Rajeswar, S., Palacios, H., Cappart, Q., Vazquez, D., Chapados, N., Gasse, M., and Lacoste, A · 2023
Earlier work this paper cites.
Mind2Web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2023
Cited alongside, same era.
Multimodal web navigation with instruction-finetuned foundation models
Furuta, H., Nachum, O., Lee, K.-H., Matsuo, Y., Gu, S. S., and Gur, I · 2023
Cited alongside, same era.
Chrome devtools protocol, 2023
Google · 2023
Cited alongside, same era.
Language models can solve computer tasks
Kim, G., Baldi, P., and McAleer, S · 2023
Cited alongside, same era.
Playwright for Python documentation, 2023
Microsoft · 2023
Cited alongside, same era.
ReAct: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2023
Later among the works it cites.
Agenttuning: Enabling generalized agent abilities for llms
Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., and Tang, J · 2023
Later among the works it cites.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Bisk, Y., Fried, D., Alon, U., and Neubig, G · 2023
Later among the works it cites.
WebVoyager: Building an end-to-end web agent with large multimodal models
He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., and Yu, D · 2024
Closest in time.
Automatic macro mining from interaction traces at scale
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
OpenAI · 2023
Cited alongside, same era.
Androidinthewild: A large-scale dataset for android device control
Rawles, C., Li, A., Rodriguez, D., Riva, O., and Lillicrap, T · 2023
Cited alongside, same era.
Vancouver release notes
ServiceNow · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Canton-Ferrer, C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Cited alongside, same era.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J · 2023
Cited alongside, same era.
A real-world webagent with planning, long context understanding, and program synthesis
Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., and Faust, A
Cited in the paper.
A real-world WebAgent with planning, long context understanding, and program synthesis
Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., and Faust, A
Cited in the paper.
Huang, F., Li, G., Li, T., and Li, Y · 2024
Closest in time.
Weblinx: Real-world website navigation with multi-turn dialogue
Lù, X. H., Kasner, Z., and Reddy, S · 2024
Closest in time.
ServiceNow joins the prestigious Fortune 500 list
Mastantuono, G · 2024
Closest in time.
Llama 3: Meta’s latest large language model
Meta · 2024
Closest in time.
A journey into the future of the translation industry, 2021
van der Meer, J · 2024
Closest in time.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., Liu, Y., Xu, Y., Zhou, S., Savarese, S., Xiong, C., Zhong, V., and Yu, T · 2024
Closest in time.