Fetching the paper…
Reading the bibliography…
Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows.
The graphical user interface
Jansen, B. J · 1998
Earlier work this paper cites.
Mapping natural language instructions to mobile ui action sequences
Li, Y., He, J., Zhou, X., Zhang, Y., and Baldridge, J · 2005
Earlier work this paper cites.
Widget captioning: Generating natural language description for mobile user interface elements
Li, Y., Li, G., He, L., Zheng, J., Li, H., and Guan, Z · 2010
Earlier work this paper cites.
From one tree to a forest: a unified solution for structured web data extraction
Hao, Q., Cai, R., Pang, Y., and Zhang, L · 2011
Earlier work this paper cites.
Rico: A mobile app dataset for building data-driven design applications
Deka, B., Huang, Z., Franzen, C., Hibschman, J., Afergan, D., Li, Y., Nichols, J., and Kumar, R · 2017
Earlier work this paper cites.
World of bits: An open-domain platform for web-based agents
Shi, T., Karpathy, A., Fan, L., Hernandez, J., and Liang, P · 2017
Earlier work this paper cites.
Introducing chatgpt
OpenAI · 2021
Earlier work this paper cites.
Androidenv: A reinforcement learning platform for android
Toyama, D., Hamel, P., Gergely, A., Comanici, G., Glaese, A., Ahmed, Z., Jackson, T., Mourad, S., and Precup, D · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Assistgpt: A general multi-modal assistant that can plan, execute, inspect, and learn
Gao, D., Ji, L., Zhou, L., Lin, K. Q., Chen, J., Fan, Z., and Shou, M. Z · 2023
Earlier work this paper cites.
Cogagent: A visual language model for gui agents
Hong, W., Wang, W., Lv, Q., Xu, J., Yu, W., Ji, J., Wang, Y., Wang, Z., Dong, Y., Ding, M., et al · 2023
Earlier work this paper cites.
Pix2struct: Screenshot parsing as pretraining for visual language understanding
Lee, K., Joshi, M., Turc, I. R., Hu, H., Liu, F., Eisenschlos, J. M., Khandelwal, U., Shaw, P., Chang, M.-W., and Toutanova, K · 2023
Earlier work this paper cites.
Sheetcopilot: Bringing software productivity to the next level through large language models
Li, H., Su, J., Chen, Y., Li, Q., and ZHANG, Z.-X · 2023
Earlier work this paper cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Earlier work this paper cites.
Android in the wild: A large-scale dataset for android device control
Rawles, C., Li, A., Rodriguez, D., Riva, O., and Lillicrap, T · 2023
Earlier work this paper cites.
Starvector: Generating scalable vector graphics code from images
Rodriguez, J. A., Agarwal, S., Laradji, I. H., Rodriguez, P., Vazquez, D., Pal, C., and Pedersoli, M · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Cited alongside, same era.
Webui: A dataset for enhancing visual ui understanding with web semantics
Wu, J., Wang, S., Shen, S., Peng, Y.-H., Nichols, J., and Bigham, J. P · 2023
Cited alongside, same era.
Gpt-4v in wonderland: Large multimodal models for zero-shot smartphone gui navigation
Yan, A., Yang, Z., Zhu, W., Lin, K., Li, L., Wang, J., Yang, J., Zhong, Y., McAuley, J., Gao, J., et al · 2023
Cited alongside, same era.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023
Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Weblinx: Real-world website navigation with multi-turn dialogue
Lù, X. H., Kasner, Z., and Reddy, S · 2024
Later among the works it cites.
Omniparser for pure vision based gui agent, 2024
Lu, Y., Yang, J., Shen, Y., and Awadallah, A · 2024
Later among the works it cites.
Introducing meta llama 3: The most capable openly available llm to date, 2024
Meta · 2024
Later among the works it cites.
Hello gpt-4o, May 2024
OpenAI · 2024
Later among the works it cites.
Rodriguez, J., Jian, X., Panigrahi, S. S., Zhang, T., Feizi, A., Puri, A., Kalkunte, A., Savard, F., Masry, A., Nayak, S., et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2023
Cited alongside, same era.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Bisk, Y., Fried, D., Alon, U., et al · 2023
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku
Anthropic, A · 2024
Cited alongside, same era.
Windows agent arena: Evaluating multi-modal os agents at scale
Bonatti, R., Zhao, D., Bonacci, F., Dupont, D., Abdali, S., Li, Y., Lu, Y., Wagle, J., Koishida, K., Bucker, A., et al · 2024
Cited alongside, same era.
Chen, Z., Wang, W., Tian, H., Ye, S., Gao, Z., Cui, E., Tong, W., Hu, K., Luo, J., Ma, Z., Ma, J., Wang, J., Dong, X., Yan, H., Guo, H., He, C., Shi, B., Jin, Z., Xu, C., Wang, B., Wei, X., Li, W., Zhang, W., Zhang, B., Cai, P., Wen, L., Yan, X., Dou, M., Lu, L., Zhu, X., Lu, T., Lin, D., Qiao, Y., Dai, J., and Wang, W · 2024
Cited alongside, same era.
Seeclick: Harnessing gui grounding for advanced visual gui agents
Cheng, K., Sun, Q., Chu, Y., Xu, F., Li, Y., Zhang, J., and Wu, Z · 2024
Cited alongside, same era.
Visual grounding for desktop graphical user interfaces
Dardouri, T., Minkova, L., Espejel, J. L., Dahhane, W., and Ettifouri, E. H · 2024
Cited alongside, same era.
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2024
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., Mariooryad, S., Ding, Y., Geng, X., Alcober, F., Frostig, R., Omernick, M., Walker, L., Paduraru, C., Sorokin, C., Tacchetti, A., Gaffney, C., Daruki, S., Sercinoglu, O., Gleicher, Z., and Others · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., Fan, Y., Dang, K., Du, M., Ren, X., Men, R., Liu, D., Zhou, C., Zhou, J., and Lin, J · 2024
Later among the works it cites.
V?: Guided visual search as a core mechanism in multimodal llms
Wu, P. and Xie, S · 2024
Later among the works it cites.
Os-atlas: A foundation action model for generalist gui agents
Wu, Z., Wu, Z., Xu, F., Wang, Y., Sun, Q., Jia, C., Cheng, K., Ding, Z., Chen, L., Liang, P. P., et al · 2024
Later among the works it cites.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., et al · 2024
Later among the works it cites.
Aguvis: Unified pure vision agents for autonomous gui interaction, 2024
Xu, Y., Wang, Z., Wang, J., Lu, D., Xie, T., Saha, A., Sahoo, D., Yu, T., and Xiong, C · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone
Yao, Y., Yu, T., Zhang, A., Wang, C., Cui, J., Zhu, H., Cai, T., Li, H., Zhao, W., He, Z., et al · 2024
Later among the works it cites.
Ferret-ui: Grounded mobile ui understanding with multimodal llms
You, K., Zhang, H., Schoop, E., Weers, F., Swearngin, A., Nichols, J., Yang, Y., and Gan, Z · 2024
Later among the works it cites.
Android in the zoo: Chain-of-action-thought for gui agents
Zhang, J., Wu, J., Teng, Y., Liao, M., Xu, N., Xiao, X., Wei, Z., and Tang, D · 2024
Later among the works it cites.
Gpt-4v(ision) is a generalist web agent, if grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y · 2024
Later among the works it cites.
Screenspot-pro: Gui grounding for professional high-resolution computer use, 2025
Li, K., Meng, Z., Lin, H., Luo, Z., Tian, Y., Ma, J., Huang, Z., and Chua, T.-S · 2025
Closest in time.
Ui-tars: Pioneering automated gui interaction with native agents
Qin, Y., Ye, Y., Fang, J., Wang, H., Liang, S., Tian, S., Zhang, J., Li, J., Li, Y., Huang, S., Zhong, W., Li, K., Yang, J., Miao, Y., Lin, W., Liu, L., Jiang, X., Ma, Q., Li, J., Xiao, X., Cai, K., Li, C., Zheng, Y., Jin, C., Li, C., Zhou, X., Wang, M., Chen, H., Li, Z., Yang, H., Liu, H., Lin, F., Peng, T., Liu, X., and Shi, G · 2025
Closest in time.