Fetching the paper…
Reading the bibliography…
Automating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 1910
Earlier work this paper cites.
Rico: A mobile app dataset for building data-driven design applications
Deka, B., Huang, Z., Franzen, C., Hibschman, J., Afergan, D., Li, Y., Nichols, J., and Kumar, R · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
Mapping natural language instructions to mobile UI action sequences
Li, Y., He, J., Zhou, X., Zhang, Y., and Baldridge, J · 2020
Earlier work this paper cites.
Widget captioning: Generating natural language description for mobile user interface elements
Li, Y., Li, G., He, L., Zheng, J., Li, H., and Guan, Z · 2020
Earlier work this paper cites.
Zero: memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Uibert: Learning generic multimodal representations for UI understanding
Bai, C., Zang, X., Xu, Y., Sunkara, S., Rastogi, A., Chen, J., and y Arcas, B. A · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Inner monologue: Embodied reasoning through planning with language models
Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., et al · 2022
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2023
Earlier work this paper cites.
Language models can solve computer tasks
Kim, G., Baldi, P., and McAleer, S · 2023
Earlier work this paper cites.
A zero-shot language agent for computer control with structured reflection, 2023
Li, T., Li, G., Deng, Z., Wang, B., and Li, Y · 2023
Earlier work this paper cites.
Webui: A dataset for enhancing visual ui understanding with web semantics
Wu, J., Wang, S., Shen, S., Peng, Y.-H., Nichols, J., and Bigham, J · 2023
Earlier work this paper cites.
Appagent: Multimodal agents as smartphone users
Zhang, C. X., Yang, Z., Liu, J., Han, Y., Chen, X., Huang, Z., Fu, B., and Yu, G · 2023
Earlier work this paper cites.
Windows agent arena: Evaluating multi-modal os agents at scale
Bonatti, R., Zhao, D., Bonacci, F., Dupont, D., Abdali, S., Li, Y., Lu, Y., Wagle, J., Koishida, K., Bucker, A. F. C., Jang, L., and Hui, Z · 2024
Cited alongside, same era.
Spider2-v: How far are multimodal agents from automating data science and engineering workflows?
Cao, R., Lei, F., Wu, H., Chen, J., Fu, Y., Gao, H., Xiong, X., Zhang, H., Mao, Y., Hu, W., Xie, T., Xu, H., Zhang, D., Wang, S., Sun, R., Yin, P., Xiong, C., Ni, A., Liu, Q., Zhong, V., Chen, L., Yu, K., and Yu, T · 2024
Cited alongside, same era.
Amex: Android multi-annotation expo dataset for mobile gui agents
Chai, Y., Huang, S., Niu, Y., Xiao, H., Liu, L., Zhang, D., Gao, P., Ren, S., and Li, H · 2024
Cited alongside, same era.
Seeclick: Harnessing GUI grounding for advanced visual GUI agents
Cheng, K., Sun, Q., Chu, Y., Xu, F., Li, Y., Zhang, J., and Wu, Z · 2024
Cited alongside, same era.
Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models
Omniparser for pure vision based gui agent, 2024
Lu, Y., Yang, J., Shen, Y., and Awadallah, A · 2024
Closest in time.
Playwright for python documentation
Microsoft · 2024
Closest in time.
Screenagent: A vision language model-driven computer control agent, 2024
Niu, R., Li, J., Wang, S., Fu, Y., Hu, X., Leng, X., Kong, H., Chang, Y., and Wang, Q · 2024
Closest in time.
Hello gpt-4o, 2024
OpenAI · 2024
Closest in time.
Webcanvas: Benchmarking web agents in online environments
Pan, Y., Kong, D., Zhou, S., Cui, C., Leng, Y., Jiang, B., Liu, H., Shang, Y., Zhou, S., Wu, T., et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deitke, M., Clark, C., Lee, S., Tripathi, R., Yang, Y., Park, J. S., Salehi, M., Muennighoff, N., Lo, K., Soldaini, L., Lu, J., Anderson, T., Bransom, E., Ehsani, K., Ngo, H., Chen, Y., Patel, A., Yatskar, M., Callison-Burch, C., Head, A., Hendrix, R., Bastani, F., VanderBilt, E., Lambert, N., Chou, Y., Chheda, A., Sparks, J., Skjonsberg, S., Schmitz, M., Sarnat, A., Bischoff, B., Walsh, P., Newell, C., Wolters, P., Gupta, T., Zeng, K., Borchardt, J., Groeneveld, D., Dumas, J., Nam, C., Lebrecht, S., Wittlif, C., Schoenick, C., Michel, O., Krishna, R., Weihs, L., Smith, N. A., Hajishirzi, H., Girshick, R. B., Farhadi, A., and Kembhavi, A · 2024
Cited alongside, same era.
Workarena: How capable are web agents at solving common knowledge work tasks?, 2024
Drouin, A., Gasse, M., Caccia, M., Laradji, I. H., Verme, M. D., Marty, T., Boisvert, L., Thakkar, M., Cappart, Q., Vazquez, D., Chapados, N., and Lacoste, A · 2024
Cited alongside, same era.
Navigating the digital world as humans do: Universal visual grounding for GUI agents
Gou, B., Wang, R., Zheng, B., Xie, Y., Chang, C., Shu, Y., Sun, H., and Su, Y · 2024
Cited alongside, same era.
A real-world webagent with planning, long context understanding, and program synthesis
Gur, I., Furuta, H., Huang, A. V., Safdari, M., Matsuo, Y., Eck, D., and Faust, A · 2024
Cited alongside, same era.
Cogagent: A visual language model for gui agents
Hong, W., Wang, W., Lv, Q., Xu, J., Yu, W., Ji, J., Wang, Y., Wang, Z., Dong, Y., Ding, M., et al · 2024
Cited alongside, same era.
Kapoor, R., Butala, Y. P., Russak, M., Koh, J. Y., Kamble, K., Alshikh, W., and Salakhutdinov, R · 2024
Cited alongside, same era.
Autowebglm: Bootstrap and reinforce a large language model-based web navigating agent
Lai, H., Liu, X., Iong, I. L., Yao, S., Chen, Y., Shen, P., Yu, H., Zhang, H., Zhang, X., Dong, Y., et al · 2024
Cited alongside, same era.
MUG: interactive multimodal grounding on user interfaces
Li, T., Li, G., Zheng, J., Wang, P., and Li, Y · 2024
Cited alongside, same era.
Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al · 2024
Closest in time.
OS-ATLAS: A foundation action model for generalist GUI agents
Wu, Z., Wu, Z., Xu, F., Wang, Y., Sun, Q., Jia, C., Cheng, K., Ding, Z., Chen, L., Liang, P. P., and Qiao, Y · 2024
Closest in time.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., et al · 2024
Closest in time.
Agent lumos: Unified and modular training for open-source language agents
Yin, D., Brahman, F., Ravichander, A., Chandu, K. R., Chang, K., Choi, Y., and Lin, B. Y · 2024
Closest in time.
You only look at screens: Multimodal chain-of-action agents
Zhang, Z. and Zhang, A · 2024
Closest in time.
Gpt-4v(ision) is a generalist web agent, if grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y · 2024
Closest in time.
Synapse: Trajectory-as-exemplar prompting with memory for computer control
Zheng, L., Wang, R., Wang, X., and An, B · 2024
Closest in time.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., Alon, U., and Neubig, G · 2024
Closest in time.
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., and Lin, J · 2025
Closest in time.