Fetching the paper…
Reading the bibliography…
Recent advancements in Large Vision Language Models (LVLMs) have enabled the development of LVLM-based Graphical User Interface (GUI) agents under various paradigms.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Mapping natural language instructions to mobile ui action sequences
Li, Y., He, J., Zhou, X., Zhang, Y., and Baldridge, J · 2005
Earlier work this paper cites.
Screenspot: multidimensional resource discovery for distributed applications in smart spaces
Jurmu, M., Boring, S., and Riekki, J · 2008
Earlier work this paper cites.
World of bits: An open-domain platform for web-based agents
Shi, T., Karpathy, A., Fan, L., Hernandez, J., and Liang, P · 2017
Earlier work this paper cites.
Gur, I., Rückert, U., Faust, A., and Hakkani-Tür, D. Z · 2018
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
Liu, E. Z., Guu, K., Pasupat, P., Shi, T., and Liang, P · 2018
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., teusz Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Unblind your apps: Predicting natural-language labels for mobile gui components by deep learning
Chen, J., Chen, C., Xing, Z., Xu, X., Zhu, L., Li, G., and Wang, J · 2020
Earlier work this paper cites.
Actionbert: Leveraging user actions for semantic understanding of user interfaces
He, Z., Sunkara, S., Zang, X., Xu, Y., Liu, L., Wichers, N., Schubiner, G., Lee, R. B., and Chen, J · 2020
Earlier work this paper cites.
Uibert: Learning generic multimodal representations for ui understanding
Bai, C., Zang, X., Xu, Y., Sunkara, S., Rastogi, A., Chen, J., and y Arcas, B. A · 2021
Earlier work this paper cites.
Screen2words: Automatic mobile ui summarization with multimodal learning
Wang, B., Li, G., Zhou, X., Chen, Z., Grossman, T., and Li, Y · 2021
Earlier work this paper cites.
Screen parsing: Towards reverse engineering of ui models from screenshots
Wu, J., Zhang, X., Nichols, J., and Bigham, J. P · 2021
Earlier work this paper cites.
Screen recognition: Creating accessibility metadata for mobile applications from pixels
Zhang, X., de Greef, L., Swearngin, A., White, S., Murray, K. I., Yu, L., Shan, Q., Nichols, J., Wu, J., Fleizach, C., Everitt, A., and Bigham, J. P · 2021
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Yao, S., Chen, H., Yang, J., and Narasimhan, K · 2022
Cited alongside, same era.
Introducing claude 3.5, 2023
Anthropic · 2023
Cited alongside, same era.
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2023
Cited alongside, same era.
Multimodal web navigation with instruction-finetuned foundation models
Furuta, H., Nachum, O., Lee, K.-H., Matsuo, Y., Gu, S. S., and Gur, I · 2023
Cited alongside, same era.
A real-world webagent with planning, long context understanding, and program synthesis
Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., and Faust, A · 2023
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J · 2023
Later among the works it cites.
Appagent: Multimodal agents as smartphone users
Zhang, C. X., Yang, Z., Liu, J., Han, Y., Chen, X., Huang, Z., Fu, B., and Yu, G · 2023
Later among the works it cites.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Bisk, Y., Fried, D., Alon, U., et al · 2023
Later among the works it cites.
Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
Bai, H., Zhou, Y., Cemri, M., Pan, J., Suhr, A., Levine, S., and Kumar, A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cogagent: A visual language model for gui agents
Hong, W., Wang, W., Lv, Q., Xu, J., Yu, W., Ji, J., Wang, Y., Wang, Z., Zhang, Y., Li, J.-Z., Xu, B., Dong, Y., Ding, M., and Tang, J · 2023
Cited alongside, same era.
Ultralytics yolov8, 2023
Jocher, G., Chaurasia, A., and Qiu, J · 2023
Cited alongside, same era.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., and Girshick, R · 2023
Cited alongside, same era.
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Cited alongside, same era.
Android in the wild: A large-scale dataset for android device control, 2023
Rawles, C., Li, A., Rodriguez, D., Riva, O., and Lillicrap, T · 2023
Cited alongside, same era.
From pixels to ui actions: Learning to follow instructions via graphical user interfaces
Shaw, P., Joshi, M., Cohan, J., Berant, J., Pasupat, P., Hu, H., Khandelwal, U., Lee, K., and Toutanova, K · 2023
Cited alongside, same era.
Hierarchical prompting assists large language model on web navigation
Sridhar, A., Lo, R., Xu, F. F., Zhu, H., and Zhou, S · 2023
Cited alongside, same era.
Cheng, K., Sun, Q., Chu, Y., Xu, F., Li, Y., Zhang, J., and Wu, Z · 2024
Later among the works it cites.
Read anywhere pointed: Layout-aware gui screen reading with tree-of-lens grounding, 2024
Fan, Y., Ding, L., Kuo, C.-C., Jiang, S., Zhao, Y., Guan, X., Yang, J., Zhang, Y., and Wang, X. E · 2024
Later among the works it cites.
Webvoyager: Building an end-to-end web agent with large multimodal models
He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., and Yu, D · 2024
Later among the works it cites.
Tree search for language model agents, 2024
Koh, J. Y., McAleer, S., Fried, D., and Salakhutdinov, R · 2024
Later among the works it cites.
Visualwebbench: How far have multimodal llms evolved in web page understanding and grounding?
Liu, J., Song, Y., Lin, B. Y., Lam, W., Neubig, G., Li, Y., and Yue, X · 2024
Later among the works it cites.
Lu, Y., Yang, J., Shen, Y., and Awadallah, A · 2024
Later among the works it cites.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., Liu, Y., Xu, Y., Zhou, S., Savarese, S., Xiong, C., Zhong, V., and Yu, T · 2024
Later among the works it cites.
Ferret-ui: Grounded mobile ui understanding with multimodal llms
You, K., Zhang, H., Schoop, E., Weers, F., Swearngin, A., Nichols, J., Yang, Y., and Gan, Z · 2024
Later among the works it cites.
Gpt-4v(ision) is a generalist web agent, if grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y · 2024
Later among the works it cites.