Fetching the paper…
Reading the bibliography…
The rapid advancement of large Vision-Language Models (VLMs) has propelled the development of pure-vision-based GUI Agents, capable of perceiving and operating Graphical User Interfaces (GUI) to autonomously fulfill user instructions.
A markovian decision process
R. Bellman · 1957
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
World of bits: An open-domain platform for web-based agents
T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang · 2017
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
E. Z. Liu, K. Guu, P. Pasupat, T. Shi, and P. Liang · 2018
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection for dimension reduction
L. McInnes, J. Healy, and J. Melville · 2018
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 2019
Earlier work this paper cites.
Mapping natural language instructions to mobile ui action sequences
Y. Li, J. He, X. Zhou, Y. Zhang, and J. Baldridge · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Screenqa: Large-scale question-answer pairs over mobile app screenshots
Y.-C. Hsiao, F. Zubach, G. Baechler, V. Carbune, J. Lin, M. Wang, S. Sunkara, Y. Zhu, and J. Chen · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Meta-gui: Towards multi-modal conversational agents on mobile gui
L. Sun, X. Chen, L. Chen, T. Dai, Z. Zhu, and K. Yu · 2022
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
S. Yao, H. Chen, J. Yang, and K. Narasimhan · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2023
Earlier work this paper cites.
Androidinthewild: A large-scale dataset for android device control
C. Rawles, A. Li, D. Rodriguez, O. Riva, and T. Lillicrap · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2023
Earlier work this paper cites.
Webarena: A realistic web environment for building autonomous agents
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, et al · 2023
Earlier work this paper cites.
Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
H. Bai, Y. Zhou, J. Pan, M. Cemri, A. Suhr, S. Levine, and A. Kumar · 2024
Earlier work this paper cites.
Windows agent arena: Evaluating multi-modal os agents at scale
R. Bonatti, D. Zhao, F. Bonacci, D. Dupont, S. Abdali, Y. Li, Y. Lu, J. Wagle, K. Koishida, A. Bucker, et al · 2024
Earlier work this paper cites.
Amex: Android multi-annotation expo dataset for mobile gui agents
Y. Chai, S. Huang, Y. Niu, H. Xiao, L. Liu, D. Zhang, P. Gao, S. Ren, and H. Li · 2024
Earlier work this paper cites.
Guicourse: From general vision language models to versatile gui agents
W. Chen, J. Cui, J. Hu, Y. Qin, J. Fang, Y. Zhao, C. Wang, J. Liu, G. Chen, Y. Huo, et al · 2024
Earlier work this paper cites.
Seeclick: Harnessing gui grounding for advanced visual gui agents
K. Cheng, Q. Sun, Y. Chu, F. Xu, Y. Li, J. Zhang, and Z. Wu · 2024
Cited alongside, same era.
Workarena: How capable are web agents at solving common knowledge work tasks?
A. Drouin, M. Gasse, M. Caccia, I. H. Laradji, M. Del Verme, T. Marty, L. Boisvert, M. Thakkar, Q. Cappart, D. Vazquez, et al · 2024
Cited alongside, same era.
Assistgui: Task-oriented pc graphical user interface automation
D. Gao, L. Ji, Z. Bai, M. Ouyang, P. Li, D. Mao, Q. Wu, W. Zhang, P. Wang, X. Guo, et al · 2024
Cited alongside, same era.
Navigating the digital world as humans do: Universal visual grounding for gui agents
B. Gou, R. Wang, B. Zheng, Y. Xie, C. Chang, Y. Shu, H. Sun, and Y. Su · 2024
Cited alongside, same era.
Webvoyager: Building an end-to-end web agent with large multimodal models
Android in the zoo: Chain-of-action-thought for gui agents
J. Zhang, J. Wu, Y. Teng, M. Liao, N. Xu, X. Xiao, Z. Wei, and D. Tang · 2024
Later among the works it cites.
Gpt-4v (ision) is a generalist web agent, if grounded
B. Zheng, B. Gou, J. Kil, H. Sun, and Y. Su · 2024
Later among the works it cites.
Introducing claude 3.5 sonnet
Anthropic · 2025
Closest in time.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin · 2025
Closest in time.
R1-v: Reinforcing super generalization ability in vision-language models with less than $3
L. Chen, L. Li, H. Zhao, Y. Song, and Vinci · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. He, W. Yao, K. Ma, W. Yu, Y. Dai, H. Zhang, Z. Lan, and D. Yu · 2024
Cited alongside, same era.
Cogagent: A visual language model for gui agents
W. Hong, W. Wang, Q. Lv, J. Xu, W. Yu, J. Ji, Y. Wang, Z. Wang, Y. Dong, M. Ding, et al · 2024
Cited alongside, same era.
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al · 2024
Cited alongside, same era.
Omniact: A dataset and benchmark for enabling multimodal generalist autonomous agents for desktop and web
R. Kapoor, Y. P. Butala, M. Russak, J. Y. Koh, K. Kamble, W. AlShikh, and R. Salakhutdinov · 2024
Cited alongside, same era.
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks
J. Y. Koh, R. Lo, L. Jang, V. Duvvur, M. C. Lim, P.-Y. Huang, G. Neubig, S. Zhou, R. Salakhutdinov, and D. Fried · 2024
Cited alongside, same era.
On the effects of data scale on ui control agents
W. Li, W. E. Bishop, A. Li, C. Rawles, F. Campbell-Ajala, D. Tyamagundlu, and O. Riva · 2024
Cited alongside, same era.
Showui: One vision-language-action model for gui visual agent
K. Q. Lin, L. Li, D. Gao, Z. Yang, S. Wu, Z. Bai, W. Lei, L. Wang, and M. Z. Shou · 2024
Cited alongside, same era.
Gui odyssey: A comprehensive dataset for cross-app gui navigation on mobile devices
Q. Lu, W. Shao, Z. Liu, F. Meng, B. Li, B. Chen, S. Huang, K. Zhang, Y. Qiao, and P. Luo · 2024
Cited alongside, same era.
Closest in time.
H. Deng, D. Zou, R. Ma, H. Luo, Y. Cao, and Y. Kang · 2025
Closest in time.
Gemini 1.5 pro | generative ai on vertex ai
Google · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Vision-r1: Incentivizing reasoning capability in multimodal large language models
W. Huang, B. Jia, Z. Zhai, S. Cao, Z. Ye, F. Zhao, Z. Xu, Y. Hu, and S. Lin · 2025
Closest in time.
Autogui: Scaling gui grounding with automatic functionality annotations from llms
H. Li, J. Chen, J. Su, Y. Chen, Q. Li, and Z. Zhang · 2025
Closest in time.
Rethinking kl divergence in rlhf: From single sample to mini-batch to expectation
Y. Liu · 2025
Closest in time.
Ui-r1: Enhancing action prediction of gui agents by reinforcement learning
Z. Lu, Y. Chai, Y. Guo, X. Yin, L. Liu, H. Wang, G. Xiong, and H. Li · 2025
Closest in time.
Introducing operator
OpenAI · 2025
Closest in time.
Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl
Y. Peng, G. Zhang, M. Zhang, Z. You, J. Liu, Q. Zhu, K. Yang, X. Xu, X. Geng, and X. Yang · 2025
Closest in time.
Ui-tars: Pioneering automated gui interaction with native agents
Y. Qin, Y. Ye, J. Fang, H. Wang, S. Liang, S. Tian, J. Zhang, J. Li, Y. Li, S. Huang, et al · 2025
Closest in time.
Approximating kl divergence
J. Schulman · 2025
Closest in time.
Vlm-r1: A stable and generalizable r1-style large vision-language model
H. Shen, P. Liu, J. Li, C. Fang, Y. Ma, J. Liao, Q. Shen, Z. Zhang, K. Zhao, Q. Zhang, et al · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al · 2025
Closest in time.
Gui-r1: A generalist r1-style vision-language action model for gui agents
X. Xia and R. Luo · 2025
Closest in time.
R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization
Y. Yang, X. He, H. Pan, X. Jiang, Y. Deng, X. Yang, H. Lu, D. Yin, F. Rao, M. Zhu, et al · 2025
Closest in time.
Worldgui: Dynamic testing for comprehensive desktop gui automation
H. H. Zhao, D. Gao, and M. Z. Shou · 2025
Closest in time.
R1-zero’s" aha moment" in visual reasoning on a 2b non-sft model
H. Zhou, X. Li, R. Wang, M. Cheng, T. Zhou, and C.-J. Hsieh · 2025
Closest in time.
Ttrl: Test-time reinforcement learning
Y. Zuo, K. Zhang, S. Qu, L. Sheng, X. Zhu, B. Qi, Y. Sun, G. Cui, N. Ding, and B. Zhou · 2025
Closest in time.