Fetching the paper…
Reading the bibliography…
Recent Graphical User Interface (GUI) agents replicate the R1-Zero paradigm, coupling online Reinforcement Learning (RL) with explicit chain-of-thought reasoning prior to object grounding and thereby achieving substantial performance gains.
Thinking, Fast and Slow
D. Kahneman · 2011
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Buy 4 reinforce samples, get a baseline for free!
W. Kool, H. van Hoof, and M. Welling · 2019
Earlier work this paper cites.
Uibert: Learning generic multimodal representations for ui understanding
C. Bai, X. Zang, Y. Xu, S. Sunkara, A. Rastogi, J. Chen, et al · 2021
Earlier work this paper cites.
Vut: Versatile ui transformer for multi-modal multi-task user interface modeling
Y. Li, G. Li, X. Zhou, M. Dehghani, and A. Gritsenko · 2021
Earlier work this paper cites.
Spotlight: Mobile ui understanding using vision-language models with a focus
G. Li and Y. Li · 2023
Earlier work this paper cites.
Reinforced ui instruction grounding: Towards a generic ui task automation api
Z. Zhang, W. Xie, X. Zhang, and Y. Lu · 2023
Earlier work this paper cites.
Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms
A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, O. Pietquin, A. Üstün, and S. Hooker · 2024
Earlier work this paper cites.
Amex: Android multi-annotation expo dataset for mobile gui agents
Y. Chai, S. Huang, Y. Niu, H. Xiao, L. Liu, D. Zhang, P. Gao, S. Ren, and H. Li · 2024
Earlier work this paper cites.
Seeclick: Harnessing gui grounding for advanced visual gui agents
K. Cheng, Q. Sun, Y. Chu, F. Xu, L. YanTao, J. Zhang, and Z. Wu · 2024
Earlier work this paper cites.
Cogagent: A visual language model for gui agents
W. Hong, W. Wang, Q. Lv, J. Xu, W. Yu, J. Ji, Y. Wang, Z. Wang, Y. Dong, M. Ding, et al · 2024
Earlier work this paper cites.
On the effects of data scale on computer control agents
W. Li, W. Bishop, A. Li, C. Rawles, F. Campbell-Ajala, D. Tyamagundlu, and O. Riva · 2024
Earlier work this paper cites.
Ferret-ui 2: Mastering universal user interface understanding across platforms
Z. Li, K. You, H. Zhang, D. Feng, H. Agrawal, X. Li, M. P. S. Moorthy, J. Nichols, Y. Yang, and Z. Gan · 2024
Earlier work this paper cites.
Showui: One vision-language-action model for gui visual agent
K. Q. Lin, L. Li, D. Gao, Z. Yang, S. Wu, Z. Bai, W. Lei, L. Wang, and M. Z. Shou · 2024
Earlier work this paper cites.
Autoglm: Autonomous foundation agents for guis
X. Liu, B. Qin, D. Liang, G. Dong, H. Lai, H. Zhang, H. Zhao, I. L. Iong, J. Sun, J. Wang, et al · 2024
Earlier work this paper cites.
Omniparser for pure vision based gui agent
Y. Lu, J. Yang, Y. Shen, and A. Awadallah · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al · 2024
Cited alongside, same era.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, et al · 2024
Cited alongside, same era.
Gui agents with foundation models: A comprehensive survey
S. Wang, W. Liu, J. Chen, Y. Zhou, W. Gan, X. Zeng, Y. Che, S. Yu, X. Hao, K. Shao, et al · 2024
Cited alongside, same era.
Aguvis: Unified pure vision agents for autonomous gui interaction
Infigui-r1: Advancing multimodal gui agents from reactive actors to deliberative reasoners
Y. Liu, P. Li, C. Xie, X. Hu, X. Han, S. Zhang, H. Yang, and F. Wu · 2025
Closest in time.
Understanding r1-zero-like training: A critical perspective
Z. Liu, C. Chen, W. Li, P. Qi, T. Pang, C. Du, W. S. Lee, and M. Lin · 2025
Closest in time.
Ui-r1: Enhancing action prediction of gui agents by reinforcement learning
Z. Lu, Y. Chai, Y. Guo, X. Yin, L. Liu, H. Wang, G. Xiong, and H. Li · 2025
Closest in time.
Mm-eureka: Exploring the frontiers of multimodal reasoning with rule-based reinforcement learning
F. Meng, L. Du, Z. Liu, Z. Zhou, Q. Lu, D. Fu, T. Han, B. Shi, W. Wang, J. He, et al · 2025
Closest in time.
Learning to reason with llms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Xu, Z. Wang, J. Wang, D. Lu, T. Xie, A. Saha, D. Sahoo, T. Yu, and C. Xiong · 2024
Cited alongside, same era.
Aria-ui: Visual grounding for gui instructions
Y. Yang, Y. Wang, D. Li, Z. Luo, B. Chen, C. Huang, and J. Li · 2024
Cited alongside, same era.
Developing a computer use model
Anthropic · 2025
Cited alongside, same era.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al · 2025
Cited alongside, same era.
An empirical study on eliciting and improving r1-like reasoning models
Z. Chen, Y. Min, B. Zhang, J. Chen, J. Jiang, D. Cheng, W. X. Zhao, Z. Liu, X. Miao, Y. Lu, et al · 2025
Cited alongside, same era.
Gemini-2.0 (project mariner)
G. DeepMind · 2025
Cited alongside, same era.
Navigating the digital world as humans do: Universal visual grounding for gui agents
B. Gou, R. Wang, B. Zheng, Y. Xie, C. Chang, Y. Shu, H. Sun, and Y. Su · 2025
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Cited alongside, same era.
OpenAI · 2025
Closest in time.
Gpt-4o, 2024
OpenAI · 2025
Closest in time.
Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl
Y. Peng, G. Zhang, M. Zhang, Z. You, J. Liu, Q. Zhu, K. Yang, X. Xu, X. Geng, and X. Yang · 2025
Closest in time.
Ui-tars: Pioneering automated gui interaction with native agents
Y. Qin, Y. Ye, J. Fang, H. Wang, S. Liang, S. Tian, J. Zhang, J. Li, Y. Li, S. Huang, et al · 2025
Closest in time.
Vlm-r1: A stable and generalizable r1-style large vision-language model
H. Shen, P. Liu, J. Li, C. Fang, Y. Ma, J. Liao, Q. Shen, Z. Zhang, K. Zhao, Q. Zhang, R. Xu, and T. Zhao · 2025
Closest in time.
K. Team, A. Du, B. Yin, B. Xing, B. Qu, B. Wang, C. Chen, C. Zhang, C. Du, C. Wei, et al · 2025
Closest in time.
Os-atlas: A foundation action model for generalist gui agents
Z. Wu, Z. Wu, F. Xu, Y. Wang, Q. Sun, C. Jia, K. Cheng, Z. Ding, L. Chen, P. P. Liang, et al · 2025
Closest in time.
Gui-r1: A generalist r1-style vision-language action model for gui agents
X. Xia and R. Luo · 2025
Closest in time.
Does chain-of-thought reasoning help mobile gui agent? an empirical study
L. Zhang, L. Gao, and M. Xu · 2025
Closest in time.
R1-zero’s" aha moment" in visual reasoning on a 2b non-sft model
H. Zhou, X. Li, R. Wang, M. Cheng, T. Zhou, and C.-J. Hsieh · 2025
Closest in time.
Chop: Mobile operating assistant with constrained high-frequency optimized subtask planning
Y. Zhou, S. Wang, S. Dai, Q. Jia, Z. Du, Z. Dong, and J. Xu · 2025
Closest in time.