Fetching the paper…
Reading the bibliography…
Comprehensive evaluation of mobile agents can significantly advance their development and real-world applicability.
“Reinforcement learning for mapping instructions to actions,”
Satchuthananthavale RK Branavan, Harr Chen, Luke Zettlemoyer, and Regina Barzilay, · 2009
Earlier work this paper cites.
“Mapping natural language instructions to mobile ui action sequences,”
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge, · 2020
Earlier work this paper cites.
“Mapping natural language instructions to mobile ui action sequences,”
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge, · 2020
Earlier work this paper cites.
“Appbuddy: Learning to accomplish tasks in mobile apps via reinforcement learning,”
Maayan Shvo, Zhiming Hu, Rodrigo Toro Icarte, Iqbal Mohomed, Allan Jepson, and Sheila A McIlraith, · 2021
Earlier work this paper cites.
“Environment generation for zero-shot compositional reinforcement learning,”
Izzeddin Gur, Natasha Jaques, Yingjie Miao, Jongwook Choi, Manoj Tiwari, Honglak Lee, and Aleksandra Faust, · 2021
Earlier work this paper cites.
“Ugif: Ui grounded instruction following,”
Sagar Gubbi Venkatesh, Partha Talukdar, and Srini Narayanan, · 2022
Earlier work this paper cites.
“A data-driven approach for learning to control computers,”
Peter C Humphreys, David Raposo, Tobias Pohlen, Gregory Thornton, Rachita Chhaparia, Alistair Muldal, Josh Abramson, Petko Georgiev, Adam Santoro, and Timothy Lillicrap, · 2022
Earlier work this paper cites.
Sunjae Lee, Junyoung Choi, Jungjae Lee, Munim Hasan Wasi, Hojun Choi, Steven Y Ko, Sangeun Oh, and Insik Shin, · 2023
Earlier work this paper cites.
“Androidinthewild: A large-scale dataset for android device control,”
Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy Lillicrap, · 2023
Earlier work this paper cites.
“Motif: Intrinsic motivation from artificial intelligence feedback,”
Martin Klissarov, Pierluca D’Oro, Shagun Sodhani, Roberta Raileanu, Pierre-Luc Bacon, Pascal Vincent, Amy Zhang, and Mikael Henaff, · 2023
Earlier work this paper cites.
“A zero-shot language agent for computer control with structured reflection,”
Tao Li, Gang Li, Zhiwei Deng, Bryan Wang, and Yang Li, · 2023
Earlier work this paper cites.
“Reflexion: Language agents with verbal reinforcement learning,”
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao, · 2023
Earlier work this paper cites.
“Gptscore: Evaluate as you desire,”
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu, · 2023
Earlier work this paper cites.
“Chateval: Towards better llm-based evaluators through multi-agent debate,”
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu, · 2023
Earlier work this paper cites.
“Judging llm-as-a-judge with mt-bench and chatbot arena,”
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al., · 2023
Cited alongside, same era.
“Gemini: a family of highly capable multimodal models,”
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al., · 2023
Cited alongside, same era.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al., · 2023
Cited alongside, same era.
“You only look at screens: Multimodal chain-of-action agents,”
Zhuosheng Zhang and Aston Zhang, · 2024
Cited alongside, same era.
“Axnav: Replaying accessibility tests from natural language,”
Maryam Taeb, Amanda Swearngin, Eldon Schoop, Ruijia Cheng, Yue Jiang, and Jeffrey Nichols, · 2024
Later among the works it cites.
“Omniparser for pure vision based gui agent,”
Yadong Lu, Jianwei Yang, Yelong Shen, and Ahmed Awadallah, · 2024
Later among the works it cites.
“Gui odyssey: A comprehensive dataset for cross-app gui navigation on mobile devices,”
Quanfeng Lu, Wenqi Shao, Zitao Liu, Fanqing Meng, Boxuan Li, Botong Chen, Siyuan Huang, Kaipeng Zhang, Yu Qiao, and Ping Luo, · 2024
Later among the works it cites.
“Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning,”
Hao Bai, Yifei Zhou, Jiayi Pan, Mert Cemri, Alane Suhr, Sergey Levine, and Aviral Kumar, · 2024
Later among the works it cites.
“Os-copilot: Towards generalist computer agents with self-improvement,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Songqin Nong, Jiali Zhu, Rui Wu, Jiongchao Jin, Shuo Shan, Xiutian Huang, and Wenhao Xu, · 2024
Cited alongside, same era.
“Autodroid: Llm-powered task automation in android,”
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu, · 2024
Cited alongside, same era.
“Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration,”
Junyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang, · 2024
Cited alongside, same era.
“Cogagent: A visual language model for gui agents,”
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, et al., · 2024
Cited alongside, same era.
“Understanding the weakness of large language model agents within a complex android environment,”
Mingzhe Xing, Rongkai Zhang, Hui Xue, Qi Chen, Fan Yang, and Zhen Xiao, · 2024
Cited alongside, same era.
“Llamatouch: A faithful and scalable testbed for mobile ui task automation,”
Li Zhang, Shihe Wang, Xianqing Jia, Zhihan Zheng, Yunhe Yan, Longxi Gao, Yuanchun Li, and Mengwei Xu, · 2024
Cited alongside, same era.
“On the effects of data scale on ui control agents,”
Wei Li, William E Bishop, Alice Li, Christopher Rawles, Folawiyo Campbell-Ajala, Divya Tyamagundlu, and Oriana Riva, · 2024
Cited alongside, same era.
“Androidlab: Training and systematic benchmarking of android autonomous agents,”
Yifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng, Hao Yu, Hanyu Lai, Shudan Zhang, Dan Zhang, Jie Tang, and Yuxiao Dong, · 2024
Cited alongside, same era.
Zhiyong Wu, Chengcheng Han, Zichen Ding, Zhenmin Weng, Zhoumianze Liu, Shunyu Yao, Tao Yu, and Lingpeng Kong, · 2024
Later among the works it cites.
“Agent-as-a-judge: Evaluate agents with agents,”
Mingchen Zhuge, Changsheng Zhao, Dylan Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, et al., · 2024
Later among the works it cites.
“Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark,”
Dongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang, Yinuo Liu, Huichi Zhou, Qihui Zhang, Yao Wan, Pan Zhou, and Lichao Sun, · 2024
Later among the works it cites.
“Auto-arena: Automating llm evaluations with agent peer battles and committee discussions,”
Ruochen Zhao, Wenxuan Zhang, Yew Ken Chia, Weiwen Xu, Deli Zhao, and Lidong Bing, · 2024
Later among the works it cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al., · 2024
Later among the works it cites.
“Deepseek-v3 technical report,”
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al., · 2024
Later among the works it cites.
“Appagent: Multimodal agents as smartphone users,”
Chi Zhang, Zhao Yang, Jiaxuan Liu, Yanda Li, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu, · 2025
Closest in time.
“Mobile-agent-e: Self-evolving mobile assistant for complex tasks,”
Zhenhailong Wang, Haiyang Xu, Junyang Wang, Xi Zhang, Ming Yan, Ji Zhang, Fei Huang, and Heng Ji, · 2025
Closest in time.
“Ui-tars: Pioneering automated gui interaction with native agents,”
Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, et al., · 2025
Closest in time.