Fetching the paper…
Reading the bibliography…
Recent advances in foundation models, particularly Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs), have facilitated the development of intelligent agents capable of performing complex tasks.
Reinforcement learning on web interfaces using workflow-guided exploration
Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang · 2018
Earlier work this paper cites.
Mapping natural language instructions to mobile ui action sequences
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge · 2020
Earlier work this paper cites.
Uibert: Learning generic multimodal representations for ui understanding, 2021
Chen Bai, Xiaoyu Zang, Yan Xu, Srinivas Sunkara, Abhinav Rastogi, and Jieshan Chen · 2021
Earlier work this paper cites.
Vut: Versatile ui transformer for multi-modal multi-task user interface modeling
Yang Li, Gang Li, Xin Zhou, Mostafa Dehghani, and Alexey Gritsenko · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, et al · 2021
Earlier work this paper cites.
Screen2words: Automatic mobile ui summarization with multimodal learning
Bryan Wang, Gang Li, Xin Zhou, Zhourong Chen, Tovi Grossman, and Yang Li · 2021
Earlier work this paper cites.
Screen recognition: Creating accessibility metadata for mobile applications from pixels
Xiaoyi Zhang, Lilian De Greef, Amanda Swearngin, et al · 2021
Earlier work this paper cites.
Meta-gui: Towards multi-modal conversational agents on mobile gui
Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, and Kai Yu · 2022
Earlier work this paper cites.
Language models can solve computer tasks
Geunwoo Kim, Pierre Baldi, and Stephen McAleer · 2023
Earlier work this paper cites.
Pix2struct: Screenshot parsing as pretraining for visual language understanding
Kenton Lee, Mandar Joshi, Iulia Raluca Turc, et al · 2023
Earlier work this paper cites.
Sunjae Lee, Junyoung Choi, Jungjae Lee, et al · 2023
Earlier work this paper cites.
Spotlight: Mobile ui understanding using vision-language models with a focus
Gang Li and Yang Li · 2023
Earlier work this paper cites.
Laser: Llm agent with state-space exploration for web navigation
Kaixin Ma, Hongming Zhang, Hongwei Wang, Xiaoman Pan, and Dong Yu · 2023
Earlier work this paper cites.
Autotask: Executing arbitrary voice commands by exploring and learning from mobile gui
Lihang Pan, Bowen Wang, Chun Yu, Yuxuan Chen, Xiangyu Zhang, and Yuanchun Shi · 2023
Earlier work this paper cites.
Androidinthewild: A large-scale dataset for android device control
Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy P Lillicrap · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, et al · 2023
Earlier work this paper cites.
Enabling conversational interaction with mobile ui using large language models
Bryan Wang, Gang Li, and Yang Li · 2023
Earlier work this paper cites.
Empowering LLM to use Smartphone for Intelligent Task Automation
Hao Wen, Yuanchun Li, Guohong Liu, et al · 2023
Earlier work this paper cites.
Droidbot-gpt: Gpt-powered ui automation for android
Hao Wen, Hongming Wang, Jiaxuan Liu, and Yuanchun Li · 2023
Earlier work this paper cites.
Gpt-4v in wonderland: Large multimodal models for zero-shot smartphone gui navigation
An Yan, Zhengyuan Yang, Wanrong Zhu, et al · 2023
Earlier work this paper cites.
You only look at screens: Multimodal chain-of-action agents
Zhuosheng Zhang and Aston Zhang · 2023
Earlier work this paper cites.
AppAgent: Multimodal Agents as Smartphone Users
Chi Zhang, Zhao Yang, Jiaxuan Liu, et al · 2023
Cited alongside, same era.
Webarena: A realistic web environment for building autonomous agents
Shuyan Zhou, Frank F Xu, Hao Zhu, et al · 2023
Cited alongside, same era.
Gpt-4 technical report, 2024
Josh Achiam, Steven Adler, et al · 2024
Cited alongside, same era.
Gui-world: A dataset for gui-oriented multimodal llm-based agents
Dongping Chen, Yue Huang, Siyuan Wu, et al · 2024
Cited alongside, same era.
Spa-bench: A comprehensive benchmark for smartphone agent evaluation
Jingxuan Chen, Derek Yuen, Bin Xie, et al · 2024
Cited alongside, same era.
Mobileagent: enhancing mobile control via human-machine interaction and sop integration
Androidworld: A dynamic benchmarking environment for autonomous agents
Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, et al · 2024
Closest in time.
Falcon-ui: Understanding gui before following user instructions
Huawen Shen, Chang Liu, Gengluo Li, et al · 2024
Closest in time.
Ugif-dataset: A new dataset for cross-lingual, cross-modal sequential actions on the ui
Sagar Gubbi Venkatesh, Partha Talukdar, and Srini Narayanan · 2024
Closest in time.
Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration
Junyang Wang, Haiyang Xu, Haitao Jia, et al · 2024
Closest in time.
Mobile-agent: Autonomous multi-modal mobile device agent with visual perception
Junyang Wang, Haiyang Xu, Jiabo Ye, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tinghe Ding · 2024
Cited alongside, same era.
Multimodal web navigation with instruction-finetuned foundation models
Hiroki Furuta, Kuang-Huei Lee, Ofir Nachum, et al · 2024
Cited alongside, same era.
Iris: Breaking gui complexity with adaptive focus and self-refining
Zhiqi Ge, Juncheng Li, Xinglei Pang, et al · 2024
Cited alongside, same era.
Webvoyager: Building an end-to-end web agent with large multimodal models
Hongliang He, Wenlin Yao, Kaixin Ma, et al · 2024
Cited alongside, same era.
Pc agent: While you sleep, ai works–a cognitive journey into digital world
Yanheng He, Jiahe Jin, Shijie Xia, et al · 2024
Cited alongside, same era.
Cogagent: A visual language model for gui agents
Wenyi Hong, Weihan Wang, Qingsong Lv, et al · 2024
Cited alongside, same era.
Dual-view visual contextualization for web navigation
Jihyung Kil, Chan Hee Song, Boyuan Zheng, Xiang Deng, Yu Su, and Wei-Lun Chao · 2024
Cited alongside, same era.
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, et al · 2024
Closest in time.
Autodroid: Llm-powered task automation in android
Hao Wen, Yuanchun Li, Guohong Liu, et al · 2024
Closest in time.
Mobilevlm: A vision-language model for better intra-and inter-ui understanding
Qinzhuo Wu, Weikai Xu, Wei Liu, et al · 2024
Closest in time.
Openagents: An open platform for language agents in the wild
Tianbao Xie, Fan Zhou, et al · 2024
Closest in time.
Aguvis: Unified pure vision agents for autonomous gui interaction
Yiheng Xu, Zekun Wang, Junli Wang, et al · 2024
Closest in time.
Qwen2 technical report, 2024
An Yang, Baosong Yang, et al · 2024
Closest in time.
Aria-ui: Visual grounding for gui instructions
Yuhao Yang, Yue Wang, Dongxu Li, et al · 2024
Closest in time.
Ferret-ui: Grounded mobile ui understanding with multimodal llms
Keen You, Haotian Zhang, Eldon Schoop, et al · 2024
Closest in time.
Ufo: A ui-focused agent for windows os interaction
Chaoyun Zhang, Liqun Li, Shilin He, et al · 2024
Closest in time.
Android in the zoo: Chain-of-action-thought for gui agents
Jiwen Zhang, Jihao Wu, Yihua Teng, et al · 2024
Closest in time.
Gui testing arena: A unified benchmark for advancing autonomous gui testing agent
Kangjia Zhao, Jiahui Song, Leigang Sha, et al · 2024
Closest in time.
Gpt-4v (ision) is a generalist web agent, if grounded
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su · 2024
Closest in time.
Moba: A two-level agent system for efficient mobile task automation
Zichen Zhu, Hao Tang, Yansi Li, et al · 2024
Closest in time.
Ui-tars: Pioneering automated gui interaction with native agents
Yujia Qin, Yining Ye, Junjie Fang, et al · 2025
Closest in time.
Mobile-agent-e: Self-evolving mobile assistant for complex tasks
Zhenhailong Wang, Haiyang Xu, Junyang Wang, et al · 2025
Closest in time.