Fetching the paper…
Reading the bibliography…
Recently, Multimodal Large Language Models (MLLMs) have been used as agents to control keyboard and mouse inputs by directly perceiving the Graphical User Interface (GUI) and generating corresponding commands.
Mapping natural language instructions to mobile ui action sequences
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge · 2005
Earlier work this paper cites.
Widget captioning: Generating natural language description for mobile user interface elements
Yang Li, Gang Li, Luheng He, Jingjie Zheng, Hong Li, and Zhiwei Guan · 2010
Earlier work this paper cites.
Gui ferret: Gui test tool to analyze complex behavior of multi-window applications
Hajime Nakajima, Takeshi Masuda, and Ikuya Takahashi · 2013
Earlier work this paper cites.
Test data generation based on gui: A systematic mapping
Rodrigo Funabashi Jorge, Márcio Eduardo Delamaro, Celso Gonçalves Camilo-Junior, and Auri Marcelo Rizzo Vincenzi · 2014
Earlier work this paper cites.
ios applications testing
Ivans Kulesovs · 2015
Earlier work this paper cites.
A key volume mining deep framework for action recognition
Wangjiang Zhu, Jie Hu, Gang Sun, Xudong Cao, and Yu Qiao · 2016
Earlier work this paper cites.
Rico: A mobile app dataset for building data-driven design applications
Biplab Deka, Zifeng Huang, Chad Franzen, Joshua Hibschman, Daniel Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar · 2017
Earlier work this paper cites.
Unsupervised video summarization with adversarial lstm networks
Behrooz Mahasseni, Michael Lam, and Sinisa Todorovic · 2017
Earlier work this paper cites.
pix2code: Generating code from a graphical user interface screenshot
Tony Beltramelli · 2018
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang · 2018
Earlier work this paper cites.
Deep keyframe detection in human action videos
Xiang Yan, Syed Zulqarnain Gilani, Hanlin Qin, Mingtao Feng, Liang Zhang, and Ajmal Mian · 2018
Earlier work this paper cites.
Understanding user interface preferences for xr environments when exploring physics and engineering principles
Brian K. Sanders, Yuzhong Shen, and Dennis A. Vincenzi · 2019
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Screen2words: Automatic mobile ui summarization with multimodal learning
Bryan Wang, Gang Li, Xin Zhou, Zhourong Chen, Tovi Grossman, and Yang Li · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Earlier work this paper cites.
Spotlight: Mobile ui understanding using vision-language models with a focus
Gang Li and Yang Li · 2022
Earlier work this paper cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang · 2022
Earlier work this paper cites.
R3m: A universal visual representation for robot manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta · 2022
Earlier work this paper cites.
What is xr? towards a framework for augmented and virtual reality
Philipp A. Rauschnabel, Reto Felix, Christian Hinsch, Hamza Shahab, and Florain Alt · 2022
Earlier work this paper cites.
Meta-gui: towards multi-modal conversational agents on mobile gui
Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, and Kai Yu · 2022
Earlier work this paper cites.
Ugif: Ui grounded instruction following
Sagar Gubbi Venkatesh, Partha Talukdar, and Srini Narayanan · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Albert Li, Pascale Fung, and Steven C. H. Hoi · 2023
Cited alongside, same era.
An embodied generalist agent in 3d world
Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang · 2023
Cited alongside, same era.
Peng Jin, Ryuichi Takanobu, Caiwan Zhang, Xiaochun Cao, and Li Yuan · 2023
Cited alongside, same era.
Vision2ui: A real-world dataset with layout for code generation from ui designs
Yi Gui, Zhen Li, Yao Wan, Yemin Shi, Hongyu Zhang, Yi Su, Shaoling Dong, Xing Zhou, and Wenbin Jiang · 2024
Closest in time.
Cogagent: A visual language model for gui agents
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, et al · 2024
Closest in time.
Raghav Kapoor, Yash Parag Butala, Melisa Russak, Jing Yu Koh, Kiran Kamble, Waseem Alshikh, and Ruslan Salakhutdinov · 2024
Closest in time.
Katna - video to keyframes and video to image crops utility
KeplerLab · 2024
Closest in time.
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A systematic review on pattern-based gui testing of android and web apps: State-of-the-art, taxonomy, challenges and future directions
Ambreen Kousar, Saif Ur Rehman Khan, Shahid Hussain, M. Abdul Basit Ur Rahim, Wen-Li Wang, and Naseem Ibrahim · 2023
Cited alongside, same era.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan · 2023
Cited alongside, same era.
Gaia: a benchmark for general ai assistants
Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom · 2023
Cited alongside, same era.
Chatgpt, 2023
OpenAI · 2023
Cited alongside, same era.
Openai models - gpt-4-vision
OpenAI · 2023
Cited alongside, same era.
Android in the wild: A large-scale dataset for android device control
Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy Lillicrap · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Cited alongside, same era.
Gpt-4v (ision) for robotics: Multimodal task planning from human demonstration
Naoki Wake, Atsushi Kanehira, Kazuhiro Sasabuchi, Jun Takamatsu, and Katsushi Ikeuchi · 2023
Cited alongside, same era.
Closest in time.
Autowebglm: Bootstrap and reinforce a large language model-based web navigating agent
Hanyu Lai, Xiao Liu, Iat Long Iong, Shuntian Yao, Yuxuan Chen, Pengbo Shen, Hao Yu, Hanchen Zhang, Xiaohan Zhang, Yuxiao Dong, et al · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee · 2024
Closest in time.
Weblinx: Real-world website navigation with multi-turn dialogue
Xing Han Lù, Zdeněk Kasner, and Siva Reddy · 2024
Closest in time.
Screenagent: A vision language model-driven computer control agent
Runliang Niu, Jindong Li, Shiqi Wang, Yali Fu, Xiyu Hu, Xueyuan Leng, He Kong, Yi Chang, and Qi Wang · 2024
Closest in time.
Hello gpt-4o, May 2024
OpenAI · 2024
Closest in time.
Videogui: A benchmark for gui automation from instructional videos
Kevin Qinghong Lin, Linjie Li, Difei Gao, Qinchen WU, Mingyi Yan, Zhengyuan Yang, Lijuan Wang, and Mike Zheng Shou · 2024
Closest in time.
Mmac-copilot: Multi-modal agent collaboration operating system copilot
Zirui Song, Yaohang Li, Meng Fang, Zhenhao Chen, Zecheng Shi, and Yuan Huang · 2024
Closest in time.
Towards general computer control: A multimodal agent for red dead redemption ii as a case study
Weihao Tan, Ziluo Ding, Wentao Zhang, Boyu Li, Bohan Zhou, Junpeng Yue, Haochong Xia, Jiechuan Jiang, Longtao Zheng, Xinrun Xu, et al · 2024
Closest in time.
Mobile-agent: Autonomous multi-modal mobile device agent with visual perception
Junyang Wang, Haiyang Xu, Jiabo Ye, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang · 2024
Closest in time.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al · 2024
Closest in time.
Justice or prejudice? quantifying biases in llm-as-a-judge
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, et al · 2024
Closest in time.
A survey on multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen · 2024
Closest in time.
Ferret-ui: Grounded mobile ui understanding with multimodal llms
Keen You, Haotian Zhang, Eldon Schoop, Floris Weers, Amanda Swearngin, Jeffrey Nichols, Yinfei Yang, and Zhe Gan · 2024
Closest in time.
Pairwise gui dataset construction between android phones and tablets
Haolan Zhan, Yujin Huang, Di Liu, et al · 2024
Closest in time.
LLM-as-a-coauthor: Can mixed human-written and machine-generated text be detected?
Qihui Zhang, Chujie Gao, Dongping Chen, Yue Huang, Yixin Huang, Zhenyang Sun, Shilin Zhang, Weiye Li, Zhengyan Fu, Yao Wan, and Lichao Sun · 2024
Closest in time.
Universal visual decomposer: Long-horizon manipulation made easy
Zichen Zhang, Yunshuang Li, Osbert Bastani, Abhishek Gupta, Dinesh Jayaraman, Yecheng Jason Ma, and Luca Weihs · 2024
Closest in time.
NL2Formula: Generating spreadsheet formulas from natural language queries
Wei Zhao, Zhitao Hou, Siyuan Wu, Yan Gao, Haoyu Dong, Yao Wan, Hongyu Zhang, Yulei Sui, and Haidong Zhang · 2024
Closest in time.
On the trustworthiness of generative foundation models: Guideline, assessment, and perspective
Yue Huang, Chujie Gao, Siyuan Wu, Haoran Wang, Xiangqi Wang, Yujun Zhou, Yanbo Wang, Jiayi Ye, Jiawen Shi, Qihui Zhang, et al · 2025
Closest in time.