Fetching the paper…
Reading the bibliography…
This paper introduces UI-TARS, a native GUI agent model that solely perceives the screenshots as input and performs human-like interactions (e.g., keyboard and mouse operations).
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Habituation: a dual-process theory
Philip M Groves and Richard F Thompson · 1970
Earlier work this paper cites.
A framework for behavioural cloning
Michael Bain and Claude Sammut · 1995
Earlier work this paper cites.
Dart: a framework for regression testing "nightly/daily builds" of gui applications
A. Memon, I. Banerjee, N. Hashmi, and A. Nagarajan · 2003
Earlier work this paper cites.
Fasttext.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov · 2016
Earlier work this paper cites.
World of bits: An open-domain platform for web-based agents
Tianlin Shi, Andrej Karpathy, Linxi Fan, Jonathan Hernandez, and Percy Liang · 2017
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang · 2018
Earlier work this paper cites.
A mobile robotic chemist
Benjamin Burger, Phillip M Maffettone, Vladimir V Gusev, Catherine M Aitchison, Yang Bai, Xiaoyan Wang, Xiaobo Li, Ben M Alston, Buyi Li, Rob Clowes, et al · 2020
Earlier work this paper cites.
Robotic process automation
Peter Hofmann, Caroline Samp, and Nils Urbach · 2020
Earlier work this paper cites.
Mapping natural language instructions to mobile UI action sequences
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge · 2020
Earlier work this paper cites.
Widget captioning: Generating natural language description for mobile user interface elements
Yang Li, Gang Li, Luheng He, Jingjie Zheng, Hong Li, and Zhiwei Guan · 2020
Earlier work this paper cites.
Roscript: A visual script driven truly non-intrusive robotic testing system for touch screen applications
Ju Qian, Zhengyu Shang, Shuoyan Yan, Yan Wang, and Lin Chen · 2020
Earlier work this paper cites.
Uibert: Learning generic multimodal representations for UI understanding
Chongyang Bai, Xiaoxue Zang, Ying Xu, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, and Blaise Agüera y Arcas · 2021
Earlier work this paper cites.
WebSRC: A dataset for web-based structural reading comprehension
Xingyu Chen, Zihan Zhao, Lu Chen, JiaBao Ji, Danyang Zhang, Ao Luo, Yuxuan Xiong, and Kai Yu · 2021
Earlier work this paper cites.
FLIN: A flexible natural language interface for web navigation
Sahisnu Mazumder and Oriana Riva · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
A dataset for interactive vision-language navigation with unknown command feasibility
Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, and Bryan A. Plummer · 2022
Earlier work this paper cites.
Robotic process automation platform uipath
Liliana Dobrica · 2022
Earlier work this paper cites.
Screenqa: Large-scale question-answer pairs over mobile app screenshots
Yu-Chung Hsiao, Fedir Zubach, Gilles Baechler, Victor Carbune, Jason Lin, Maria Wang, Srinivas Sunkara, Yun Zhu, and Jindong Chen · 2022
Earlier work this paper cites.
Learning to denoise raw mobile UI layouts for improving datasets at scale
Gang Li, Gilles Baechler, Manuel Tragut, and Yang Li · 2022
Earlier work this paper cites.
Towards better semantic understanding of mobile interfaces
Srinivas Sunkara, Maria Wang, Lijuan Liu, Gilles Baechler, Yu-Chung Hsiao, Jindong Chen, Abhanshu Sharma, and James W. W. Stout · 2022
Earlier work this paper cites.
System design for an integrated lifelong reinforcement learning agent for real-time strategy games
Indranil Sur, Zachary Daniels, Abrar Rahman, Kamil Faber, Gianmarco Gallardo, Tyler Hayes, Cameron Taylor, Mustafa Burak Gurbuz, James Smith, Sahana Joshi, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Introducing our multimodal models, 2023
Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, and Sağnak Taşırlar · 2023
Earlier work this paper cites.
Dynamic planning with a llm, 2023
Gautier Dagan, Frank Keller, and Alex Lascarides · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samual Stevens, Boshi Wang, Huan Sun, and Yu Su · 2023
Earlier work this paper cites.
Assistgui: Task-oriented desktop graphical user interface automation, 2023
Difei Gao, Lei Ji, Zechen Bai, Mingyu Ouyang, Peiran Li, Dongxing Mao, Qinchen Wu, Weichen Zhang, Peiyi Wang, Xiangwu Guo, Hengxu Wang, Luowei Zhou, and Mike Zheng Shou · 2023
Cited alongside, same era.
A real-world webagent with planning, long context understanding, and program synthesis
Izzeddin Gur, Hiroki Furuta, Austin Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust · 2023
Cited alongside, same era.
Sheetcopilot: Bringing software productivity to the next level through large language models
Hongxin Li, Jingran Su, Yuntao Chen, Qing Li, and Zhaoxiang Zhang · 2023
Cited alongside, same era.
A zero-shot language agent for computer control with structured reflection
Tao Li, Gang Li, Zhiwei Deng, Bryan Wang, and Yang Li · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Webvoyager: Building an end-to-end web agent with large multimodal models
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu · 2024
Later among the works it cites.
Cogagent: A visual language model for gui agents
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, et al · 2024
Later among the works it cites.
Os agents: A survey on mllm-based agents for general computing devices use
Xueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, et al · 2024
Later among the works it cites.
Understanding the planning of llm agents: A survey, 2024
Xu Huang, Weiwen Liu, Xiaolong Chen, Xingmei Wang, Hao Wang, Defu Lian, Yasheng Wang, Ruiming Tang, and Enhong Chen · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn · 2023
Cited alongside, same era.
Androidinthewild: A large-scale dataset for android device control
Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy P. Lillicrap · 2023
Cited alongside, same era.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Noah Shinn, Beck Labash, and Ashwin Gopinath · 2023
Cited alongside, same era.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Chan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao, Clayton Washington, and Yu Su · 2023
Cited alongside, same era.
Webwise: Web interface control and sequential exploration with large language models
Heyi Tao, Sethuraman TV, Michal Shlapentokh-Rothman, and Derek Hoiem · 2023
Cited alongside, same era.
Enabling conversational interaction with mobile UI using large language models
Bryan Wang, Gang Li, and Yang Li · 2023
Cited alongside, same era.
Droidbot-gpt: Gpt-powered ui automation for android
Hao Wen, Hongming Wang, Jiaxuan Liu, and Yuanchun Li · 2023
Cited alongside, same era.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al · 2023
Cited alongside, same era.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Later among the works it cites.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Later among the works it cites.
Raghav Kapoor, Yash Parag Butala, Melisa Russak, Jing Yu Koh, Kiran Kamble, Waseem Alshikh, and Ruslan Salakhutdinov · 2024
Later among the works it cites.
Dual-view visual contextualization for web navigation
Jihyung Kil, Chan Hee Song, Boyuan Zheng, Xiang Deng, Yu Su, and Wei-Lun Chao · 2024
Later among the works it cites.
Benchmarking mobile device control agents across diverse configurations, 2024
Juyong Lee, Taywon Min, Minyong An, Dongyoon Hahm, Haeone Lee, Changyeon Kim, and Kimin Lee · 2024
Later among the works it cites.
MUG: Interactive multimodal grounding on user interfaces
Tao Li, Gang Li, Jingjie Zheng, Purple Wang, and Yang Li · 2024
Later among the works it cites.
Comprehensive cognitive llm agent for smartphone gui automation
Xinbei Ma, Zhuosheng Zhang, and Hai Zhao · 2024
Later among the works it cites.
Aios: Llm agent operating system
Kai Mei, Zelong Li, Shuyuan Xu, Ruosong Ye, Yingqiang Ge, and Yongfeng Zhang · 2024
Later among the works it cites.
Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, et al · 2024
Later among the works it cites.
Agent q: Advanced reasoning and learning for autonomous ai agents
Pranav Putta, Edmund Mills, Naman Garg, Sumeet Motwani, Chelsea Finn, Divyansh Garg, and Rafael Rafailov · 2024
Later among the works it cites.
Tool learning with foundation models
Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Xuanhe Zhou, Yufei Huang, Chaojun Xiao, Chi Han, Yi Ren Fung, Yusheng Su, Huadong Wang, Cheng Qian, Runchu Tian, Kunlun Zhu, Shihao Liang, Xingyu Shen, Bokai Xu, Zhen Zhang, Yining Ye, Bowen Li, Ziwei Tang, Jing Yi, Yuzhang Zhu, Zhenning Dai, Lan Yan, Xin Cong, Yaxi Lu, Weilin Zhao, Yuxiang Huang, Junxi Yan, Xu Han, Xian Sun, Dahai Li, Jason Phang, Cheng Yang, Tongshuang Wu, Heng Ji, Guoliang Li, Zhiyuan Liu, and Maosong Sun · 2024
Later among the works it cites.
Hariprasauth Ramamoorthy, Shubhankar Gupta, and Suresh Sundaram · 2024
Later among the works it cites.
Self-reflection in llm agents: Effects on problem-solving performance, 2024
Matthew Renze and Erhan Guven · 2024
Later among the works it cites.
A collective ai via lifelong learning and sharing at the edge
Andrea Soltoggio, Eseoghene Ben-Iwhiwhu, Vladimir Braverman, Eric Eaton, Benjamin Epstein, Yunhao Ge, Lucy Halperin, Jonathan How, Laurent Itti, Michael A Jacobs, et al · 2024
Later among the works it cites.
Beyond browsing: Api-based web agents
Yueqi Song, Frank F Xu, Shuyan Zhou, and Graham Neubig · 2024
Later among the works it cites.
Towards general computer control: A multimodal agent for red dead redemption ii as a case study
Weihao Tan, Ziluo Ding, Wentao Zhang, Boyu Li, Bohan Zhou, Junpeng Yue, Haochong Xia, Jiechuan Jiang, Longtao Zheng, Xinrun Xu, et al · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al · 2024
Later among the works it cites.
Oscar: Operating system control via state-aware reasoning and re-planning
Xiaoqiang Wang and Bang Liu · 2024
Later among the works it cites.
Agentless: Demystifying llm-based software engineering agents
Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang · 2024
Later among the works it cites.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu · 2024
Later among the works it cites.
Aguvis: Unified pure vision agents for autonomous gui interaction
Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu, and Caiming Xiong · 2024
Later among the works it cites.
Screenspot-pro: Gui grounding for professional high-resolution computer use, 2025
Kaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo, Yuchen Tian, Jing Ma, Zhiyong Huang, and Tat-Seng Chua · 2025
Closest in time.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al · 2025
Closest in time.