Fetching the paper…
Reading the bibliography…
A Large Language Model (LLM) powered GUI agent is a specialized autonomous system that performs tasks on the user's behalf according to high-level instructions.
The theory of affordances:(1979)
James J Gibson. 2014 · 1979
Earlier work this paper cites.
Cognitive load during problem solving: Effects on learning
John Sweller. 1988 · 1988
Earlier work this paper cites.
In two minds: dual-process accounts of reasoning
Jonathan St BT Evans. 2003 · 2003
Earlier work this paper cites.
Visual attention: the where, what, how and why of saliency
Stefan Treue. 2003 · 2003
Earlier work this paper cites.
Privacy as contextual integrity
Helen Nissenbaum. 2004 · 2004
Earlier work this paper cites.
Handbook of inter-rater reliability: The definitive guide to measuring the extent of agreement among raters
Kilem L Gwet. 2014 · 2014
Earlier work this paper cites.
Membership inference attacks against machine learning models. In 2017 IEEE symposium on security and privacy (SP) . IEEE, 3–18
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Earlier work this paper cites.
The dark (patterns) side of UX design. In Proceedings of the 2018 CHI conference on human factors in computing systems . 1–14
Colin M Gray, Yubo Kou, Bryan Battles, Joseph Hoggatt, and Austin L Toombs. 2018 · 2018
Earlier work this paper cites.
Extracting training data from large language models. In 30th USENIX security symposium (USENIX Security 21) . 2633–2650
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Extracting training data from diffusion models. In 32nd USENIX Security Symposium (USENIX Security 23) . 5253–5270
Nicolas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramer, Borja Balle, Daphne Ippolito, and Eric Wallace. 2023 · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023 · 2023
Cited alongside, same era.
Assistgui: Task-oriented desktop graphical user interface automation
Difei Gao, Lei Ji, Zechen Bai, Mingyu Ouyang, Peiran Li, Dongxing Mao, Qinchen Wu, Weichen Zhang, Peiyi Wang, Xiangwu Guo, et al · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2023 · 2023
Cited alongside, same era.
Agentbench: Evaluating llms as agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al · 2023
Cited alongside, same era.
Eia: Environmental injection attack on generalist web agents for privacy leakage
Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang, Jiawei Zhang, Chaowei Xiao, Yuan Tian, Bo Li, and Huan Sun. 2024 · 2024
Later among the works it cites.
Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, Xintong Li, Jing Shi, Hongjie Chen, Viet Dac Lai, Zhouhang Xie, Sungchul Kim, Ruiyi Zhang, Tong Yu, Mehrab Tanjim, Nesreen K. Ahmed, Puneet Mathur, Seunghyun Yoon, Lina Yao, Branislav Kveton, Thien Huu Nguyen, Trung Bui, Tianyi Zhou, Ryan A. Rossi, and Franck Dernoncourt. 2024 · 2024
Later among the works it cites.
Privacylens: Evaluating privacy norm awareness of language models in action
Yijia Shao, Tianshi Li, Weiyan Shi, Yanchen Liu, and Diyi Yang. 2024 · 2024
Later among the works it cites.
Security matrix for multimodal agents on mobile devices: A systematic and proof of concept study
Yulong Yang, Xinshan Yang, Shuaidong Li, Chenhao Lin, Zhengyu Zhao, Chao Shen, and Tianwei Zhang. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi. 2023 · 2023
Cited alongside, same era.
Gpt-4v in wonderland: Large multimodal models for zero-shot smartphone gui navigation
An Yan, Zhengyuan Yang, Wanrong Zhu, Kevin Lin, Linjie Li, Jianfeng Wang, Jianwei Yang, Yiwu Zhong, Julian McAuley, Jianfeng Gao, et al · 2023
Cited alongside, same era.
Appagent: Multimodal agents as smartphone users
Chi Zhang, Zhao Yang, Jiaxuan Liu, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu. 2023 · 2023
Cited alongside, same era.
AirGapAgent: Protecting privacy-conscious conversational agents. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security . 3868–3882
Eugene Bagdasarian, Ren Yi, Sahra Ghalebikesabi, Peter Kairouz, Marco Gruteser, Sewoong Oh, Borja Balle, and Daniel Ramage. 2024 · 2024
Cited alongside, same era.
Generalist virtual agents: A survey on autonomous agents across digital platforms
Minghe Gao, Wendong Bu, Bingchen Miao, Yang Wu, Yunfei Li, Juncheng Li, Siliang Tang, Qi Wu, Yueting Zhuang, and Meng Wang. 2024 · 2024
Cited alongside, same era.
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
Siyuan Hu, Mingyu Ouyang, Difei Gao, and Mike Zheng Shou. 2024 · 2024
Cited alongside, same era.
Dissecting Adversarial Robustness of Multimodal LM Agents. In The Thirteenth International Conference on Learning Representations
Chen Henry Wu, Rishi Rajesh Shah, Jing Yu Koh, Russ Salakhutdinov, Daniel Fried, and Aditi Raghunathan. [n. d.]
Cited in the paper.
Later among the works it cites.
Ufo: A ui-focused agent for windows os interaction
Chaoyun Zhang, Liqun Li, Shilin He, Xu Zhang, Bo Qiao, Si Qin, Minghua Ma, Yu Kang, Qingwei Lin, Saravan Rajmohan, et al · 2024
Later among the works it cites.
Attacking Vision-Language Computer Agents via Pop-ups
Yanzhe Zhang, Tao Yu, and Diyi Yang. 2024b · 2024
Later among the works it cites.
Gpt-4v (ision) is a generalist web agent, if grounded
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. 2024 · 2024
Later among the works it cites.
Computer use (beta)
Anthropic. 2024 · 2025
Closest in time.
Introducing Operator-Safety and privacy
OpenAI. 2025 · 2025
Closest in time.