Fetching the paper…
Reading the bibliography…
The rise of Large Language Models (LLMs) has revolutionized Graphical User Interface (GUI) automation through LLM-powered GUI agents, yet their ability to process sensitive data with limited human oversight raises significant privacy and security risks.
Consumer privacy: Balancing economic and justice considerations
Mary J Culnan and Robert J Bies. 2003 · 2003
Earlier work this paper cites.
Privacy as contextual integrity
Helen Nissenbaum. 2004 · 2004
Earlier work this paper cites.
Web application tests with selenium
Andreas Bruns, Andreas Kornstadt, and Dennis Wichmann. 2009 · 2009
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel. 2020 · 2020
Earlier work this paper cites.
Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI conference on human factors in computing systems . 1–16
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021 · 2021
Earlier work this paper cites.
To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making
Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021 · 2021
Earlier work this paper cites.
How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection
Maia Jacobs, Melanie F Pradier, Thomas H McCoy Jr, Roy H Perlis, Finale Doshi-Velez, and Krzysztof Z Gajos. 2021 · 2021
Earlier work this paper cites.
Everyday algorithm auditing: Understanding the power of everyday users in surfacing harmful algorithmic behaviors
Hong Shen, Alicia DeVos, Motahhare Eslami, and Kenneth Holstein. 2021 · 2021
Earlier work this paper cites.
Toward User-Driven Algorithm Auditing: Investigating users’ strategies for uncovering harmful algorithmic behavior. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–19
Alicia DeVos, Aditi Dhabalia, Hong Shen, Kenneth Holstein, and Motahhare Eslami. 2022 · 2022
Earlier work this paper cites.
End-user audits: A system empowering communities to lead large-scale investigations of harmful algorithmic behavior
Michelle S Lam, Mitchell L Gordon, Danaë Metaxa, Jeffrey T Hancock, James A Landay, and Michael S Bernstein. 2022 · 2022
Earlier work this paper cites.
Reflections on the human-algorithm complex duality perspectives in the auditing process
Adriana Tiron-Tudor and Delia Deliu. 2022 · 2022
Earlier work this paper cites.
GenAIPABench: A benchmark for generative AI-based privacy assistants
Aamir Hamid, Hemanth Reddy Samidi, Tim Finin, Primal Pappachan, and Roberto Yus. 2023 · 2023
Earlier work this paper cites.
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi. 2023 · 2023
Earlier work this paper cites.
Explanations can reduce overreliance on ai systems during decision-making
Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S Bernstein, and Ranjay Krishna. 2023 · 2023
Earlier work this paper cites.
Pre-made Empowering Artificial Intelligence and ChatGPT: The Growing Importance of Human AI-Experts. In 2023 14th International Conference on Information, Intelligence, Systems & Applications (IISA) . IEEE, 1–8
Maria Virvou and George A Tsihrintzis. 2023 · 2023
Earlier work this paper cites.
Human oversight done right: The AI Act should use humans to monitor AI only when effective
Johannes Walter. 2023 · 2023
Earlier work this paper cites.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models.. In NeurIPS
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al · 2023
Earlier work this paper cites.
Auto-gpt for online decision making: Benchmarks and additional opinions
Hui Yang, Sifu Yue, and Yunzhong He. 2023 · 2023
Earlier work this paper cites.
Appagent: Multimodal agents as smartphone users
Chi Zhang, Zhao Yang, Jiaxuan Liu, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu. 2023 · 2023
Cited alongside, same era.
Exploring Autonomous Agents through the Lens of Large Language Models: A Review
Saikat Barua. 2024 · 2024
Cited alongside, same era.
Auditing large language models for privacy compliance with specially crafted prompts
Simon Chard, Brent Johnson, and Daniel Lewis. 2024 · 2024
Cited alongside, same era.
An Empathy-Based Sandbox Approach to Bridge the Privacy Gap among Attitudes, Goals, Knowledge, and Behaviors. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24) . Association for Computing Machinery, New York, NY, USA, Article 234, 28 pages
Chaoran Chen, Weijun Li, Wenxin Song, Yanfang Ye, Yaxing Yao, and Toby Jia-Jun Li. 2024 · 2024
Cited alongside, same era.
PrivLM-Bench: A Multi-Level Privacy Evaluation Benchmark for Language Models
Amazon Science. 2024 · 2024
Later among the works it cites.
Privacylens: Evaluating privacy norm awareness of language models in action
Yijia Shao, Tianshi Li, Weiyan Shi, Yanchen Liu, and Diyi Yang. 2024 · 2024
Later among the works it cites.
On the Quest for Effectiveness in Human Oversight: Interdisciplinary Perspectives. In The 2024 ACM Conference on Fairness, Accountability, and Transparency . 2495–2507
Sarah Sterz, Kevin Baum, Sebastian Biewer, Holger Hermanns, Anne Lauber-Rönsberg, Philip Meinel, and Markus Langer. 2024 · 2024
Later among the works it cites.
Gui agents with foundation models: A comprehensive survey
Shuai Wang, Weiwen Liu, Jingxuan Chen, Weinan Gan, Xingshan Zeng, Shuai Yu, Xinlong Hao, Kun Shao, Yasheng Wang, and Ruiming Tang. 2024 · 2024
Later among the works it cites.
AutoDroid-V2: Boosting SLM-based GUI Agents via Code Generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi. 2024 · 2024
Cited alongside, same era.
Privacy preserving prompt engineering: A survey
Kennedy Edemacu and Xintao Wu. 2024 · 2024
Cited alongside, same era.
Is Human Oversight to AI Systems still possible?
Andreas Holzinger, Kurt Zatloukal, and Heimo Müller. 2024 · 2024
Cited alongside, same era.
Trustllm: Trustworthiness in large language models
Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, et al · 2024
Cited alongside, same era.
soid: A Tool for Legal Accountability for Automated Decision Making. In Computer Aided Verification , Arie Gurfinkel and Vijay Ganesh (Eds.). Springer Nature Switzerland, Cham, 233–246
Samuel Judson, Matthew Elacqua, Filip Cano, Timos Antonopoulos, Bettina Könighofer, Scott J. Shapiro, and Ruzica Piskac. 2024 · 2024
Cited alongside, same era.
Trust and reliance on AI—An experimental study on the extent and costs of overreliance on AI
Artur Klingbeil, Cassandra Grützner, and Philipp Schreck. 2024 · 2024
Cited alongside, same era.
Towards algorithm auditing: managing legal, ethical and technological risks of AI, ML and associated algorithms
Adriano Koshiyama, Emre Kazim, Philip Treleaven, Pete Rai, Lukasz Szpruch, Giles Pavey, Ghazi Ahamat, Franziska Leutner, Randy Goebel, Andrew Knight, et al · 2024
Cited alongside, same era.
Effective Human Oversight of AI-Based Systems: A Signal Detection Perspective on the Detection of Inaccurate and Unfair Outputs
Markus Langer, Kevin Baum, and Nadine Schlicker. 2024 · 2024
Cited alongside, same era.
Hao Wen, Shizuo Tian, Borislav Pavlov, Wenjie Du, Yixuan Li, Ge Chang, Shanhui Zhao, Jiacheng Liu, Yunxin Liu, Ya-Qin Zhang, and Yuanchun Li. 2024 · 2024
Later among the works it cites.
Large language model-brained gui agents: A survey
Chaoyun Zhang, Shilin He, Jiaxu Qian, Bowen Li, Liqun Li, Si Qin, Yu Kang, Minghua Ma, Qingwei Lin, Saravan Rajmohan, et al · 2024
Later among the works it cites.
Ufo: A ui-focused agent for windows os interaction
Chaoyun Zhang, Liqun Li, Shilin He, Xu Zhang, Bo Qiao, Si Qin, Minghua Ma, Yu Kang, Qingwei Lin, Saravan Rajmohan, et al · 2024
Later among the works it cites.
Agent Security Bench: A Benchmark for Evaluating Security and Privacy in LLM-Based Agents
Hanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao, Zhenting Wang, Chenlu Zhan, Hongwei Wang, and Yongfeng Zhang. 2024c · 2024
Later among the works it cites.
Attacking Vision-Language Computer Agents via Pop-ups
Yanzhe Zhang, Tao Yu, and Diyi Yang. 2024g · 2024
Later among the works it cites.
Zhiping Zhang, Bingcan Guo, and Tianshi Li. 2024a · 2024
Later among the works it cites.
Gpt-4v (ision) is a generalist web agent, if grounded
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. 2024 · 2024
Later among the works it cites.
lm-extraction-benchmark
2023 · 2025
Closest in time.
Computer use (beta)
Anthropic. 2024 · 2025
Closest in time.
Yurun Chen, Xueyu Hu, Keting Yin, Juncheng Li, and Shengyu Zhang. 2025 · 2025
Closest in time.
Artificial Intelligence Act: Article 14 - Human Oversight
European Union. 2024 · 2025
Closest in time.
Introducing Operator-Safety and privacy
OpenAI. 2025 · 2025
Closest in time.
Beyond Browsing: API-Based Web Agents
Yueqi Song, Frank Xu, Shuyan Zhou, and Graham Neubig. 2025 · 2025
Closest in time.
Users’ Mental Models of Generative AI Chatbot Ecosystems
Xingyi Wang, Xiaozheng Wang, Sunyup Park, and Yaxing Yao. 2025 · 2025
Closest in time.