Fetching the paper…
Reading the bibliography…
With the advancement of web techniques, they have significantly revolutionized various aspects of people's lives.
Q-learning
Christopher JCH Watkins and Peter Dayan. 1992 · 1992
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore. 1996 · 1996
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour. 1999 · 1999
Earlier work this paper cites.
Information retrieval on the web
Mei Kobayashi and Koichi Takeda. 2000 · 2000
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
Satinder Singh, Tommi Jaakkola, Michael L Littman, and Csaba Szepesvári. 2000 · 2000
Earlier work this paper cites.
Value-function reinforcement learning in Markov games
Michael L Littman. 2001 · 2001
Earlier work this paper cites.
Drivers of Internet shopping
Mohamed Khalifa and Moez Limayem. 2003 · 2003
Earlier work this paper cites.
Semantic wikipedia. In Proceedings of the 15th international conference on World Wide Web . 585–594
Max Völkel, Markus Krötzsch, Denny Vrandecic, Heiko Haller, and Rudi Studer. 2006 · 2006
Earlier work this paper cites.
Autonomously semantifying wikipedia. In Proceedings of the sixteenth ACM conference on Conference on information and knowledge management . 41–50
Fei Wu and Daniel S Weld. 2007 · 2007
Earlier work this paper cites.
Kernelized value function approximation for reinforcement learning. In Proceedings of the 26th annual international conference on machine learning . 1017–1024
Gavin Taylor and Ronald Parr. 2009 · 2009
Earlier work this paper cites.
A survey of monte carlo tree search methods
Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton. 2012 · 2012
Earlier work this paper cites.
Structured communication-centered programming for web services
Marco Carbone, Kohei Honda, and Nobuko Yoshida. 2012 · 2012
Earlier work this paper cites.
A survey on solutions and main free tools for privacy enhancing Web communications
Antonio Ruiz-Martínez. 2012 · 2012
Earlier work this paper cites.
A survey of OCR applications
Amarjot Singh, Ketan Bacchuwar, and Akshay Bhasin. 2012 · 2012
Earlier work this paper cites.
Online shopping
Yi Cai and Brenda J Cude. 2016 · 2016
Earlier work this paper cites.
Web news mining in an evolving framework
José Antonio Iglesias, Alexandra Tiemblo, Agapito Ledezma, and Araceli Sanchis. 2016 · 2016
Earlier work this paper cites.
A survey of Web crawlers for information retrieval
Manish Kumar, Rajesh Bhatia, and Dhavleesh Rattan. 2017 · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Deep learning for computer vision: A brief review
Athanasios Voulodimos, Nikolaos Doulamis, Anastasios Doulamis, and Eftychios Protopapadakis. 2018 · 2018
Earlier work this paper cites.
On the use of arxiv as a dataset
Colin B Clement, Matthew Bierbaum, Kevin P O’Keeffe, and Alexander A Alemi. 2019 · 2019
Earlier work this paper cites.
Deep reinforcement learning for optimizing finance portfolio management. In 2019 amity international conference on artificial intelligence (AICAI) . IEEE, 14–20
Yuh-Jong Hu and Shang-Jen Lin. 2019 · 2019
Earlier work this paper cites.
An overview of online fake news: Characterization, detection, and discussion
Xichen Zhang and Ali A Ghorbani. 2020 · 2020
Earlier work this paper cites.
Attacking black-box recommendations via copying cross-domain user profiles. In 2021 IEEE 37th international conference on data engineering (ICDE) . IEEE, 1583–1594
Wenqi Fan, Tyler Derr, Xiangyu Zhao, Yao Ma, Hui Liu, Jianping Wang, Jiliang Tang, and Qing Li. 2021 · 2021
Earlier work this paper cites.
Deep reinforcement learning for autonomous driving: A survey
B Ravi Kiran, Ibrahim Sobh, Victor Talpaert, Patrick Mannion, Ahmad A Al Sallab, Senthil Yogamani, and Patrick Pérez. 2021 · 2021
Earlier work this paper cites.
Policy learning with constraints in model-free reinforcement learning: A survey. In The 30th international joint conference on artificial intelligence (ijcai)
Yongshuai Liu, Avishai Halev, and Xin Liu. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PmLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar. 2021 · 2021
Earlier work this paper cites.
Knowledge-enhanced black-box attacks for recommendations. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 108–117
Jingfan Chen, Wenqi Fan, Guanghui Zhu, Xiangyu Zhao, Chunfeng Yuan, Qing Li, and Yihua Huang. 2022 · 2022
Earlier work this paper cites.
A data-driven approach for learning to control computers. In International Conference on Machine Learning . PMLR, 9466–9482
Peter C Humphreys, David Raposo, Tobias Pohlen, Gregory Thornton, Rachita Chhaparia, Alistair Muldal, Josh Abramson, Petko Georgiev, Adam Santoro, and Timothy Lillicrap. 2022 · 2022
Earlier work this paper cites.
Spotlight: Mobile ui understanding using vision-language models with a focus
Gang Li and Yang Li. 2022 · 2022
Earlier work this paper cites.
Trustworthy AI: A computational perspective
Haochen Liu, Yiqi Wang, Wenqi Fan, Xiaorui Liu, Yaxin Li, Shaili Jain, Yunhao Liu, Anil Jain, and Jiliang Tang. 2022 · 2022
Earlier work this paper cites.
Reinforcement learning in robotic applications: a comprehensive survey
Bharat Singh, Rajesh Kumar, and Vinay Pratap Singh. 2022 · 2022
Earlier work this paper cites.
Deep reinforcement learning: A survey
Xu Wang, Sen Wang, Xingxing Liang, Dawei Zhao, Jincai Huang, Xin Xu, Bin Dai, and Qiguang Miao. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
AlphaHoldem: High-performance artificial intelligence for heads-up no-limit poker via end-to-end reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 36. 4689–4697
Enmin Zhao, Renye Yan, Jinqiu Li, Kai Li, and Junliang Xing. 2022 · 2022
Earlier work this paper cites.
Lexi: Self-supervised learning of the ui language
Pratyay Banerjee, Shweti Mahajan, Kushal Arora, Chitta Baral, and Oriana Riva. 2023 · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023 · 2023
Earlier work this paper cites.
Adversarial attacks for black-box recommender systems via copying transferable cross-domain user profiles
Wenqi Fan, Xiangyu Zhao, Qing Li, Tyler Derr, Yao Ma, Hui Liu, Jianping Wang, and Jiliang Tang. 2023 · 2023
Earlier work this paper cites.
Multimodal web navigation with instruction-finetuned foundation models
Hiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo, Aleksandra Faust, Shixiang Shane Gu, and Izzeddin Gur. 2023a · 2023
Earlier work this paper cites.
Exposing limitations of language model agents in sequential-task compositions on the web
Hiroki Furuta, Yutaka Matsuo, Aleksandra Faust, and Izzeddin Gur. 2023b · 2023
Earlier work this paper cites.
Assistgui: Task-oriented desktop graphical user interface automation
Difei Gao, Lei Ji, Zechen Bai, Mingyu Ouyang, Peiran Li, Dongxing Mao, Qinchen Wu, Weichen Zhang, Peiyi Wang, Xiangwu Guo, et al · 2023
Earlier work this paper cites.
Intelligent virtual assistants with llm-based process automation
Yanchu Guan, Dong Wang, Zhixuan Chu, Shiyu Wang, Feiyue Ni, Ruihua Song, Longfei Li, Jinjie Gu, and Chenyi Zhuang. 2023 · 2023
Earlier work this paper cites.
A real-world webagent with planning, long context understanding, and program synthesis
Izzeddin Gur, Hiroki Furuta, Austin Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust. 2023 · 2023
Earlier work this paper cites.
Iluvui: Instruction-tuned language-vision modeling of uis from machine conversations
Yue Jiang, Eldon Schoop, Amanda Swearngin, and Jeffrey Nichols. 2023 · 2023
Earlier work this paper cites.
Language models can solve computer tasks
Geunwoo Kim, Pierre Baldi, and Stephen McAleer. 2023 · 2023
Earlier work this paper cites.
Pix2struct: Screenshot parsing as pretraining for visual language understanding. In ICML . PMLR, 18893–18912
Kenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu, Fangyu Liu, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. 2023 · 2023
Earlier work this paper cites.
A Zero-Shot Language Agent for Computer Control with Structured Reflection. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 11261–11274
Tao Li, Gang Li, Zhiwei Deng, Bryan Wang, and Yang Li. 2023b · 2023
Earlier work this paper cites.
UINav: A practical approach to train on-device automation agents
Wei Li, Fu-Lin Hsu, Will Bishop, Folawiyo Campbell-Ajala, Max Lin, and Oriana Riva. 2023a · 2023
Earlier work this paper cites.
Hierarchical prompting assists large language model on web navigation. In Findings of the Association for Computational Linguistics: EMNLP 2023
Robert Lo, Abishek Sridhar, Frank F Xu, Hao Zhu, and Shuyan Zhou. 2023 · 2023
Earlier work this paper cites.
LASER: LLM Agent with State-Space Exploration for Web Navigation. In NeurIPS 2023 Foundation Models for Decision Making Workshop
Kaixin Ma, Hongming Zhang, Hongwei Wang, Xiaoman Pan, and Dong Yu. [n. d.] · 2023
Earlier work this paper cites.
Android in the wild: A large-scale dataset for android device control
Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy Lillicrap. 2023 · 2023
Earlier work this paper cites.
Enhancing trust in llm-based ai automation agents: New considerations and future challenges
Sivan Schwartz, Avi Yaeli, and Segev Shlomov. 2023 · 2023
Earlier work this paper cites.
Step: Stacked llm policies for web actions
Paloma Sodhi, SRK Branavan, Yoav Artzi, and Ryan McDonald. 2023 · 2023
Earlier work this paper cites.
Webwise: Web interface control and sequential exploration with large language models
Heyi Tao, Sethuraman TV, Michal Shlapentokh-Rothman, and Derek Hoiem. 2023 · 2023
Earlier work this paper cites.
Empowering llm to use smartphone for intelligent task automation
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2023 · 2023
Earlier work this paper cites.
Autogen: Enabling next-gen llm applications via multi-agent conversation
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, et al · 2023
Earlier work this paper cites.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao. 2023b · 2023
Earlier work this paper cites.
The dawn of lmms: Preliminary explorations with gpt-4v (ision)
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. 2023a · 2023
Earlier work this paper cites.
Reinforced ui instruction grounding: Towards a generic ui task automation api
Zhizheng Zhang, Wenxuan Xie, Xiaoyi Zhang, and Yan Lu. 2023 · 2023
Earlier work this paper cites.
Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control. In NeurIPS 2023 Foundation Models for Decision Making Workshop
Longtao Zheng, Rundong Wang, Xinrun Wang, and Bo An. [n. d.] · 2023
Earlier work this paper cites.
Webarena: A realistic web environment for building autonomous agents
Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al · 2023
Earlier work this paper cites.
Cooperation, competition, and maliciousness: LLM-stakeholders interactive negotiation
Sahar Abdelnabi, Amr Gomaa, Sarath Sivaprasad, Lea Schönherr, and Mario Fritz. 2024 · 2024
Cited alongside, same era.
Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems. In NeurIPS 2024 Workshop on Open-World Agents
Tamer Abuelsaad, Deepak Akkil, Prasenjit Dey, Ashish Jagmohan, Aditya Vempaty, and Ravi Kokku. [n. d.] · 2024
Cited alongside, same era.
Dynamic fairness-aware recommendation through multi-agent social choice
Amanda Aird, Paresha Farastu, Joshua Sun, Elena Stefancová, Cassidy All, Amy Voida, Nicholas Mattei, and Robin Burke. 2024 · 2024
Cited alongside, same era.
ScreenAI: A vision-language model for ui and infographics understanding
Gilles Baechler, Srinivas Sunkara, Maria Wang, Fedir Zubach, Hassan Mansoor, Vincent Etter, Victor Cărbune, Jason Lin, Jindong Chen, and Abhanshu Sharma. 2024 · 2024
Cited alongside, same era.
Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks
Falcon-UI: Understanding GUI Before Following User Instructions
Huawen Shen, Chang Liu, Gengluo Li, Xinlong Wang, Yu Zhou, Can Ma, and Xiangyang Ji. 2024b · 2024
Later among the works it cites.
ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data
Junhong Shen, Atishay Jain, Zedian Xiao, Ishan Amlekar, Mouad Hadji, Aaron Podolny, and Ameet Talwalkar. 2024a · 2024
Later among the works it cites.
From grounding to planning: Benchmarking bottlenecks in web agents
Segev Shlomov, Aviad Sela, Ido Levy, Liane Galanti, Roy Abitbol, et al · 2024
Later among the works it cites.
Beyond Browsing: API-Based Web Agents
Yueqi Song, Frank Xu, Shuyan Zhou, and Graham Neubig. 2024b · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault de Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. 2024 · 2024
Cited alongside, same era.
Tell Me What’s Next: Textual Foresight for Generic UI Representations
Andrea Burns, Kate Saenko, and Bryan A Plummer. 2024 · 2024
Cited alongside, same era.
Large language models empowered personalized web agents
Hongru Cai, Yongqi Li, Wenjie Wang, Fengbin Zhu, Xiaoyu Shen, Wenjie Li, and Tat-Seng Chua. 2024 · 2024
Cited alongside, same era.
Web agents with world models: Learning and leveraging environment dynamics in web navigation. In The Thirteenth International Conference on Learning Representations
Hyungjoo Chae, Namyoung Kim, Kai Tzu-iunn Ong, Minju Gwak, Gwanwoo Song, Jihoon Kim, Sunghwan Kim, Dongha Lee, and Jinyoung Yeo. 2024 · 2024
Cited alongside, same era.
CoMM: Collaborative multi-agent, multi-reasoning-path prompting for complex problem solving
Pei Chen, Boran Han, and Shuai Zhang. 2024a · 2024
Cited alongside, same era.
EDGE: Enhanced grounded gui understanding with enriched multi-granularity synthetic data
Xuetian Chen, Hangcheng Li, Jiaqing Liang, Sihang Jiang, and Deqing Yang. 2024b · 2024
Cited alongside, same era.
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 9313–9332
Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Li YanTao, Jianbing Zhang, and Zhiyong Wu. 2024 · 2024
Cited alongside, same era.
Caap: Context-aware action planning prompting to solve computer tasks with front-end ui only
Junhee Cho, Jihoon Kim, Daseul Bae, Jinho Choo, Youngjune Gwon, and Yeong-Dae Kwon. 2024 · 2024
Cited alongside, same era.
Zirui Song, Yaohang Li, Meng Fang, Zhenhao Chen, Zecheng Shi, Yuan Huang, and Ling Chen. 2024a · 2024
Later among the works it cites.
Cradle: Empowering Foundation Agents towards General Computer Control. In NeurIPS 2024 Workshop on Open-World Agents
Weihao Tan, Wentao Zhang, Xinrun Xu, Haochong Xia, Gang Ding, Boyu Li, Bohan Zhou, Junpeng Yue, Jiechuan Jiang, Yewen Li, et al · 2024
Later among the works it cites.
Steward: Natural language web automation
Brian Tang and Kang G Shin. 2024 · 2024
Later among the works it cites.
Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment
Hao Tang, Darren Key, and Kevin Ellis. 2024 · 2024
Later among the works it cites.
Navigating WebAI: Training agents to complete web tasks with large language models and reinforcement learning. In Proceedings of the 39th ACM/SIGAPP Symposium on Applied Computing . 866–874
Lucas-Andrei Thil, Mirela Popa, and Gerasimos Spanakis. 2024 · 2024
Later among the works it cites.
ALI-Agent: Assessing LLMs’ Alignment with Human Values via Agent-based Evaluation
Han Wang, An Zhang, Nguyen Duy Tai, Jun Sun, Tat-Seng Chua, et al · 2024
Later among the works it cites.
Multi-agent attacks for black-box social recommendations
Shijie Wang, Wenqi Fan, Xiao-Yong Wei, Xiaowei Mei, Shanru Lin, and Qing Li. 2024a · 2024
Later among the works it cites.
Oscar: Operating system control via state-aware reasoning and re-planning
Xiaoqiang Wang and Bang Liu. 2024 · 2024
Later among the works it cites.
Ponder & press: Advancing visual gui agent towards general computer control
Yiqin Wang, Haoji Zhang, Jingqi Tian, and Yansong Tang. 2024d · 2024
Later among the works it cites.
Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig. 2024b · 2024
Later among the works it cites.
Wipi: A new web threat for llm-driven web agents
Fangzhou Wu, Shutong Wu, Yulong Cao, and Chaowei Xiao. 2024b · 2024
Later among the works it cites.
A survey on large language models for recommendation
Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, et al · 2024
Later among the works it cites.
MobileVLM: A vision-language model for better intra-and inter-ui understanding
Qinzhuo Wu, Weikai Xu, Wei Liu, Tao Tan, Jianfeng Liu, Ang Li, Jian Luan, Bin Wang, and Shuo Shang. 2024d · 2024
Later among the works it cites.
OS-Copilot: Towards Generalist Computer Agents with Self-Improvement. In ICLR 2024 Workshop on Large Language Model (LLM) Agents
Zhiyong Wu, Chengcheng Han, Zichen Ding, Zhenmin Weng, Zhoumianze Liu, Shunyu Yao, Tao Yu, and Lingpeng Kong. 2024a · 2024
Later among the works it cites.
OS-Atlas: A foundation action model for generalist GUI agents
Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, et al · 2024
Later among the works it cites.
Advweb: Controllable black-box attacks on vlm-powered web agents
Chejian Xu, Mintong Kang, Jiawei Zhang, Zeyi Liao, Lingbo Mo, Mengqi Yuan, Huan Sun, and Bo Li. 2024a · 2024
Later among the works it cites.
Tur [k] ingbench: A challenge benchmark for web agents
Kevin Xu, Yeganeh Kordi, Tanay Nayak, Ado Asija, Yizhong Wang, Kate Sanders, Adam Byerly, Jingyu Zhang, Benjamin Van Durme, and Daniel Khashabi. 2024b · 2024
Later among the works it cites.
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
Yiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang, Zekun Wang, Yuchen Mao, Caiming Xiong, and Tao Yu. 2024c · 2024
Later among the works it cites.
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu, and Caiming Xiong. 2024d · 2024
Later among the works it cites.
Agentoccam: A simple yet strong baseline for llm-based web agents
Ke Yang, Yao Liu, Sapana Chaudhary, Rasool Fakoor, Pratik Chaudhari, George Karypis, and Huzefa Rangwala. 2024a · 2024
Later among the works it cites.
R-judge: Benchmarking safety risk awareness for llm agents
Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, et al · 2024
Later among the works it cites.
Ufo: A ui-focused agent for windows os interaction
Chaoyun Zhang, Liqun Li, Shilin He, Xu Zhang, Bo Qiao, Si Qin, Minghua Ma, Yu Kang, Qingwei Lin, Saravan Rajmohan, et al · 2024
Later among the works it cites.
Mm1. 5: Methods, analysis & insights from multimodal llm fine-tuning
Haotian Zhang, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, Bowen Zhang, Yanghao Li, et al · 2024
Later among the works it cites.
Android in the zoo: Chain-of-action-thought for gui agents
Jiwen Zhang, Jihao Wu, Yihua Teng, Minghui Liao, Nuo Xu, Xiao Xiao, Zhongyu Wei, and Duyu Tang. 2024d · 2024
Later among the works it cites.
Ui-hawk: Unleashing the screen stream understanding for gui agents
Jiwen Zhang, Yaqi Yu, Minghui Liao, Wentao Li, Jihao Wu, and Zhongyu Wei. 2024f · 2024
Later among the works it cites.
Dynamic Planning for LLM-based Graphical User Interface Automation. In Findings of the Association for Computational Linguistics: EMNLP 2024 . 1304–1320
Shaoqing Zhang, Zhuosheng Zhang, Kehai Chen, Xinbei Ma, Muyun Yang, Tiejun Zhao, and Min Zhang. 2024g · 2024
Later among the works it cites.
Privacyasst: Safeguarding user privacy in tool-using large language model agents
Xinyu Zhang, Huiyu Xu, Zhongjie Ba, Zhibo Wang, Yuan Hong, Jian Liu, Zhan Qin, and Kui Ren. 2024e · 2024
Later among the works it cites.
Yao Zhang, Zijian Ma, Yunpu Ma, Zhen Han, Yu Wu, and Volker Tresp. 2024c · 2024
Later among the works it cites.
You Only Look at Screens: Multimodal Chain-of-Action Agents. In Findings of the Association for Computational Linguistics ACL 2024 . 3132–3149
Zhuosheng Zhang and Aston Zhang. 2024a · 2024
Later among the works it cites.
You Only Look at Screens: Multimodal Chain-of-Action Agents. In Findings of the Association for Computational Linguistics ACL 2024 . 3132–3149
Zhuosheng Zhang and Aston Zhang. 2024b · 2024
Later among the works it cites.
Recommender systems in the era of large language models (llms)
Zihuai Zhao, Wenqi Fan, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Zhen Wen, Fei Wang, Xiangyu Zhao, Jiliang Tang, et al · 2024
Later among the works it cites.
GPT-4V (ision) is a Generalist Web Agent, if Grounded. In International Conference on Machine Learning . PMLR, 61349–61385
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. 2024 · 2024
Later among the works it cites.
The Design and application of RAG-based conversational agents for collaborative problem solving. In Proceedings of the 2024 9th International Conference on Distance Education and Learning . 62–68
Xuanyan Zhong, Haiyang Xin, Wenfeng Li, Zehui Zhan, and May-hung Cheng. 2024b · 2024
Later among the works it cites.
Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
Hao Bai, Yifei Zhou, Jiayi Pan, Mert Cemri, Alane Suhr, Sergey Levine, and Aviral Kumar. 2025 · 2025
Closest in time.
DoomArena: A framework for Testing AI Agents Against Evolving Security Threats. In ICML 2025 Workshop on Computer Use Agents
Léo Boisvert, Abhay Puri, Gabriel Huang, Mihir Bansal, Chandra Kiran Reddy Evuru, Avinandan Bose, Maryam Fazel, Quentin Cappart, Alexandre Lacoste, Alexandre Drouin, et al · 2025
Closest in time.
PLAN-AND-ACT: Improving Planning of Agents for Long-Horizon Tasks
Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami. 2025 · 2025
Closest in time.
Computational Protein Science in the Era of Large Language Models (LLMs)
Wenqi Fan, Yi Zhou, Shijie Wang, Yuyao Yan, Hui Liu, Qian Zhao, Le Song, and Qing Li. 2025 · 2025
Closest in time.
AgentRefine: Enhancing Agent Generalization through Refinement Tuning
Dayuan Fu, Keqing He, Yejie Wang, Wentao Hong, Zhuoma Gongque, Weihao Zeng, Wei Wang, Jingang Wang, Xunliang Cai, and Weiran Xu. 2025 · 2025
Closest in time.
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
Chengquan Guo, Xun Liu, Chulin Xie, Andy Zhou, Yi Zeng, Zinan Lin, Dawn Song, and Bo Li. 2025 · 2025
Closest in time.
From RAG to Memory: Non-Parametric Continual Learning for Large Language Models
Bernal Jiménez Gutiérrez, Yiheng Shu, Weijian Qi, Sizhe Zhou, and Yu Su. 2025 · 2025
Closest in time.
Jiani Huang, Shijie Wang, Liang-bo Ning, Wenqi Fan, Shuaiqiang Wang, Dawei Yin, and Qing Li. 2025b · 2025
Closest in time.
R2D2: Remembering, Reflecting and Dynamic Decision Making for Web Agents
Tenghao Huang, Kinjal Basu, Ibrahim Abdelaziz, Pavan Kapanipathi, Jonathan May, and Muhao Chen. 2025a · 2025
Closest in time.
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
Faria Huq, Zora Zhiruo Wang, Frank F Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P Bigham, and Graham Neubig. 2025 · 2025
Closest in time.
Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents
Dongjun Lee, Juyong Lee, Kyuyoung Kim, Jihoon Tack, Jinwoo Shin, Yee Whye Teh, and Kimin Lee. 2025 · 2025
Closest in time.
On the effects of data scale on ui control agents
Wei Li, William E Bishop, Alice Li, Christopher Rawles, Folawiyo Campbell-Ajala, Divya Tyamagundlu, and Oriana Riva. 2025a · 2025
Closest in time.
GenRL: Multimodal-foundation world models for generalization in embodied agents
Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Aaron C Courville, and Sai Rajeswar Mudumba. 2025 · 2025
Closest in time.
From Documents to Dialogue: Building KG-RAG Enhanced AI Assistants
Manisha Mukherjee, Sungchul Kim, Xiang Chen, Dan Luo, Tong Yu, and Tung Mai. 2025 · 2025
Closest in time.
Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey
Bo Ni, Zheyuan Liu, Leyao Wang, Yongjia Lei, Yuying Zhao, Xueqi Cheng, Qingkai Zeng, Luna Dong, Yinglong Xia, Krishnaram Kenthapadi, et al · 2025
Closest in time.
Synatra: Turning indirect knowledge into direct demonstrations for digital agents at scale
Tianyue Ou, Frank F Xu, Aman Madaan, Jiarui Liu, Robert Lo, Abishek Sridhar, Sudipta Sengupta, Dan Roth, Graham Neubig, and Shuyan Zhou. 2025 · 2025
Closest in time.
Manish Sanwal. 2025 · 2025
Closest in time.
Privacylens: Evaluating privacy norm awareness of language models in action
Yijia Shao, Tianshi Li, Weiyan Shi, Yanchen Liu, and Diyi Yang. 2025b · 2025
Closest in time.
Unveiling Privacy Risks in LLM Agent Memory
Bo Wang, Weiyi He, Pengfei He, Shenglai Zeng, Zhen Xiang, Yue Xing, and Jiliang Tang. 2025b · 2025
Closest in time.
Knowledge Graph Retrieval-Augmented Generation for LLM-based Recommendation
Shijie Wang, Wenqi Fan, Yue Feng, Xinyu Ma, Shuaiqiang Wang, and Dawei Yin. 2025a · 2025
Closest in time.
Dissecting Adversarial Robustness of Multimodal LM Agents. In ICLR
Chen Henry Wu, Rishi Rajesh Shah, Jing Yu Koh, Russ Salakhutdinov, Daniel Fried, and Aditi Raghunathan. 2025 · 2025
Closest in time.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al · 2025
Closest in time.
Watch out for your agents! investigating backdoor threats to llm-based agents
Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, and Xu Sun. 2025 · 2025
Closest in time.
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
Peter Yong Zhong, Siyuan Chen, Ruiqi Wang, McKenna McCall, Ben L Titzer, and Heather Miller. 2025 · 2025
Closest in time.