Fetching the paper…
Reading the bibliography…
Agents for computer use (ACUs) are an emerging class of systems capable of executing complex tasks on digital devices -- such as desktops, mobile phones, and web platforms -- given instructions in natural language.
Language Models are Few-Shot Learners. In Proc. of the 33rd Int. Conf. on NeurIPS , Vol. 33. Curran Associates, Inc., Vancouver, Canada, 1877–1901
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Fine-Tuning Language Models from Human Preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2020 · 1909
Earlier work this paper cites.
ALVINN: An autonomous land vehicle in a neural network. In Proc. of the 2nd Int. Conf. on NeurIPS , Vol. 1. Morgan Kaufmann, Denver, CO, USA
Dean A. Pomerleau. 1988 · 1988
Earlier work this paper cites.
Human error
James Reason. 1990 · 1990
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments. In 1990 IJCNN international joint conference on neural networks . IEEE, IEEE, San Diego, CA, USA, 253–258
Jürgen Schmidhuber. 1990 · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S. Sutton. 1991 · 1991
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping. In Proc. of the 16th ICML . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 278–287
Andrew Y Ng, Daishi Harada, and Stuart Russell. 1999 · 1999
Earlier work this paper cites.
Metrics for multi-class classification: an overview
Margherita Grandini, Enrico Bagli, and Giorgio Visani. 2020 · 2008
Earlier work this paper cites.
Curriculum learning. In Proc. of the 26th ICML . PMLR, Montreal, QC, Canada, 41–48
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
Reinforcement Learning for Mapping Instructions to Actions. In Proc. of the Joint Conf. of the 47th Annual Meeting of the ACL and the 4th IJCNLP . ACL, Suntec, Singapore, 82–90
S.R.K. Branavan, Harr Chen, Luke Zettlemoyer, and Regina Barzilay. 2009 · 2009
Earlier work this paper cites.
Continuous delivery: reliable software releases through build, test, and deployment automation (7th edition ed.)
Jez Humble and David Farley. 2011 · 2011
Earlier work this paper cites.
Playing Atari with Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Earlier work this paper cites.
Deep reinforcement learning: A brief survey
Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage, and Anil Anthony Bharath. 2017 · 2017
Earlier work this paper cites.
SUGILITE: Creating multimodal smartphone automation by demonstration. In Proc. of the Conf. on CHI . ACM, Denver, CO, USA, 6038–6049
Toby Jia-Jun Li, Amos Azaria, and Brad A. Myers. 2017 · 2017
Earlier work this paper cites.
World of Bits: An open-domain platform for web-based agents. In Proc. of the 34th ICML . PMLR, Sydney, NSW, Australia, 3135–3144
Tianlin Shi, Andrej Karpathy, Linxi Fan, Jonathan Hernandez, and Percy Liang. 2017 · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. 2017 · 2017
Earlier work this paper cites.
Recurrent World Models Facilitate Policy Evolution. In Proc. of the 32st Int. Conf. on NeurIPS , Vol. 31. Curran Associates, Inc., Montréal, Quebec, Canada
David Ha and Jürgen Schmidhuber. 2018 · 2018
Earlier work this paper cites.
QBE: QLearning-based exploration of Android applications. In Proc. of the 11th ICST . IEEE, New York, NY, USA, 105–115
Yavuz Koroglu, Alper Sen, Ozlem Muslu, Yunus Mete, Ceyda Ulker, Tolga Tanriverdi, and Yunus Donmez. 2018 · 2018
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration. In Proc. of the 6th ICLR . OpenReview.net, Vancouver, BC, Canada
E. Z. Liu, K. Guu, P. Pasupat, T. Shi, and P. Liang. 2018 · 2018
Earlier work this paper cites.
Mapping natural language commands to web elements. In Proc. of the Conf. on EMNLP . ACL, Brussels, Belgium, 4970–4976
Panupong Pasupat, Tian-Shun Jiang, Evan Liu, Kelvin Guu, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction (second edition ed.)
Richard S. Sutton and Andrew G. Barto. 2018 · 2018
Earlier work this paper cites.
Learning user interface element interactions. In Proc. of the 28th ACM SIGSOFT Int. Symposium on Software Testing and Analysis . ACM, Beijing China, 296–306
Christian Degott, Nataniel P. Borges Jr., and Andreas Zeller. 2019 · 2019
Earlier work this paper cites.
Learning to navigate the web. In Proc. of the 7th ICLR . OpenReview.net, New Orleans, LA, USA
Izzeddin Gur, Ulrich Rückert, Aleksandra Faust, and Dilek Hakkani-Tür. 2019 · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proc. of the IEEE/CVF Conf. on CVPR . IEEE, Long Beach, CA, USA, 6700–6709
Drew A Hudson and Christopher D Manning. 2019 · 2019
Earlier work this paper cites.
DOM-Q-NET: Grounded RL on structured language. In Proc. of the 7th ICLR . OpenReview.net, New Orleans, LA, USA
S. Jia, J. Kiros, and J. Ba. 2019 · 2019
Earlier work this paper cites.
DeepEE: Joint optimization of job scheduling and cooling control for data center energy efficiency using deep reinforcement learning. In Proc. of the 39th ICDCS . IEEE, Dallas, TX, USA, 645–655
Yongyi Ran, Han Hu, Xin Zhou, and Yonggang Wen. 2019 · 2019
Earlier work this paper cites.
Robotic Process Automation: Contemporary Themes and Challenges
Rehan Syed, Suriadi Suriadi, Michael Adams, Wasana Bandara, Sander J. J. Leemans, Chun Ouyang, Arthur H. M. Ter Hofstede, Inge Van De Weerd, Moe Thandar Wynn, and Hajo A. Reijers. 2020 · 2019
Earlier work this paper cites.
From Robotic Process Automation to Intelligent Process Automation: Emerging Trends
Tathagata Chakraborti, Vatche Isahagian, Rania Khalaf, Yasaman Khazaeni, Vinod Muthusamy, Yara Rizk, and Merve Unuvar. 2020 · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020 · 2020
Earlier work this paper cites.
A Survey of Deep Learning Techniques for Autonomous Driving
Sorin Grigorescu, Bogdan Trasnea, Tiberiu Cocias, and Gigel Macesanu. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks. In Proc. of the 34th Int. Conf. on NeurIPS , Vol. 33. Curran Associates, Inc., virtual, 9459–9474
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Interactive task learning from GUI-grounded natural language instructions and demonstrations. In Proc. of the 58th Annual Meeting of the ACL: System Demonstrations . ACL, Online, 215–223
Toby Jia-Jun Li, Tom Mitchell, and Brad Myers. 2020b · 2020
Earlier work this paper cites.
Mapping natural language instructions to mobile UI action sequences. In Proc. of the 58th Annual Meeting of the ACL . ACL, Online, 8198–8210
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge. 2020a · 2020
Earlier work this paper cites.
Reinforcement learning based curiosity-driven testing of Android applications. In Proc. of the 29th ACM SIGSOFT Int. Symposium on Software Testing and Analysis . ACM, New York, NY, USA, 153–164
Minxue Pan, An Huang, Guoxin Wang, Tian Zhang, and Xuandong Li. 2020 · 2020
Earlier work this paper cites.
Data efficient reinforcement learning for legged robots. In Proc. of the Conf. on Robot Learning , Vol. 100. PMLR, Cambridge, MA, USA, 1–10
Yuxiang Yang, Ken Caluwaerts, Atil Iscen, Tingnan Zhang, Jie Tan, and Vikas Sindhwani. 2020 · 2020
Earlier work this paper cites.
WebSRC: A Dataset for Web-Based Structural Reading Comprehension. In Proc. of the Conf. on EMNLP . ACL, Punta Cana, Dominican Republic, 4173–4185
Xingyu Chen, Zihan Zhao, Lu Chen, JiaBao Ji, Danyang Zhang, Ao Luo, Yuxuan Xiong, and Kai Yu. 2021 · 2021
Earlier work this paper cites.
Environment Generation for Zero-Shot Compositional Reinforcement Learning. In Proc. of the 34th Int. Conf. on NeurIPS , Vol. 34. Curran Associates, Inc., virtual, 4157–4169
Izzeddin Gur, Natasha Jaques, Yingjie Miao, Jongwook Choi, Manoj Tiwari, Honglak Lee, and Aleksandra Faust. 2021 · 2021
Earlier work this paper cites.
ActionBert: Leveraging User Actions for Semantic Understanding of User Interfaces
Zecheng He, Srinivas Sunkara, Xiaoxue Zang, Ying Xu, Lijuan Liu, Nevan Wichers, Gabriel Schubiner, Ruby Lee, and Jindong Chen. 2021 · 2021
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Learning UI Navigation through Demonstrations composed of Macro Actions
Wei Li. 2021 · 2021
Earlier work this paper cites.
Glider: A reinforcement learning approach to extract UI scripts from websites. In Proc. of the 44th Int. ACM SIGIR Conf. on Research and Development in Information Retrieval . ACM, New York, NY, USA, 1420–1430
Yuanchun Li and Oriana Riva. 2021 · 2021
Earlier work this paper cites.
Automating web-based infrastructure management via contextual imitation learning. In Proc. of the 22nd Asia-Pacific Network Operations and Management Symposium . IEEE, Tainan, Taiwan, 184–189
Jieyu Lin, Hongxiang Geng, and Alberto Leon-Garcia. 2021 · 2021
Earlier work this paper cites.
FLIN: A Flexible Natural Language Interface for Web Navigation. In Proc. of the NAACL: Human Language Technologies (NAACL) . ACL, Online, 2777–2788
Sahisnu Mazumder and Oriana Riva. 2021 · 2021
Earlier work this paper cites.
AppBuddy: Learning to accomplish tasks in mobile apps via reinforcement learning
Maayan Shvo, Zhiming Hu, Rodrigo Toro Icarte, Iqbal Mohomed, Allan Jepson, and Sheila A. McIlraith. 2021 · 2021
Earlier work this paper cites.
AndroidEnv: A Reinforcement Learning Platform for Android
Daniel Toyama, Philippe Hamel, Anita Gergely, Gheorghe Comanici, Amelia Glaese, Zafarali Ahmed, Tyler Jackson, Shibl Mourad, and Doina Precup. 2021 · 2021
Earlier work this paper cites.
Grounding Open-Domain Instructions to Automate Web Support Tasks. In Proc. of the NAACL: Human Language Technologies (NAACL) . ACL, Online, 1022–1032
Nancy Xu, Sam Masling, Michael Du, Giovanni Campagna, Larry Heck, James Landay, and Monica Lam. 2021 · 2021
Earlier work this paper cites.
Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos. In Proc. of the 36th Int. Conf. on NeurIPS , Vol. 35. Curran Associates, Inc., New Orleans, LA, USA, 24639–24654
Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune. 2022 · 2022
Earlier work this paper cites.
A dataset for interactive vision-language navigation with unknown command feasibility
Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, and Bryan A. Plummer. 2022 · 2022
Earlier work this paper cites.
Optimal energy management for air cooled server fans using deep reinforcement learning control method
Yogesh Fulpagare, Kuei-Ru Huang, Ying-Hao Liao, and Chi-Chuan Wang. 2022 · 2022
Earlier work this paper cites.
A data-driven approach for learning to control computers. In Proc. of the 39th ICML . PMLR, Baltimore, Maryland, USA, 9466–9482
Peter C Humphreys, David Raposo, Tobias Pohlen, Gregory Thornton, Rachita Chhaparia, Alistair Muldal, Josh Abramson, Petko Georgiev, Adam Santoro, and Timothy Lillicrap. 2022 · 2022
Earlier work this paper cites.
Do BERTs learn to use browser user interface? Exploring multi-step tasks with unified vision-and-language BERTs
Taichi Iki and Akiko Aizawa. 2022 · 2022
Cited alongside, same era.
A Path Towards Autonomous Machine Intelligence
Yann LeCun. 2022 · 2022
Cited alongside, same era.
WebGPT: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback. In Proc. of the 36th Int. Conf. on NeurIPS , Vol. 35. Curran Associates, Inc., New Orleans, LA, USA, 27730–27744
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Cited alongside, same era.
MobileAgent: Enhancing mobile control via human-machine interaction and SOP integration
Tinghe Ding. 2024 · 2024
Later among the works it cites.
Training a Vision Language Model as Smartphone Assistant
Nicolai Dorka, Janusz Marecki, and Ammar Anwar. 2024 · 2024
Later among the works it cites.
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?. In Proc. of the 41st ICML . PMLR, Vienna, Austria, 11642–11662
Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste. 2024 · 2024
Later among the works it cites.
Search beyond queries: Training smaller language models for web interactions via reinforcement learning
Moghis Fereidouni and A. B. Siddique. 2024 · 2024
Later among the works it cites.
Multimodal web navigation with instruction-finetuned foundation models. In Proc. of the 12th ICLR . OpenReview.net, Singapore
H. Furuta, K.-H. Lee, O. Nachum, Y. Matsuo, A. Faust, S. S. Gu, and I. Gur. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stuart J. Russell and Peter Norvig. 2022 · 2022
Cited alongside, same era.
META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI. In Proc. of the Conf. on EMNLP . ACL, Abu Dhabi, United Arab Emirates, 6699–6712
Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, and Kai Yu. 2022 · 2022
Cited alongside, same era.
Emergent Abilities of Large Language Models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022 · 2022
Cited alongside, same era.
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents. In Proc. of the 36th Int. Conf. on NeurIPS , Vol. 35. Curran Associates, Inc., New Orleans, LA, USA, 20744–20757
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022 · 2022
Cited alongside, same era.
Learning to Navigate Wikipedia by Taking Random Walks. In Proc. of the 35th Int. Conf. on NeurIPS , Vol. 35. Curran Associates, Inc., New Orleans, LA, USA, 1529–1541
Manzil Zaheer, Kenneth Marino, Will Grathwohl, John Schultz, Wendy Shang, Sheila Babayan, Arun Ahuja, Ishita Dasgupta, Christine Kaeser-Chen, and Rob Fergus. 2022 · 2022
Cited alongside, same era.
The unsolved challenges of LLMs as generalist web agents: A case study. In Proc. of the 37th Int. Conf. on NeurIPS: Foundation Models for Decision Making Workshop . Curran Associates, Inc., New Orleans, LA, USA
Rim Assouel, Tom Marty, Massimo Caccia, Issam H. Laradji, Alexandre Drouin, Sai Rajeswar, Hector Palacios, Quentin Cappart, David Vazquez, Nicolas Chapados, Maxime Gasse, and Alexandre Lacoste. 2023 · 2023
Cited alongside, same era.
Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. 2023 · 2023
Cited alongside, same era.
Mind2Web: Towards a generalist agent for the web. In Proc. of the 37th Int. Conf. on NeurIPS . Curran Associates, Inc., New Orleans, LA, USA, 28091–28114
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
PPTC benchmark: Evaluating large language models for PowerPoint task completion. In Findings of the ACL . ACL, Bangkok, Thailand, 8682–8701
Yiduo Guo, Zekai Zhang, Yaobo Liang, Dongyan Zhao, and Nan Duan. 2024b · 2024
Later among the works it cites.
StableToolBench: Towards stable large-scale benchmarking on tool learning of large language models. In Findings of the ACL . ACL, Bangkok, Thailand, 11143–11156
Zhicheng Guo, Sijie Cheng, Hao Wang, Shihao Liang, Yujia Qin, Peng Li, Zhiyuan Liu, Maosong Sun, and Yang Liu. 2024a · 2024
Later among the works it cites.
A real-world webagent with planning, long context understanding, and program synthesis. In Proc. of the 12th ICLR . OpenReview.net, Singapore
I. Gur, H. Furuta, A. V. Huang, M. Safdari, Y. Matsuo, D. Eck, and A. Faust. 2024 · 2024
Later among the works it cites.
WebVoyager: Building an end-to-end web agent with large multimodal models. In Proc. of the 62nd Annual Meeting of the ACL . ACL, Bangkok, Thailand, 6864–6890
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024 · 2024
Later among the works it cites.
CogAgent: A visual language model for GUI agents. In Proc. of the IEEE/CVF Conf. on CVPR . IEEE, Seattle, WA, USA, 14281–14290
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, and Jie Tang. 2024 · 2024
Later among the works it cites.
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
Siyuan Hu, Mingyu Ouyang, Difei Gao, and Mike Zheng Shou. 2024 · 2024
Later among the works it cites.
Can Large Language Models Reason and Plan?
Subbarao Kambhampati. 2024 · 2024
Later among the works it cites.
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web. In Proceedings of the ECCV . Springer-Verlag, Milan, Italy, 161–178
Raghav Kapoor, Yash Parag Butala, Melisa Russak, Jing Yu Koh, Kiran Kamble, Waseem Alshikh, and Ruslan Salakhutdinov. 2024 · 2024
Later among the works it cites.
Dual-view visual contextualization for web navigation. In Proc. of the IEEE/CVF Conf. on CVPR . IEEE, Seattle WA, USA, 14445–14454
Jihyung Kil, Chan Hee Song, Boyuan Zheng, Xiang Deng, Yu Su, and Wei-Lun Chao. 2024 · 2024
Later among the works it cites.
VisualWebArena: Evaluating multimodal agents on realistic visual web tasks. In Proc. of the 62nd Annual Meeting of the ACL . ACL, Bangkok, Thailand, 881–905
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. 2024a · 2024
Later among the works it cites.
TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Systems. In Proc. of the Conf. on EMNLP: Industry Track . ACL, Singapore, 371–385
Yilun Kong, Jingqing Ruan, Yihong Chen, Bin Zhang, Tianpeng Bao, Shiwei Shi, Guoqing Du, Xiaoru Hu, Hangyu Mao, Ziyue Li, Xingyu Zeng, and Rui Zhao. 2023 · 2024
Later among the works it cites.
AutoWebGLM: Bootstrap and reinforce a large language model-based web navigating agent
Hanyu Lai, Xiao Liu, Iat Long Iong, Shuntian Yao, Yuxuan Chen, Pengbo Shen, Hao Yu, Hanchen Zhang, Xiaohan Zhang, Yuxiao Dong, and Jie Tang. 2024 · 2024
Later among the works it cites.
Benchmarking Mobile Device Control Agents across Diverse Configurations
Juyong Lee, Taywon Min, Minyong An, Dongyoon Hahm, Haeone Lee, Changyeon Kim, and Kimin Lee. 2024 · 2024
Later among the works it cites.
MUG: Interactive multimodal grounding on user interfaces. In Findings of the ACL: EACL 2024 . ACL, St. Julian’s, Malta, 231–251
T. Li, G. Li, J. Zheng, P. Wang, and Y. Li. 2024d · 2024
Later among the works it cites.
UINav: A Practical Approach to Train On-Device Automation Agents. In Proc. of the NAACL: Human Language Technologies (NAACL, Vol. 6) . ACL, Rochester, New York, USA, 36–51
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen. 2024b · 2024
Later among the works it cites.
AgentBench: Evaluating LLMs as Agents. In Proc. of the 12th ICLR . OpenReview.net, Singapore
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. 2024 · 2024
Later among the works it cites.
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue. In Proc. of the 41st ICML . PMLR, Vienna, Austria, 33007–33056
Xing Han Lù, Zdeněk Kasner, and Siva Reddy. 2024 · 2024
Later among the works it cites.
WILBUR: Adaptive in-context learning for robust and accurate web agents
Michael Lutz, Arth Bohra, Manvel Saroyan, Artem Harutyunyan, and Giovanni Campagna. 2024 · 2024
Later among the works it cites.
CoCo-Agent: A comprehensive cognitive MLLM agent for smartphone GUI automation. In Findings of the ACL . ACL, Bangkok, Thailand, 9097–9110
Xinbei Ma, Zhuosheng Zhang, and Hai Zhao. 2024a · 2024
Later among the works it cites.
BAGEL: Bootstrapping agents by guiding exploration with language
Shikhar Murty, Christopher Manning, Peter Shaw, Mandar Joshi, and Kenton Lee. 2024 · 2024
Later among the works it cites.
Privacy Issues in Large Language Models: A Survey
Seth Neel and Peter Chang. 2024 · 2024
Later among the works it cites.
ScreenAgent: A Vision Language Model-driven Computer Control Agent. In Proc. of the 33rd IJCAI . IJCAI, Jeju, Korea, 6433–6441
Runliang Niu, Jindong Li, Shiqi Wang, Yali Fu, Xiyu Hu, Xueyuan Leng, He Kong, Yi Chang, and Qi Wang. 2024 · 2024
Later among the works it cites.
MobileFlow: A Multimodal LLM For Mobile GUI Agent
Songqin Nong, Jiali Zhu, Rui Wu, Jiongchao Jin, Shuo Shan, Xiutian Huang, and Wenhao Xu. 2024 · 2024
Later among the works it cites.
GPT-4 Technical Report
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, C. J. Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. 2024 · 2024
Later among the works it cites.
Autonomous evaluation and refinement of digital agents
Jiayi Pan, Yichi Zhang, Nicholas Tomlin, Yifei Zhou, Sergey Levine, and Alane Suhr. 2024 · 2024
Later among the works it cites.
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
Pranav Putta, Edmund Mills, Naman Garg, Sumeet Motwani, Chelsea Finn, Divyansh Garg, and Rafael Rafailov. 2024 · 2024
Later among the works it cites.
ChatDev: Communicative agents for software development. In Proc. of the 62nd Annual Meeting of the ACL . ACL, Bangkok, Thailand, 15174–15186
Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2024 · 2024
Later among the works it cites.
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. In Proc. of the 12th ICLR . OpenReview.net, Singapore
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2024 · 2024
Later among the works it cites.
V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM
Abdur Rahman, Rajat Chawla, Muskaan Kumar, Arkajit Datta, Adarsh Jha, Mukunda NS, and Ishaan Bhola. 2024 · 2024
Later among the works it cites.
AndroidWorld: A dynamic benchmarking environment for autonomous agents
Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William Bishop, Wei Li, Folawiyo Campbell-Ajala, Daniel Toyama, Robert Berry, Divya Tyamagundlu, Timothy Lillicrap, and Oriana Riva. 2024 · 2024
Later among the works it cites.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024 · 2024
Later among the works it cites.
Towards General Computer Control: A Multimodal Agent for Red Dead Redemption II as a Case Study
Weihao Tan, Ziluo Ding, Wentao Zhang, Boyu Li, Bohan Zhou, Junpeng Yue, Haochong Xia, Jiechuan Jiang, Longtao Zheng, Xinrun Xu, Yifei Bi, Pengjie Gu, Xinrun Wang, Börje F. Karlsson, Bo An, and Zongqing Lu. 2024 · 2024
Later among the works it cites.
WebWISE: Web Interface Control and Sequential Exploration with Large Language Models. In Proc. of the NAACL: Human Language Technologies . ACL, Mexico City, Mexico, 3693–3711
Heyi Tao, Sethuraman T V, Michal Shlapentokh-Rothman, and Derek Hoiem. 2024 · 2024
Later among the works it cites.
So you want your private LLM at home? A survey and benchmark of methods for efficient GPTs. In Proc. of the 11th SDS . IEEE, Zurich, Switzerland, 205–212
Lukas Tuggener, Pascal Sager, Yassine Taoudi-Benchekroun, Benjamin F. Grewe, and Thilo Stadelmann. 2024 · 2024
Later among the works it cites.
LLMs Still Can’t Plan; Can LRMs? A Preliminary Evaluation of OpenAI’s o1 on PlanBench
Karthik Valmeekam, Kaya Stechly, and Subbarao Kambhampati. 2024 · 2024
Later among the works it cites.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen. 2024c · 2024
Later among the works it cites.
MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding
Qinzhuo Wu, Weikai Xu, Wei Liu, Tao Tan, Jianfeng Liu, Ang Li, Jian Luan, Bin Wang, and Shuo Shang. 2024c · 2024
Later among the works it cites.
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. 2024 · 2024
Later among the works it cites.
Android in the Zoo: Chain-of-Action-Thought for GUI Agents. In Findings of the ACL: EMNLP 2024 . ACL, Miami, Florida, USA, 12016–12031
Jiwen Zhang, Jihao Wu, Yihua Teng, Minghui Liao, Nuo Xu, Xiao Xiao, Zhongyu Wei, and Duyu Tang. 2024e · 2024
Later among the works it cites.
You only look at screens: Multimodal chain-of-action agents. In Findings of the ACL . ACL, Bangkok, Thailand, 3132–3149
Zhuosheng Zhang and Aston Zhang. 2024 · 2024
Later among the works it cites.
WebArena: A Realistic Web Environment for Building Autonomous Agents. In Proc. of the 12th ICLR . OpenReview.net, Singapore
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024 · 2024
Later among the works it cites.
Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
Emilia David. 2025 · 2025
Later among the works it cites.
Project Mariner: A research prototype exploring the future of human-agent interaction, starting with your browser
Google Deepmind. 2024 · 2025
Later among the works it cites.
Mastering diverse control tasks through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. 2025 · 2025
Later among the works it cites.