Fetching the paper…
Reading the bibliography…
The advent of large language models (LLMs) has opened up new opportunities in the field of mobile task automation.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Sus: a “quick and dirty’usability
John Brooke. 1996 · 1996
Earlier work this paper cites.
Determining what individual SUS scores mean: Adding an adjective rating scale
Aaron Bangor, Philip Kortum, and James Miller. 2009 · 2009
Earlier work this paper cites.
How do users like this feature? a fine grained sentiment analysis of app reviews. In 2014 IEEE 22nd international requirements engineering conference (RE) . Ieee, 153–162
Emitza Guzman and Walid Maalej. 2014 · 2014
Earlier work this paper cites.
Understanding usage states on mobile devices. In Proceedings of the 2015 ACM international joint conference on pervasive and ubiquitous computing . 1221–1225
Chakajkla Jesdabodi and Walid Maalej. 2015 · 2015
Earlier work this paper cites.
SUGILITE: Creating Multimodal Smartphone Automation by Demonstration. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (Denver, Colorado, USA) (CHI ’17) . Association for Computing Machinery, New York, NY, USA, 6038–6049
Toby Jia-Jun Li, Amos Azaria, and Brad A. Myers. 2017a · 2017
Earlier work this paper cites.
Programming IoT devices by demonstration using mobile apps. In End-User Development: 6th International Symposium, IS-EUD 2017, Eindhoven, The Netherlands, June 13-15, 2017, Proceedings 6 . Springer, 3–17
Toby Jia-Jun Li, Yuanchun Li, Fanglin Chen, and Brad A Myers. 2017b · 2017
Earlier work this paper cites.
KITE: Building conversational bots from mobile apps. In Proceedings of the 16th Annual International Conference on Mobile Systems, Applications, and Services . 96–109
Toby Jia-Jun Li and Oriana Riva. 2018 · 2018
Earlier work this paper cites.
Resource-rational task decomposition to minimize planning costs
Carlos G Correa, Mark K Ho, Fred Callaway, and Thomas L Griffiths. 2020 · 2020
Earlier work this paper cites.
The Digital Burnout Scale Development Study
Pınar Erten and Oguzhan Ozdemir. 2020 · 2020
Earlier work this paper cites.
Mapping natural language instructions to mobile UI action sequences
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge. 2020 · 2020
Earlier work this paper cites.
Human skill learning: expansion, exploration, selection, and refinement
Martin Lövdén, Benjamín Garzón, and Ulman Lindenberger. 2020 · 2020
Earlier work this paper cites.
Which app features are being used? Learning app feature usages from interaction data. In 2020 IEEE 28th International Requirements Engineering Conference (RE) . IEEE, 66–77
Christoph Stanik, Marlo Haering, Chakajkla Jesdabodi, and Walid Maalej. 2020 · 2020
Earlier work this paper cites.
Screen2vec: Semantic embedding of gui screens and gui components. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–15
Toby Jia-Jun Li, Lindsay Popowski, Tom Mitchell, and Brad A Myers. 2021 · 2021
Earlier work this paper cites.
Glider: A reinforcement learning approach to extract UI scripts from websites. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1420–1430
Yuanchun Li and Oriana Riva. 2021 · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
Androidenv: A reinforcement learning platform for android
Daniel Toyama, Philippe Hamel, Anita Gergely, Gheorghe Comanici, Amelia Glaese, Zafarali Ahmed, Tyler Jackson, Shibl Mourad, and Doina Precup. 2021 · 2021
Earlier work this paper cites.
Screen recognition: Creating accessibility metadata for mobile applications from pixels. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–15
Xiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White, Kyle Murray, Lisa Yu, Qi Shan, Jeffrey Nichols, Jason Wu, Chris Fleizach, et al · 2021
Earlier work this paper cites.
A-mash: providing single-app illusion for multi-app use through user-centric UI mashup. In Proceedings of the 28th Annual International Conference on Mobile Computing And Networking . 690–702
Sunjae Lee, Hoyoung Kim, Sijung Kim, Sangwook Lee, Hyosu Kim, Jean Young Song, Steven Y Ko, Sangeun Oh, and Insik Shin. 2022 · 2022
Earlier work this paper cites.
MUG: Interactive Multimodal Grounding on User Interfaces
Tao Li, Gang Li, Jingjie Zheng, Purple Wang, and Yang Li. 2022 · 2022
Earlier work this paper cites.
Mining detailed information from the description for App functions comparison
Huaxiao Liu, Xinglong Yin, Shanshan Song, Shanquan Gao, and Mengxi Zhang. 2022 · 2022
Earlier work this paper cites.
META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, and Kai Yu. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Talk to Claude
anthropic. 2023 · 2023
Earlier work this paper cites.
Use Siri on all your Apple devices
Apple. 2023 · 2023
Cited alongside, same era.
Humans decompose tasks by trading off utility and computational cost
Carlos G Correa, Mark K Ho, Frederick Callaway, Nathaniel D Daw, and Thomas L Griffiths. 2023 · 2023
Cited alongside, same era.
Prompt cache: Modular attention reuse for low-latency inference
In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong. 2023 · 2023
Cited alongside, same era.
Create your own accessibility service
Google. 2023a · 2023
Cited alongside, same era.
Hey Google
Google. 2023b · 2023
Cited alongside, same era.
Critic: Large language models can self-correct with tool-interactive critiquing
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2023 · 2023
Cutting-Edge PII Scanner
Endpoint Protector. 2023 · 2023
Closest in time.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al · 2023
Closest in time.
Cache & distil: Optimising API calls to large language models
Guillem Ramírez, Matthias Lindemann, Alexandra Birch, and Ivan Titov. 2023 · 2023
Closest in time.
Android in the wild: A large-scale dataset for android device control
Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy Lillicrap. 2023 · 2023
Closest in time.
Tptu: Task planning and tool usage of large language model-based ai agents
Jingqing Ruan, Yihong Chen, Bin Zhang, Zhiwei Xu, Tianpeng Bao, Guoqing Du, Shiwei Shi, Hangyu Mao, Xingyu Zeng, and Rui Zhao. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Digital dependence: Online fatigue and coping strategies during the COVID-19 lockdown
Emilie Munch et al Gregersen. 2023 · 2023
Cited alongside, same era.
Automatic Macro Mining from Interaction Traces at Scale
Forrest Huang, Gang Li, Tao Li, and Yang Li. 2023 · 2023
Cited alongside, same era.
Zero-shot compositional reinforcement learning in humans
Akshay Kumar Jagadish, Marcel Binz, Tankred Saanum, Jane X Wang, and Eric Schulz. 2023 · 2023
Cited alongside, same era.
Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu, Julian Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. 2023 · 2023
Cited alongside, same era.
Spotlight: Mobile UI understanding using vision-language models with a focus
Gang Li and Yang Li. 2023 · 2023
Cited alongside, same era.
Api-bank: A benchmark for tool-augmented llms
Minghao Li, Feifan Song, Bowen Yu, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. 2023 · 2023
Cited alongside, same era.
Closest in time.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023 · 2023
Closest in time.
AutoGPT: the heart of the open-source agent ecosystem
Significant-Gravitas. 2023 · 2023
Closest in time.
Yifan Song, Weimin Xiong, Dawei Zhu, Cheng Li, Ke Wang, Ye Tian, and Sujian Li. [n. d.] · 2023
Closest in time.
Ilias Stogiannidis, Stavros Vassos, Prodromos Malakasiotis, and Ion Androutsopoulos. 2023 · 2023
Closest in time.
Enabling conversational interaction with mobile ui using large language models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–17
Bryan Wang, Gang Li, and Yang Li. 2023b · 2023
Closest in time.
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023c · 2023
Closest in time.
Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023d · 2023
Closest in time.
Visionllm: Large language model is also an open-ended decoder for vision-centric tasks
Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al · 2023
Closest in time.
Empowering llm to use smartphone for intelligent task automation
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2023a · 2023
Closest in time.
DroidBot-GPT: GPT-powered UI Automation for Android
Hao Wen, Hongming Wang, Jiaxuan Liu, and Yuanchun Li. 2023b · 2023
Closest in time.
WebUI: A Dataset for Enhancing Visual UI Understanding with Web Semantics. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–14
Jason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng, Jeffrey Nichols, and Jeffrey P Bigham. 2023 · 2023
Closest in time.
Appagent: Multimodal agents as smartphone users
Zhao Yang, Jiaxuan Liu, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu. 2023 · 2023
Closest in time.
Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, et al · 2023
Closest in time.
Profile your app performance
Google. 2024 · 2024
Closest in time.
AppAgent-TencentQQGYLab
mnotgod96. 2024 · 2024
Closest in time.
AutoDroid
MobileLLM. 2024 · 2024
Closest in time.
Efficient Prompt Caching via Embedding Similarity
Hanlin Zhu, Banghua Zhu, and Jiantao Jiao. 2024 · 2024
Closest in time.