Fetching the paper…
Reading the bibliography…
Mobile task automation is an emerging field that leverages AI to streamline and optimize the execution of routine tasks on mobile devices, thereby enhancing efficiency and productivity.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Laws of organization in perceptual forms
Max Wertheimer. 1938 · 1938
Earlier work this paper cites.
Enhancing the explanatory power of usability heuristics. In Proceedings of the SIGCHI conference on Human Factors in Computing Systems . 152–158
Jakob Nielsen. 1994 · 1994
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd international conference on Machine learning . 369–376
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
Rico: A mobile app dataset for building data-driven design applications. In Proceedings of the 30th annual ACM symposium on user interface software and technology . 845–854
Biplab Deka, Zifeng Huang, Chad Franzen, Joshua Hibschman, Daniel Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar. 2017 · 2017
Earlier work this paper cites.
SUGILITE: creating multimodal smartphone automation by demonstration. In Proceedings of the 2017 CHI conference on human factors in computing systems . 6038–6049
Toby Jia-Jun Li, Amos Azaria, and Brad A Myers. 2017 · 2017
Earlier work this paper cites.
Appinite: A multi-modal interface for specifying data descriptions in programming by demonstration using natural language instructions. In 2018 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) . IEEE, 105–114
Toby Jia-Jun Li, Igor Labutov, Xiaohan Nancy Li, Xiaoyi Zhang, Wenze Shi, Wanling Ding, Tom M Mitchell, and Brad A Myers. 2018 · 2018
Earlier work this paper cites.
KITE: Building conversational bots from mobile apps. In Proceedings of the 16th Annual International Conference on Mobile Systems, Applications, and Services . 96–109
Toby Jia-Jun Li and Oriana Riva. 2018 · 2018
Earlier work this paper cites.
Examining image-based button labeling for accessibility in Android apps through large-scale analysis. In Proceedings of the 20th International ACM SIGACCESS Conference on Computers and Accessibility . 119–130
Anne Spencer Ross, Xiaoyi Zhang, James Fogarty, and Jacob O Wobbrock. 2018 · 2018
Earlier work this paper cites.
Object detection for graphical user interface: Old fashioned or deep learning or a combination?. In proceedings of the 28th ACM joint meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 1202–1214
Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen, Xiwei Xu, Liming Zhu, and Guoqiang Li. 2020 · 2020
Earlier work this paper cites.
Gtc: Guided training of ctc towards efficient and accurate scene text recognition. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 11005–11012
Wenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi, and Zhiping Lin. 2020 · 2020
Earlier work this paper cites.
Mapping natural language instructions to mobile UI action sequences
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge. 2020a · 2020
Earlier work this paper cites.
Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Association for Computational Linguistics, Online, 5495–5510
Yang Li, Gang Li, Luheng He, Jingjie Zheng, Hong Li, and Zhiwei Guan. 2020b · 2020
Earlier work this paper cites.
On Faithfulness and Factuality in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 1906–1919
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 2020
Earlier work this paper cites.
VASTA: a vision and language-assisted smartphone task automation system. In Proceedings of the 25th international conference on intelligent user interfaces . 22–32
Alborz Rezazadeh Sereshkeh, Gary Leung, Krish Perumal, Caleb Phillips, Minfan Zhang, Afsaneh Fazly, and Iqbal Mohomed. 2020 · 2020
Earlier work this paper cites.
Fuyu-8B: A Multimodal Architecture for AI Agents
ADEPT. 2021 · 2021
Earlier work this paper cites.
Vins: Visual search for mobile user interface design. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–14
Sara Bunian, Kai Li, Chaima Jemmali, Casper Harteveld, Yun Fu, and Magy Seif Seif El-Nasr. 2021 · 2021
Earlier work this paper cites.
Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, and Bryan A Plummer. 2021 · 2021
Earlier work this paper cites.
Understanding Mobile GUI: from Pixel-Words to Screen-Sentences
Jingwen Fu, Xiaoyi Zhang, Yuwang Wang, Wenjun Zeng, Sam Yang, and Grayson Hilliard. 2021 · 2021
Earlier work this paper cites.
Actionbert: Leveraging user actions for semantic understanding of user interfaces. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 5931–5938
Zecheng He, Srinivas Sunkara, Xiaoxue Zang, Ying Xu, Lijuan Liu, Nevan Wichers, Gabriel Schubiner, Ruby Lee, and Jindong Chen. 2021 · 2021
Earlier work this paper cites.
Screen2vec: Semantic embedding of gui screens and gui components. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–15
Toby Jia-Jun Li, Lindsay Popowski, Tom Mitchell, and Brad A Myers. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Etna: Harvesting action graphs from websites. In The 34th Annual ACM Symposium on User Interface Software and Technology . 312–331
Oriana Riva and Jason Kace. 2021 · 2021
Cited alongside, same era.
Screen2words: Automatic mobile UI summarization with multimodal learning. In The 34th Annual ACM Symposium on User Interface Software and Technology . 498–510
Bryan Wang, Gang Li, Xin Zhou, Zhourong Chen, Tovi Grossman, and Yang Li. 2021 · 2021
Cited alongside, same era.
Screen parsing: Towards reverse engineering of UI models from screenshots. In The 34th Annual ACM Symposium on User Interface Software and Technology . 470–483
Jason Wu, Xiaoyi Zhang, Jeff Nichols, and Jeffrey P Bigham. 2021 · 2021
Cited alongside, same era.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Closest in time.
Spotlight: Mobile UI understanding using vision-language models with a focus
Gang Li and Yang Li. 2023 · 2023
Closest in time.
A Zero-Shot Language Agent for Computer Control with Structured Reflection
Tao Li, Gang Li, Zhiwei Deng, Bryan Wang, and Yang Li. 2023a · 2023
Closest in time.
Chatting with GPT-3 for Zero-Shot Human-Like Mobile Automated GUI Testing
Zhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen, Boyu Wu, Xing Che, Dandan Wang, and Qing Wang. 2023 · 2023
Closest in time.
AndroidInTheWild: A Large-Scale Dataset For Android Device Control. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Google is trying to limit what apps can use an Accessibility Service (again)
XDA. 2021 · 2021
Cited alongside, same era.
Screen recognition: Creating accessibility metadata for mobile applications from pixels. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–15
Xiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White, Kyle Murray, Lisa Yu, Qi Shan, Jeffrey Nichols, Jason Wu, Chris Fleizach, et al · 2021
Cited alongside, same era.
Learning Semantically Rich Network-based Multi-modal Mobile User Interface Embeddings
Gary Ang and Ee-Peng Lim. 2022 · 2022
Cited alongside, same era.
A dataset for interactive vision-language navigation with unknown command feasibility. In European Conference on Computer Vision . Springer, 312–328
Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, and Bryan A Plummer. 2022 · 2022
Cited alongside, same era.
Towards complete icon labeling in mobile applications. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–14
Jieshan Chen, Amanda Swearngin, Jason Wu, Titus Barik, Jeffrey Nichols, and Xiaoyi Zhang. 2022 · 2022
Cited alongside, same era.
SVTR: Scene Text Recognition with a Single Visual Model. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22 , Lud De Raedt (Ed.). International Joint Conferences on Artificial Intelligence Organization, 884–890
Yongkun Du, Zhineng Chen, Caiyan Jia, Xiaoting Yin, Tianlun Zheng, Chenxia Li, Yuning Du, and Yu-Gang Jiang. 2022a · 2022
Cited alongside, same era.
ParamMacros: Creating UI Automation Leveraging End-User Natural Language Parameterization. In 2022 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) . IEEE, 1–10
Rebecca Krosnick and Steve Oney. 2022 · 2022
Cited alongside, same era.
Describing ui screenshots in natural language
Luis A Leiva, Asutosh Hota, and Antti Oulasvirta. 2022 · 2022
Cited alongside, same era.
Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy P Lillicrap. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Voicify Your UI: Towards Android App Control with Voice Commands
Minh Duc Vu, Han Wang, Zhuang Li, Gholamreza Haffari, Zhenchang Xing, and Chunyang Chen. 2023 · 2023
Closest in time.
Enabling conversational interaction with mobile ui using large language models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–17
Bryan Wang, Gang Li, and Yang Li. 2023 · 2023
Closest in time.
Empowering llm to use smartphone for intelligent task automation
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2023 · 2023
Closest in time.
GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
An Yan, Zhengyuan Yang, Wanrong Zhu, Kevin Lin, Linjie Li, Jianfeng Wang, Jianwei Yang, Yiwu Zhong, Julian McAuley, Jianfeng Gao, et al · 2023
Closest in time.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao. 2023 · 2023
Closest in time.
You Only Look at Screens: Multimodal Chain-of-Action Agents
Zhuosheng Zhan and Aston Zhang. 2023 · 2023
Closest in time.
Responsible Task Automation: Empowering Large Language Models as Responsible Task Automators
Zhizheng Zhang, Xiaoyi Zhang, Wenxuan Xie, and Yan Lu. 2023 · 2023
Closest in time.
Screenai: A vision-language model for ui and infographics understanding
Gilles Baechler, Srinivas Sunkara, Maria Wang, Fedir Zubach, Hassan Mansoor, Vincent Etter, Victor Cărbune, Jason Lin, Jindong Chen, and Abhanshu Sharma. 2024 · 2024
Closest in time.
Prompting is all you need: Automated android bug replay with large language models. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering . 1–13
Sidong Feng and Chunyang Chen. 2024 · 2024
Closest in time.
Cogagent: A visual language model for gui agents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 14281–14290
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, et al · 2024
Closest in time.
GPT-4V: Enhancing Vision-Based Tasks
OpenAI. 2023 · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024 · 2024
Closest in time.
Axnav: Replaying accessibility tests from natural language. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–16
Maryam Taeb, Amanda Swearngin, Eldon Schoop, Ruijia Cheng, Yue Jiang, and Jeffrey Nichols. 2024 · 2024
Closest in time.
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
Keen You, Haotian Zhang, Eldon Schoop, Floris Weers, Amanda Swearngin, Jeffrey Nichols, Yinfei Yang, and Zhe Gan. 2024 · 2024
Closest in time.