Fetching the paper…
Reading the bibliography…
UI task automation enables efficient task execution by simulating human interactions with graphical user interfaces (GUIs), without modifying the existing application code.
APPINITE: A Multi-Modal Interface for Specifying Data Descriptions in Programming by Demonstration Using Natural Language Instructions. In 2018 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) . 105–114
Toby Jia-Jun Li, Igor Labutov, Xiaohan Nancy Li, Xiaoyi Zhang, Wenze Shi, Wanling Ding, Tom M. Mitchell, and Brad A. Myers. 2018 · 1943
Earlier work this paper cites.
Introduction to the ISO specification language LOTOS
Tommaso Bolognesi and Ed Brinksma. 1987 · 1987
Earlier work this paper cites.
ConcurTaskTrees: A Diagrammatic Notation for Specifying Task Models. In Proceedings of the IFIP TC13 Interantional Conference on Human-Computer Interaction (INTERACT ’97) . Chapman & Hall, Ltd., GBR, 362–369
Fabio Paternò, Cristiano Mancini, and Silvia Meniconi. 1997 · 1997
Earlier work this paper cites.
Real life information retrieval: a study of user queries on the Web
Bernard J. Jansen, Amanda Spink, Judy Bateman, and Tefko Saracevic. 1998 · 1998
Earlier work this paper cites.
Clustering User Queries of a Search Engine. In Proceedings of the 10th International Conference on World Wide Web (Hong Kong, Hong Kong) (WWW ’01) . Association for Computing Machinery, New York, NY, USA, 162–168
Ji-Rong Wen, Jian-Yun Nie, and Hong-Jiang Zhang. 2001 · 2001
Earlier work this paper cites.
Interactive Machine Learning. In Proceedings of the 8th International Conference on Intelligent User Interfaces (Miami, Florida, USA) (IUI ’03) . Association for Computing Machinery, New York, NY, USA, 39–45
Jerry Alan Fails and Dan R. Olsen. 2003 · 2003
Earlier work this paper cites.
Comparing task models for user interface design
Quentin Limbourg and Jean Vanderdonckt. 2004 · 2004
Earlier work this paper cites.
Cooperative multi-agent learning: The state of the art
Liviu Panait and Sean Luke. 2005 · 2005
Earlier work this paper cites.
Process Modeling using Event-Driven Process Chains
August-Wilhelm Scheer, Oliver Thomas, and Otmar Adam. 2005 · 2005
Earlier work this paper cites.
Hierarchical task analysis: Developments, applications, and extensions
Neville A. Stanton. 2006 · 2005
Earlier work this paper cites.
Icon-function relationship in toolbar icons
Stefano Passini, Filiberto Strazzari, and Annamaria Borghi. 2008 · 2008
Earlier work this paper cites.
Reinforcement Learning for Mapping Instructions to Actions. In Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP . Association for Computational Linguistics, Suntec, Singapore, 82–90
S.R.K. Branavan, Harr Chen, Luke Zettlemoyer, and Regina Barzilay. 2009 · 2009
Earlier work this paper cites.
Mobile application and its global impact
Rashedul Islam, Rofiqul Islam, and Tohidul Mazumder. 2010 · 2010
Earlier work this paper cites.
Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements
Yang Li, Gang Li, Luheng He, Jingjie Zheng, Hong Li, and Zhiwei Guan. 2020b · 2010
Earlier work this paper cites.
Pause-and-Play: Automatically Linking Screencast Video Tutorials with Applications. In Proceedings of the 24th Annual ACM Symposium on User Interface Software and Technology (Santa Barbara, California, USA) (UIST ’11) . Association for Computing Machinery, New York, NY, USA, 135–144
Suporn Pongnumkul, Mira Dontcheva, Wilmot Li, Jue Wang, Lubomir Bourdev, Shai Avidan, and Michael F. Cohen. 2011 · 2011
Earlier work this paper cites.
Touch-Based Mobile Phone Interface Guidelines and Design Recommendations for Elderly People: A Survey of the Literature. In Neural Information Processing , Tingwen Huang, Zhigang Zeng, Chuandong Li, and Chi Sing Leung (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 568–574
Muna S. Al-Razgan, Hend S. Al-Khalifa, Mona D. Al-Shahrani, and Hessah H. AlAjmi. 2012 · 2012
Earlier work this paper cites.
Power to the People: The Role of Humans in Interactive Machine Learning
Saleema Amershi, Maya Cakmak, William Bradley Knox, and Todd Kulesza. 2014 · 2014
Earlier work this paper cites.
Pixel-based methods for widget state and style in a runtime implementation of sliding widgets. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Toronto, Ontario, Canada) (CHI ’14) . Association for Computing Machinery, New York, NY, USA, 2231–2240
Morgan Dixon, Gierad Laput, and James Fogarty. 2014 · 2014
Earlier work this paper cites.
EverTutor: automatically creating interactive guided tutorials on smartphones by user demonstration. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . ACM, Toronto Ontario Canada, 4027–4036
Cheng-Yao Wang, Wei-Chen Chu, Hou-Ren Chen, Chun-Yen Hsu, and Mike Y. Chen. 2014 · 2014
Earlier work this paper cites.
Intelligently Creating Contextual Tutorials for GUI Applications. In 2015 IEEE 12th Intl Conf on Ubiquitous Intelligence and Computing and 2015 IEEE 12th Intl Conf on Autonomic and Trusted Computing and 2015 IEEE 15th Intl Conf on Scalable Computing and Communications and Its Associated Workshops (UIC-ATC-ScalCom) . IEEE, Beijing, 187–196
Guo Li, Tun Lu, Jiang Yang, Xiaomu Zhou, Xianghua Ding, and Ning Gu. 2015 · 2015
Earlier work this paper cites.
CoFaçade: A Customizable Assistive Approach for Elders and Their Helpers. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (Seoul, Republic of Korea) (CHI ’15) . Association for Computing Machinery, New York, NY, USA, 1583–1592
Jason Chen Zhao, Richard C. Davis, Pin Sym Foong, and Shengdong Zhao. 2015 · 2015
Earlier work this paper cites.
Automation of a Business Process Using Robotic Process Automation (RPA): A Case Study. In Applied Computer Sciences in Engineering , Juan Carlos Figueroa-García, Eduyn Ramiro López-Santana, José Luis Villa-Ramírez, and Roberto Ferro-Escobar (Eds.). Springer International Publishing, Cham, 65–71
Santiago Aguirre and Alejandro Rodriguez. 2017 · 2017
Earlier work this paper cites.
Design of Interactive Tutorials on Mobile Applications for Chinese Middle-Aged and Older Adults
Xiaoou Chen, Fei Wang, Zhenwei You, Xiaochun Wang, Chunjing Tao, and Jian Liu. 2017 · 2017
Earlier work this paper cites.
Rico: A Mobile App Dataset for Building Data-Driven Design Applications. In Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology . ACM, Québec City QC Canada, 845–854
Biplab Deka, Zifeng Huang, Chad Franzen, Joshua Hibschman, Daniel Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar. 2017 · 2017
Earlier work this paper cites.
SUGILITE: Creating Multimodal Smartphone Automation by Demonstration. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (Denver, Colorado, USA) (CHI ’17) . Association for Computing Machinery, New York, NY, USA, 6038–6049
Toby Jia-Jun Li, Amos Azaria, and Brad A. Myers. 2017 · 2017
Earlier work this paper cites.
Multi-Agent Systems: A Survey
Ali Dorri, Salil S. Kanhere, and Raja Jurdak. 2018 · 2018
Earlier work this paper cites.
A Review of User Interface Design for Interactive Machine Learning
John J. Dudley and Per Ola Kristensson. 2018 · 2018
Earlier work this paper cites.
Robotic process automation: overview and opportunities
Stefan Z Jovanović, Jelena S Đurić, and Tatjana V Šibalija. 2018 · 2018
Earlier work this paper cites.
Kite: Building Conversational Bots from Mobile Apps. In Proceedings of the 16th Annual International Conference on Mobile Systems, Applications, and Services (Munich, Germany) (MobiSys ’18) . Association for Computing Machinery, New York, NY, USA, 96–109
Toby Jia-Jun Li and Oriana Riva. 2018 · 2018
Earlier work this paper cites.
Learning Design Semantics for Mobile Apps. In Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology (Berlin, Germany) (UIST ’18) . Association for Computing Machinery, New York, NY, USA, 569–579
Thomas F. Liu, Mark Craft, Jason Situ, Ersin Yumer, Radomir Mech, and Ranjitha Kumar. 2018 · 2018
Cited alongside, same era.
Mapping Natural Language Commands to Web Elements
Panupong Pasupat, Tian-Shun Jiang, Evan Zheran Liu, Kelvin Guu, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Communication Breakdowns Between Families and Alexa. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19) . Association for Computing Machinery, New York, NY, USA, 1–13
Erin Beneteau, Olivia K. Richards, Mingrui Zhang, Julie A. Kientz, Jason Yip, and Alexis Hiniker. 2019 · 2019
Cited alongside, same era.
Ilastik: interactive machine learning for (bio) image analysis
Stuart Berg, Dominik Kutra, Thorben Kroeger, Christoph N Straehle, Bernhard X Kausler, Carsten Haubold, Martin Schiegg, Janez Ales, Thorsten Beier, Markus Rudy, et al · 2019
UGIF: UI Grounded Instruction Following
Sagar Gubbi Venkatesh, Partha Talukdar, and Srini Narayanan. 2022 · 2022
Later among the works it cites.
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022 · 2022
Later among the works it cites.
Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu. 2023 · 2023
Later among the works it cites.
Improving Factuality and Reasoning in Language Models through Multiagent Debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
RePlay: Contextually Presenting Learning Videos Across Software Applications. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19) . Association for Computing Machinery, New York, NY, USA, 1–13
C. Ailie Fraser, Tricia J. Ngoon, Mira Dontcheva, and Scott Klemmer. 2019 · 2019
Cited alongside, same era.
Chatbots, Humbots, and the Quest for Artificial General Intelligence. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19) . Association for Computing Machinery, New York, NY, USA, 1–11
Jonathan Grudin and Richard Jacques. 2019 · 2019
Cited alongside, same era.
VIANA: Visual Interactive Annotation of Argumentation. In 2019 IEEE Conference on Visual Analytics Science and Technology (VAST) . 11–22
Fabian Sperrle, Rita Sevastjanova, Rebecca Kehlbeck, and Mennatallah El-Assady. 2019 · 2019
Cited alongside, same era.
Improving random GUI testing with image-based widget detection. In Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis (Beijing, China) (ISSTA 2019) . Association for Computing Machinery, New York, NY, USA, 307–317
Thomas D. White, Gordon Fraser, and Guy J. Brown. 2019 · 2019
Cited alongside, same era.
Teachable Machine: Approachable Web-Based Tool for Exploring Machine Learning Classification. In Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI EA ’20) . Association for Computing Machinery, New York, NY, USA, 1–8
Michelle Carney, Barron Webster, Irene Alvarado, Kyle Phillips, Noura Howell, Jordan Griffith, Jonas Jongejan, Amit Pitaru, and Alexander Chen. 2020 · 2020
Cited alongside, same era.
Interactive machine learning for soybean seed and seedling quality classification
André Dantas de Medeiros, Nayara Pereira Capobiango, Jose Maria da Silva, Laercio Junio da Silva, Clissia Barboza da Silva, and Denise Cunha Fernandes dos Santos Dias. 2020 · 2020
Cited alongside, same era.
Integrating Machine Learning with Human Knowledge
Changyu Deng, Xunbi Ji, Colton Rainey, Jianyu Zhang, and Wei Lu. 2020 · 2020
Cited alongside, same era.
Mapping Natural Language Instructions to Mobile UI Action Sequences. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Online, 8198–8210
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge. 2020a · 2020
Cited alongside, same era.
Lukas-Valentin Herm, Christian Janiesch, Alexander Helm, Florian Imgrund, Adrian Hofmann, and Axel Winkelmann. 2023 · 2023
Later among the works it cites.
Interaction Proxy Manager: Semantic Model Generation and Run-Time Support for Reconstructing Ubiquitous User Interfaces of Mobile Services
Tian Huang, Chun Yu, Weinan Shi, Bowen Wang, David Yang, Yihao Zhu, Zhaoheng Li, and Yuanchun Shi. 2023 · 2023
Later among the works it cites.
Survey of Hallucination in Natural Language Generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Later among the works it cites.
Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding. In Proceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 202) , Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 18893–18912
Kenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu, Fangyu Liu, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. 2023 · 2023
Later among the works it cites.
Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus
Gang Li and Yang Li. 2023 · 2023
Later among the works it cites.
Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2023 · 2023
Later among the works it cites.
Towards A Unified Agent with Foundation Models
Norman Di Palo, Arunkumar Byravan, Leonard Hasenclever, Markus Wulfmeier, Nicolas Heess, and Martin Riedmiller. 2023 · 2023
Later among the works it cites.
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
ChatGPT and Other Large Language Models Are Double-edged Swords
Yiqiu Shen, Laura Heacock, Jonathan Elias, Keith Hentel, Beatriu Reig, George Shih, and Linda Moy. 2023 · 2023
Later among the works it cites.
ProgPrompt: Generating Situated Robot Task Plans using Large Language Models. In 2023 IEEE International Conference on Robotics and Automation (ICRA) . 11523–11530
Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. 2023 · 2023
Later among the works it cites.
Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents
Yashar Talebirad and Amirhossein Nadiri. 2023 · 2023
Later among the works it cites.
Voicify Your UI: Towards Android App Control with Voice Commands
Minh Duc Vu, Han Wang, Zhuang Li, Gholamreza Haffari, Zhenchang Xing, and Chunyang Chen. 2023 · 2023
Later among the works it cites.
Enabling Conversational Interaction with Mobile UI using Large Language Models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 432, 17 pages
Bryan Wang, Gang Li, and Yang Li. 2023 · 2023
Later among the works it cites.
Multi-Party Chat: Conversational Agents in Group Settings with Humans and Models
Jimmy Wei, Kurt Shuster, Arthur Szlam, Jason Weston, Jack Urbanek, and Mojtaba Komeili. 2023 · 2023
Later among the works it cites.
Never-Ending Learning of User Interfaces. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (UIST ’23) . Association for Computing Machinery, New York, NY, USA, Article 113, 13 pages
Jason Wu, Rebecca Krosnick, Eldon Schoop, Amanda Swearngin, Jeffrey P Bigham, and Jeffrey Nichols. 2023 · 2023
Later among the works it cites.
ExpeL: LLM Agents Are Experiential Learners
Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. 2023 · 2023
Later among the works it cites.
Prompting Is All You Need: Automated Android Bug Replay with Large Language Models. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (Lisbon, Portugal) (ICSE ’24) . Association for Computing Machinery, New York, NY, USA, Article 67, 13 pages
Sidong Feng and Chunyang Chen. 2024 · 2024
Closest in time.
Automatic Macro Mining from Interaction Traces at Scale. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24) . Association for Computing Machinery, New York, NY, USA, Article 1038, 16 pages
Forrest Huang, Gang Li, Tao Li, and Yang Li. 2024 · 2024
Closest in time.
MUG: Interactive Multimodal Grounding on User Interfaces. In Findings of the Association for Computational Linguistics: EACL 2024 , Yvette Graham and Matthew Purver (Eds.). Association for Computational Linguistics, St. Julian’s, Malta, 231–251
Tao Li, Gang Li, Jingjie Zheng, Purple Wang, and Yang Li. 2024a · 2024
Closest in time.
VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task Planning. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology (Pittsburgh, PA, USA) (UIST ’24) . Association for Computing Machinery, New York, NY, USA, Article 49, 17 pages
Yunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma, and Zhongmin Cai. 2024 · 2024
Closest in time.
AXNav: Replaying Accessibility Tests from Natural Language. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24) . Association for Computing Machinery, New York, NY, USA, Article 962, 16 pages
Maryam Taeb, Amanda Swearngin, Eldon Schoop, Ruijia Cheng, Yue Jiang, and Jeffrey Nichols. 2024 · 2024
Closest in time.
GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology (UIST ’24) . ACM, 1–17
Minh Duc Vu, Han Wang, Jieshan Chen, Zhuang Li, Shengdong Zhao, Zhenchang Xing, and Chunyang Chen. 2024 · 2024
Closest in time.
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
Junyang Wang, Haiyang Xu, Jiabo Ye, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang. 2024 · 2024
Closest in time.
AutoDroid: LLM-powered Task Automation in Android
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2024 · 2024
Closest in time.
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs. In Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part LXIV (Milan, Italy). Springer-Verlag, Berlin, Heidelberg, 240–255
Keen You, Haotian Zhang, Eldon Schoop, Floris Weers, Amanda Swearngin, Jeffrey Nichols, Yinfei Yang, and Zhe Gan. 2024 · 2024
Closest in time.