Fetching the paper…
Reading the bibliography…
Multimodal Vision-Language Models (VLMs) enable powerful applications from their fused understanding of images and language, but many perform poorly on UI tasks due to the lack of UI training data.
EAGER: Programming Repetitive Tasks by Example. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (New Orleans, Louisiana, USA) (CHI ’91) . Association for Computing Machinery, New York, NY, USA, 33–39
Allen Cypher. 1991 · 1991
Earlier work this paper cites.
Access to Graphical Interfaces for Blind Users
W. Keith Edwards, Elizabeth D. Mynatt, and Kathryn Stockton. 1995 · 1995
Earlier work this paper cites.
Rule-Based Detection for Reverse Engineering User Interfaces. In Proceedings of the 3rd Working Conference on Reverse Engineering (WCRE ’96) (WCRE ’96) . IEEE Computer Society, USA, 42
Melody M. Moore. 1996 · 1996
Earlier work this paper cites.
Using Knowledge Representation to Understand Interactive Systems. In Proceedings of the 5th International Workshop on Program Comprehension (WPC ’97) (WPC ’97) . IEEE Computer Society, USA, 60
Melody Moore and Spencer Rugaber. 1997 · 1997
Earlier work this paper cites.
User Interface Reengineering
Melody Marie Moore, James D. Foley, and Spencer Rugaber. 1998 · 1998
Earlier work this paper cites.
A Visual Medium for Programmatic Control of Interactive Applications. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Pittsburgh, Pennsylvania, USA) (CHI ’99) . Association for Computing Machinery, New York, NY, USA, 199–206
Luke S. Zettlemoyer and Robert St. Amant. 1999 · 1999
Earlier work this paper cites.
User Interface Reverse Engineering in Support of Interface Migration to the Web
E. Stroulia, M. El-Ramly, P. Iglinski, and P. Sorenson. 2003 · 2003
Earlier work this paper cites.
Reverse engineering Web applications: the WARE approach
G.A. Di Lucca, P. Fasolino, A.R.and IGLINSKI, and P. Tramontana. 2004 · 2004
Earlier work this paper cites.
Application Modeling using Reverse Engineering Techniques. In Proceedings of the 2006 ACM symposium on applied computing . ACM, 1250–1255
T. Katsimpa, Y. Panagis, E. Sakkopoulos, G. Tzimas, and A. Tsakalidis. 2006 · 2006
Earlier work this paper cites.
CoScripter: Automating & Sharing How-to Knowledge in the Enterprise. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Florence, Italy) (CHI ’08) . Association for Computing Machinery, New York, NY, USA, 1719–1728
Gilly Leshed, Eben M. Haber, Tara Matthews, and Tessa Lau. 2008 · 2008
Earlier work this paper cites.
Automated Reverse Engineering of Hard-Coded GUI Layouts. In Proceedings of the Ninth Conference on Australasian User Interface - Volume 76 (Wollongong, Australia) (AUIC ’08) . Australian Computer Society, Inc., AUS, 65–73
Christof Lutteroth. 2008 · 2008
Earlier work this paper cites.
RE-UWA approach to recover user centered conceptual models from Web applications
M. L. BERNARDI, G. A. DI LUCCA, and D. DISTANTE. 2009 · 2009
Earlier work this paper cites.
Sikuli: Using GUI Screenshots for Search and Automation. In Proceedings of the 22nd Annual ACM Symposium on User Interface Software and Technology (Victoria, BC, Canada) (UIST ’09) . Association for Computing Machinery, New York, NY, USA, 183–192
Tom Yeh, Tsung-Hsiang Chang, and Robert C. Miller. 2009 · 2009
Earlier work this paper cites.
Prefab: Implementing Advanced Behaviors Using Pixel-Based Reverse Engineering of Interface Structure. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Atlanta, Georgia, USA) (CHI ’10) . Association for Computing Machinery, New York, NY, USA, 1525–1534
Morgan Dixon and James Fogarty. 2010 · 2010
Earlier work this paper cites.
Associating the Visual Representation of User Interfaces with Their Internal Structures and Metadata. In Proceedings of the 24th Annual ACM Symposium on User Interface Software and Technology (Santa Barbara, California, USA) (UIST ’11) . Association for Computing Machinery, New York, NY, USA, 245–256
Tsung-Hsiang Chang, Tom Yeh, and Rob Miller. 2011 · 2011
Earlier work this paper cites.
Content and Hierarchy in Pixel-Based Methods for Reverse Engineering Interface Structure. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Vancouver, BC, Canada) (CHI ’11) . Association for Computing Machinery, New York, NY, USA, 969–978
Morgan Dixon, Daniel Leventhal, and James Fogarty. 2011 · 2011
Earlier work this paper cites.
Pause-and-Play: Automatically Linking Screencast Video Tutorials with Applications. In Proceedings of the 24th Annual ACM Symposium on User Interface Software and Technology (Santa Barbara, California, USA) (UIST ’11) . Association for Computing Machinery, New York, NY, USA, 135–144
Suporn Pongnumkul, Mira Dontcheva, Wilmot Li, Jue Wang, Lubomir Bourdev, Shai Avidan, and Michael F. Cohen. 2011 · 2011
Earlier work this paper cites.
Waken: Reverse Engineering Usage Information and Interface Structure from Software Videos. In Proceedings of the 25th Annual ACM Symposium on User Interface Software and Technology (Cambridge, Massachusetts, USA) (UIST ’12) . Association for Computing Machinery, New York, NY, USA, 83–92
Nikola Banovic, Tovi Grossman, Justin Matejka, and George Fitzmaurice. 2012 · 2012
Earlier work this paper cites.
Model-driven reverse engineering of legacy graphical user interfaces
A. Sanchez Ramon, J. Sanchez Cuadrado, and J. Garcia Molina. 2012 · 2012
Earlier work this paper cites.
Prefab Layers and Prefab Annotations: Extensible Pixel-Based Interpretation of Graphical Interfaces. In Proceedings of the 27th Annual ACM Symposium on User Interface Software and Technology (Honolulu, Hawaii, USA) (UIST ’14) . Association for Computing Machinery, New York, NY, USA, 221–230
Morgan Dixon, Alexander Nied, and James Fogarty. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context. In Computer Vision – ECCV 2014 , David Fleet, Tomas Pajdla, Bernt Schiele, and Tinne Tuytelaars (Eds.). Springer International Publishing, Cham, 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Reverse Engineering Mobile Application User Interfaces with REMAUI. In Proceedings of the 30th IEEE/ACM International Conference on Automated Software Engineering (Lincoln, Nebraska) (ASE ’15) . IEEE Press, 248–259
Tuan Anh Nguyen and Christoph Csallner. 2015 · 2015
Earlier work this paper cites.
Rico: A Mobile App Dataset for Building Data-Driven Design Applications. In Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology (Québec City, QC, Canada) (UIST ’17) . Association for Computing Machinery, New York, NY, USA, 845–854
Biplab Deka, Zifeng Huang, Chad Franzen, Joshua Hibschman, Daniel Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar. 2017 · 2017
Earlier work this paper cites.
SUGILITE: Creating Multimodal Smartphone Automation by Demonstration. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (Denver, Colorado, USA) (CHI ’17) . Association for Computing Machinery, New York, NY, USA, 6038–6049
Toby Jia-Jun Li, Amos Azaria, and Brad A. Myers. 2017 · 2017
Cited alongside, same era.
Genie: Input Retargeting on the Web through Command Reverse Engineering. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (Denver, Colorado, USA) (CHI ’17) . Association for Computing Machinery, New York, NY, USA, 4703–4714
Amanda Swearngin, Amy J. Ko, and James Fogarty. 2017 · 2017
Cited alongside, same era.
Pix2code: Generating Code from a Graphical User Interface Screenshot. In Proceedings of the ACM SIGCHI Symposium on Engineering Interactive Computing Systems (Paris, France) (EICS ’18) . Association for Computing Machinery, New York, NY, USA, Article 3, 6 pages
Tony Beltramelli. 2018 · 2018
Cited alongside, same era.
Robust Relational Layout Synthesis from Examples for Android
Pavol Bielik, Marc Fischer, and Martin Vechev. 2018 · 2018
OpenCLIP
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. 2021 · 2021
Later among the works it cites.
ReverseORC: Reverse Engineering of Resizable User Interface Layouts with OR-Constraints
Yue Jiang, Wolfgang Stuerzlinger, and Christof Lutteroth. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Screen Parsing: Towards Reverse Engineering of UI Models from Screenshots. In The 34th Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA) (UIST ’21) . Association for Computing Machinery, New York, NY, USA, 470–483
Jason Wu, Xiaoyi Zhang, Jeff Nichols, and Jeffrey P Bigham. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
From UI Design Image to GUI Skeleton: A Neural Machine Translator to Bootstrap Mobile GUI Implementation. In Proceedings of the 40th International Conference on Software Engineering (Gothenburg, Sweden) (ICSE ’18) . Association for Computing Machinery, New York, NY, USA, 665–676
Chunyang Chen, Ting Su, Guozhu Meng, Zhenchang Xing, and Yang Liu. 2018 · 2018
Cited alongside, same era.
Caption Crawler: Enabling Reusable Alternative Text Descriptions Using Reverse Image Search. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18) . Association for Computing Machinery, New York, NY, USA, 1–11
Darren Guinness, Edward Cutrell, and Meredith Ringel Morris. 2018 · 2018
Cited alongside, same era.
Expresso: Building Responsive Interfaces with Keyframes. In 2018 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) . 39–47
R. Krosnick, S. W. Lee, W. S. Laseck, and S. Onev. 2018 · 2018
Cited alongside, same era.
Machine Learning-Based Prototyping of Graphical User Interfaces for Mobile Apps
Kevin Moran, Carlos Bernal-Cárdenas, Michael Curcio, Richard Bonett, and Denys Poshyvanyk. 2020 · 2018
Cited alongside, same era.
Rewire: Interface Design Assistance from Examples. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18) . Association for Computing Machinery, New York, NY, USA, 1–12
Amanda Swearngin, Mira Dontcheva, Wilmot Li, Joel Brandt, Morgan Dixon, and Amy J. Ko. 2018 · 2018
Cited alongside, same era.
Gallery D.C.: Design Search and Knowledge Discovery through Auto-Created GUI Component Gallery
Chunyang Chen, Sidong Feng, Zhenchang Xing, Linda Liu, Shengdong Zhao, and Jinshui Wang. 2019b · 2019
Cited alongside, same era.
GUI-Squatting Attack: Automated Generation of Android Phishing Apps
Sen Chen, Lingling Fan, Chunyang Chen, Minhui Xue, Yang Liu, and Lihua Xu. 2021 · 2019
Cited alongside, same era.
Automated Cross-Platform GUI Code Generation for Mobile Apps. In 2019 IEEE 1st International Workshop on Artificial Intelligence for Mobile (AI4Mobile) . 13–16
Sen Chen, Lingling Fan, Ting Su, Lei Ma, Yang Liu, and Lihua Xu. 2019a · 2019
Cited alongside, same era.
Screen recognition: Creating accessibility metadata for mobile applications from pixels. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–15
Xiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White, Kyle Murray, Lisa Yu, Qi Shan, Jeffrey Nichols, Jason Wu, Chris Fleizach, et al · 2021
Later among the works it cites.
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. 2022 · 2022
Later among the works it cites.
Understanding Screen Relationships from Screenshots of Smartphone Applications. In 27th International Conference on Intelligent User Interfaces . 447–458
Shirin Feiz, Jason Wu, Xiaoyi Zhang, Amanda Swearngin, Titus Barik, and Jeffrey Nichols. 2022 · 2022
Later among the works it cites.
Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus
Gang Li and Yang Li. 2022 · 2022
Later among the works it cites.
Rediscovering Affordance: A Reinforcement Learning Perspective. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22) . Association for Computing Machinery, New York, NY, USA, Article 362, 15 pages
Yi-Chi Liao, Kashyap Todi, Aditya Acharya, Antti Keurulainen, Andrew Howes, and Antti Oulasvirta. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Later among the works it cites.
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22) . Association for Computing Machinery, New York, NY, USA, Article 36, 21 pages
Eldon Schoop, Xin Zhou, Gang Li, Zhourong Chen, Bjoern Hartmann, and Yang Li. 2022 · 2022
Later among the works it cites.
PaLI: A Jointly-Scaled Multilingual Language-Image Model
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut. 2023 · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Closest in time.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023 · 2023
Closest in time.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023b · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023 · 2023
Closest in time.
Stanford Alpaca: An Instruction-following LLaMA model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Enabling Conversational Interaction with Mobile UI Using Large Language Models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 432, 17 pages
Bryan Wang, Gang Li, and Yang Li. 2023 · 2023
Closest in time.
Empowering LLM to use Smartphone for Intelligent Task Automation
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2023 · 2023
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Closest in time.