Fetching the paper…
Reading the bibliography…
Recent advances in Large Vision-Language Models (LVLMs) have significantly improve performance in image comprehension tasks, such as formatted charts and rich-content images.
Using a goal-driven approach to generate test cases for guis
Atif M. Memon, Martha E. Pollack, and Mary Lou Soffa. 1999 · 1999
Earlier work this paper cites.
Exploiting cloze questions for few shot text classification and natural language inference
Timo Schick and Hinrich Schütze. 2020 · 2001
Earlier work this paper cites.
Modeling and testing hierarchical guis
Ana CR Paiva, Nikolai Tillmann, João CP Faria, and Raul FAM Vidal. 2005 · 2005
Earlier work this paper cites.
Object detection for graphical user interface: Old fashioned or deep learning or a combination?
Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen, Xiwei Xu, Liming Zhu, and Guoqiang Li. 2020b · 2008
Earlier work this paper cites.
Multilayer perceptron and neural networks
Marius-Constantin Popescu, Valentina E Balas, Liliana Perescu-Popescu, and Nikos Mastorakis. 2009 · 2009
Earlier work this paper cites.
A systematic capture and replay strategy for testing complex gui based java applications
Omar El Ariss, Dianxiang Xu, Santosh Dandey, Brad Vender, Phil McClean, and Brian Slator. 2010 · 2010
Earlier work this paper cites.
Graphical user interface (gui) testing: Systematic mapping and repository
Ishan Banerjee, Bao Nguyen, Vahid Garousi, and Atif Memon. 2013 · 2013
Earlier work this paper cites.
Rico: A mobile app dataset for building data-driven design applications
Biplab Deka, Zifeng Huang, Chad Franzen, Joshua Hibschman, Daniel Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar. 2017 · 2017
Earlier work this paper cites.
Alpaca: intermittent execution without checkpoints
Kiwan Maeng, Alexei Colin, and Brandon Lucia. 2017 · 2017
Earlier work this paper cites.
Yauhen Leanidavich Arnatovich and Lipo Wang. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018 · 2018
Earlier work this paper cites.
Continuous, evolutionary and large-scale: A new perspective for automated mobile app testing
Mario Linares Vásquez, Kevin Moran, and Denys Poshyvanyk. 2018 · 2018
Earlier work this paper cites.
Apply computer vision in gui automation for industrial applications
Yung-Pin Cheng, Ching-Wei Li, and Yi-Cheng Chen. 2019 · 2019
Earlier work this paper cites.
Unblind your apps: Predicting natural-language labels for mobile gui components by deep learning
Jieshan Chen, Chunyang Chen, Zhenchang Xing, Xiwei Xu, Liming Zhut, Guoqiang Li, and Jinshui Wang. 2020a · 2020
Earlier work this paper cites.
Gui testing for mobile applications: objectives, approaches and challenges
Kabir Sulaiman Said, Liming Nie, Adekunle Akinjobi Ajibode, and Xueyi Zhou. 2020 · 2020
Earlier work this paper cites.
A replication package for it takes two to tango: Combining visual and textual information for detecting duplicate video-based bug reports
Nathan Cooper, Carlos Bernal-Cárdenas, Oscar Chaparro, Kevin Moran, and Denys Poshyvanyk. 2021 · 2021
Earlier work this paper cites.
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Earlier work this paper cites.
Layout and image recognition driving cross-platform automated mobile testing
Shengcheng Yu, Chunrong Fang, Yexiao Yun, Yang Feng, and Zhenyu Chen. 2020 · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022 · 2022
Cited alongside, same era.
A survey on the use of computer vision to improve software engineering tasks
M. Bajammal, A. Stocco, D. Mazinanian, and A. Mesbah. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023 · 2023
Wizardlm: Empowering large language models to follow complex instructions
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023 · 2023
Later among the works it cites.
Empowering llm-based machine translation with cultural awareness
Binwei Yao, Ming Jiang, Diyi Yang, and Junjie Hu. 2023 · 2023
Later among the works it cites.
A survey on multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. 2023 · 2023
Later among the works it cites.
Llm for test script generation and migration: Challenges, capabilities, and opportunities
Shengcheng Yu, Chunrong Fang, Yuchen Ling, Chentian Wu, and Zhenyu Chen. 2023a · 2023
Later among the works it cites.
Grounding visual illusions in language: Do vision-language models perceive illusions like humans?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pro-cap: Leveraging a frozen vision-language model for hateful meme detection
Rui Cao, Ming Shan Hee, Adriel Kuek, Wen-Haw Chong, Roy Ka-Wei Lee, and Jing Jiang. 2023 · 2023
Cited alongside, same era.
Sharegpt4v: Improving large multi-modal models with better captions
Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. 2023 · 2023
Cited alongside, same era.
Adapting large language models via reading comprehension
Daixuan Cheng, Shaohan Huang, and Furu Wei. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023 · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023 · 2023
Cited alongside, same era.
Eva: Exploring the limits of masked visual representation learning at scale
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. 2023 · 2023
Cited alongside, same era.
Chatgpt outperforms crowd workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Cited alongside, same era.
Yichi Zhang, Jiayi Pan, Yuchen Zhou, Rui Pan, and Joyce Chai. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Later among the works it cites.
CogAgent: A Visual Language Model for GUI Agents
Hong, W., Wang, W., Lv, Q., Xu, J., Yu, W., Ji, J., Wang, Y., Wang, Z., Zhang, Y., Li, J., Xu, B., Dong, Y., Ding, M. & Tang, J. 2023 · 2023
Later among the works it cites.
Leancontext: Cost-efficient domain-specific question answering using llms
Md Adnan Arefeen, Biplob Debnath, and Srimat Chakradhar. 2024 · 2024
Closest in time.
Allava: Harnessing gpt4v-synthesized data for a lite vision-language model
Guiming Hardy Chen, Shunian Chen, Ruifei Zhang, Junying Chen, Xiangbo Wu, Zhiyi Zhang, Zhihong Chen, Jianquan Li, Xiang Wan, and Benyou Wang. 2024 · 2024
Closest in time.
Large language models for mobile gui text input generation: An empirical study
Chenhui Cui, Tao Li, Junjie Wang, Chunyang Chen, Dave Towey, and Rubing Huang. 2024 · 2024
Closest in time.
Minicpm: Unveiling the potential of small language models with scalable training strategies
Shengding Hu, Yuge Tu, Xu Han, and Chaoqun He. 2024 · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024 · 2024
Closest in time.
What matters when building vision-language models?
Hugo Laurençon, Léo Tronchon, Matthieu Cord, and Victor Sanh. 2024 · 2024
Closest in time.
OpenAI. 2024 · 2024
Closest in time.
Illusionvqa: A challenging optical illusion dataset for vision language models
Haz Sameen Shahgir, Khondker Salman Sayeed, Abhik Bhattacharjee, Wasi Uddin Ahmad, Yue Dong, and Rifat Shahriyar. 2024 · 2024
Closest in time.
Large language model driven automated software application testing
Fei Wang, Kamakshi Kodur, Michael Micheletti, Shu-Wei Cheng, and Yogalakshmi Sadasivam. 2024 · 2024
Closest in time.
Vision-language models for vision tasks: A survey
Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024 · 2024
Closest in time.
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
You, K., Zhang, H., Schoop, E., Weers, F., Swearngin, A., Nichols, J., Yang, Y. & Gan, Z. 2024 · 2024
Closest in time.