Fetching the paper…
Reading the bibliography…
Existing Visual Language Modelsoften struggle with information loss and limited reasoning abilities when handling high-resolution web interfaces that combine complex visual, textual, and interactive elements.
“Continuous, evolutionary and large-scale: A new perspective for automated mobile app testing,”
Mario Linares-Vásquez, Kevin Moran, and Denys Poshyvanyk, · 2017
Earlier work this paper cites.
“Reinforcement learning on web interfaces using workflow-guided exploration,”
Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang, · 2018
Earlier work this paper cites.
“Learning to navigate the web,”
Izzeddin Gur, Ulrich Rueckert, Aleksandra Faust, and Dilek Hakkani-Tur, · 2019
Earlier work this paper cites.
“Texts as images in prompt tuning for multi-label image recognition,”
Zixian Guo, Bowen Dong, Zhilong Ji, Jinfeng Bai, Yiwen Guo, and Wangmeng Zuo, · 2022
Earlier work this paper cites.
“Prompting visual-language models for efficient video understanding,”
Chen Ju, Tengda Han, Kunhao Zheng, and Zhang, · 2022
Earlier work this paper cites.
“Conditional prompt learning for vision-language models,”
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu, · 2022
Earlier work this paper cites.
“Ocr-free document understanding transformer,”
Geewook Kim, Teakgyu Hong, and Moonbin Yim, · 2022
Earlier work this paper cites.
“Visual instruction tuning,”
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee, · 2023
Earlier work this paper cites.
“T-rex: Counting by visual prompting,” 2023
Qing Jiang, Feng Li, Tianhe Ren, Shilong Liu, Zhaoyang Zeng, Kent Yu, and Lei Zhang, · 2023
Earlier work this paper cites.
“Appagent: Multimodal agents as smartphone users,”
China. Xiaoyan Zhang, Zhao Yang, Jiaxuan Liu, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu, · 2023
Earlier work this paper cites.
“RWKV: Reinventing RNNs for the transformer era,”
Bo Peng, Eric Alcaide, Quentin Anthony, and Alonz Albalak, · 2023
Earlier work this paper cites.
“Webui: A dataset for enhancing visual ui understanding with web semantics,”
Jason Wu, Siyan Wang, Siman Shen, and Peng, · 2023
Cited alongside, same era.
“Ocr-idl: Ocr annotations for industry document library dataset,”
Ali Furkan Biten, Rubèn Tito, and Lluis Gomez, · 2023
Cited alongside, same era.
“Deepseek-vl: Towards real-world vision-language understanding,”
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, and Bo Liu, · 2024
Cited alongside, same era.
“mplug-owi2: Revolutionizing multi-modal large language model with modality collaboration,”
Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Anwen Hu, Haowei Liu, Qi Qian, Ji Zhang, and Fei Huang, · 2024
Cited alongside, same era.
“Grounding multimodal large language models to the world,”
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, Qixiang Ye, and Furu Wei, · 2024
Cited alongside, same era.
“Doclayout-yolo: Enhancing document layout analysis through diverse synthetic data and global-to-local adaptive perception,” 2024
Zhiyuan Zhao, Hengrui Kang, Bin Wang, and Conghui He, · 2024
Later among the works it cites.
“Unlocking the conversion of web screenshots into HTML code with the websight dataset,”
Hugo Laurençon, Léo Tronchon, and Victor Sanh, · 2024
Later among the works it cites.
“Web2code: A large-scale webpage-to-code dataset and evaluation framework for multimodal llms,”
Sukmin Yun, haokun lin, Rusiru Thushara, and Mohammad Bhat, · 2024
Later among the works it cites.
“Screenqa: Large-scale question-answer pairs over mobile app screenshots,” 2024
Yu-Chung Hsiao, Fedir Zubach, Gilles Baechler, Victor Carbune, Jason Lin, Maria Wang, Srinivas Sunkara, Yun Zhu, and Jindong Chen, · 2024
Later among the works it cites.
“Enhanced visual instruction tuning for text-rich image understanding,” 2024
Yanzhe Zhang, Ruiyi Zhang, Jiuxiang Gu, Yufan Zhou, Nedim Lipka, Diyi Yang, and Tong Sun, · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Efficient multimodal learning from data-centric perspective,” 2024
Muyang He, Yexin Liu, Boya Wu, Jianhao Yuan, Yueze Wang, Tiejun Huang, and Bo Zhao, · 2024
Cited alongside, same era.
“Screenai: A vision-language model for ui and infographics understanding,” 2024
Gilles Baechler, Srinivas Sunkara, Maria Wang, and Fedir Zubach, · 2024
Cited alongside, same era.
“Mobile-agent: Autonomous multi-modal mobile device agent with visual perception,”
Junyang Wang, Haiyang Xu, and Jiabo Ye, · 2024
Cited alongside, same era.
“Cpt: Colorful prompt tuning for pre-trained vision-language models,”
Yuan Yao, Ao Zhang, and Zhengyan Zhang, · 2024
Cited alongside, same era.
“Mineru: An open-source solution for precise document content extraction,”
Bin Wang, Chao Xu, Xiaomeng Zhao, Linke Ouyang, Fan Wu, Zhiyuan Zhao, Rui Xu, Kaiwen Liu, Yuan Qu, Fukai Shang, et al., · 2024
Cited alongside, same era.
Later among the works it cites.
“Llava-next: What else influences visual instruction tuning beyond data?,” May 2024
Bo Li, Hao Zhang, Kaichen Zhang, Dong Guo, Yuanhan Zhang, Renrui Zhang, Feng Li, Ziwei Liu, and Chunyuan Li, · 2024
Later among the works it cites.
“Visualwebbench: How far have multimodal llms evolved in web page understanding and grounding?,” 2024
Junpeng Liu, Yifan Song, Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li, and Xiang Yue, · 2024
Later among the works it cites.
“Ferret-ui: Grounded mobile ui understanding with multimodal llms,”
Keen You, Haotian Zhang, and Eldon Schoop, · 2025
Closest in time.
“T-rex2: Towards generic object detection via text-visual prompt synergy,”
Qing Jiang, Feng Li, Zhaoyang Zeng, Tianhe Ren, Shilong Liu, and Lei Zhang, · 2025
Closest in time.
“VisualRWKV: Exploring recurrent neural networks for visual language models,”
Haowen Hou, Peigen Zeng, Fei Ma, and Fei Richard Yu, · 2025
Closest in time.