Fetching the paper…
Reading the bibliography…
Recent advancements in multimodal large language models (MLLMs) have been noteworthy, yet, these general-domain MLLMs often fall short in their ability to comprehend and interact effectively with user interface (UI) screens.
Edwards, W.K., Mynatt, E.D., Stockton, K.: Access to graphical interfaces for blind users. Interactions 2
1995
Earlier work this paper cites.
Amalfitano, D., Fasolino, A.R., Tramontana, P.: A gui crawling-based technique for android mobile application testing. In: 2011 IEEE fourth international conference on software testing, verification and validation workshops. pp. 252–261. IEEE (2011)
2011
Earlier work this paper cites.
Deka, B., Huang, Z., Franzen, C., Hibschman, J., Afergan, D., Li, Y., Nichols, J., Kumar, R.: Rico: A mobile app dataset for building data-driven design applications. In: Proceedings of the 30th annual ACM symposium on user interface software and technology. pp. 845–854 (2017)
2017
Earlier work this paper cites.
Linares-Vásquez, M., Moran, K., Poshyvanyk, D.: Continuous, evolutionary and large-scale: A new perspective for automated mobile app testing. In: 2017 IEEE International Conference on Software Maintenance and Evolution (ICSME). pp. 399–410. IEEE (2017)
2017
Earlier work this paper cites.
Shi, T., Karpathy, A., Fan, L., Hernandez, J., Liang, P.: World of bits: An open-domain platform for web-based agents. In: International Conference on Machine Learning. pp. 3135–3144. PMLR (2017)
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Jiang, Z., Kuang, R., Gong, J., Yin, H., Lyu, Y., Zhang, X.: What makes a great mobile app? a quantitative study using a new mobile crawler. In: 2018 IEEE Symposium on Service-Oriented System Engineering (SOSE). pp. 222–227. IEEE (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2020
Earlier work this paper cites.
Bai, C., Zang, X., Xu, Y., Sunkara, S., Rastogi, A., Chen, J., y Arcas, B.A.: Uibert: Learning generic multimodal representations for ui understanding (2021)
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning. pp. 8748–8763. PMLR (2021)
2021
Earlier work this paper cites.
Zhang, X., de Greef, L., Swearngin, A., White, S., Murray, K., Yu, L., Shan, Q., Nichols, J., Wu, J., Fleizach, C., Everitt, A., Bigham, J.P.: Screen recognition: Creating accessibility metadata for mobile applications from pixels (2021)
2021
Earlier work this paper cites.
Burns, A., Arsan, D., Agrawal, S., Kumar, R., Saenko, K., Plummer, B.A.: A dataset for interactive vision-language navigation with unknown command feasibility. In: European Conference on Computer Vision. pp. 312–328. Springer (2022)
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Bavishi, R., Elsen, E., Hawthorne, C., Nye, M., Odena, A., Somani, A., Taşırlar, S.: Introducing our multimodal models (2023), https://www.adept.ai/blog/fuyu-8b
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Huang, S., Dong, L., Wang, W., Hao, Y., Singhal, S., Ma, S., Lv, T., Cui, L., Mohammed, O.K., Patra, B., Liu, Q., Aggarwal, K., Chi, Z., Bjorck, J., Chaudhary, V., Som, S., Song, X., Wei, F.: Language is not all you need: Aligning perception with language models (2023)
2023
Cited alongside, same era.
2023
2023
Later among the works it cites.
You, H., Zhang, H., Gan, Z., Du, X., Zhang, B., Wang, Z., Cao, L., Chang, S.F., Yang, Y.: Ferret: Refer and ground anything anywhere at any granularity (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Zhang, C., Yang, Z., Liu, J., Han, Y., Chen, X., Huang, Z., Fu, B., Yu, G.: Appagent: Multimodal agents as smartphone users (2023)
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Jiang, Y., Schoop, E., Swearngin, A., Nichols, J.: Iluvui: Instruction-tuned language-vision modeling of uis from machine conversations (2023)
2023
Cited alongside, same era.
Kim, G., Baldi, P., McAleer, S.: Language models can solve computer tasks (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Lee, K., Joshi, M., Turc, I.R., Hu, H., Liu, F., Eisenschlos, J.M., Khandelwal, U., Shaw, P., Chang, M.W., Toutanova, K.: Pix2struct: Screenshot parsing as pretraining for visual language understanding. In: ICML (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Li, G., Li, Y.: Spotlight: Mobile ui understanding using vision-language models with a focus (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Zheng, L., Chiang, W.L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E.P., Zhang, H., Gonzalez, J.E., Stoica, I.: Judging llm-as-a-judge with mt-bench and chatbot arena (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
Cheng, K., Sun, Q., Chu, Y., Xu, F., Li, Y., Zhang, J., Wu, Z.: Seeclick: Harnessing gui grounding for advanced visual gui agents (2024)
2024
Closest in time.
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., Su, Y.: Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Gao, P., Zhang, R., Liu, C., Qiu, L., Huang, S., Lin, W., Zhao, S., Geng, S., Lin, Z., Jin, P., Zhang, K., Shao, W., Xu, C., He, C., He, J., Shao, H., Lu, P., Li, H., Qiao, Y.: Sphinx-x: Scaling data and parameters for a family of multi-modal large language models (2024)
2024
Closest in time.
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., Lee, Y.J.: Llava-next: Improved reasoning, ocr, and world knowledge (January 2024), https://llava-vl.github.io/blog/2024-01-30-llava-next/
2024
Closest in time.
Liu, Y., Li, Z., Yang, B., Li, C., Yin, X., lin Liu, C., Jin, L., Bai, X.: On the hidden mystery of ocr in large multimodal models (2024)
2024
Closest in time.
2024
Closest in time.
OpenAI, :, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al.: Gpt-4 technical report (2024)
2024
Closest in time.
Shaw, P., Joshi, M., Cohan, J., Berant, J., Pasupat, P., Hu, H., Khandelwal, U., Lee, K., Toutanova, K.N.: From pixels to ui actions: Learning to follow instructions via graphical user interfaces. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Wang, J., Xu, H., Ye, J., Yan, M., Shen, W., Zhang, J., Huang, F., Sang, J.: Mobile-agent: Autonomous multi-modal mobile device agent with visual perception (2024)
2024
Closest in time.
2024
Closest in time.
Zheng, L., Wang, R., Wang, X., An, B.: Synapse: Trajectory-as-exemplar prompting with memory for computer control (2024)
2024
Closest in time.