Fetching the paper…
Reading the bibliography…
For decades, human-computer interaction has fundamentally been manual.
Papineni, K., Roukos, S., Ward, T., Zhu, W.J.: Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting of the Association for Computational Linguistics. pp. 311–318 (2002)
2002
Earlier work this paper cites.
Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P.: Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13
2004
Earlier work this paper cites.
Deka, B., Huang, Z., Franzen, C., Hibschman, J., Afergan, D., Li, Y., Nichols, J., Kumar, R.: Rico: A mobile app dataset for building data-driven design applications. In: Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology. p. 845–854. UIST ’17, Association for Computing Machinery, New York, NY, USA (2017). https://doi.org/10.1145/3126594.3126651, https://doi.org/10.1145/3126594.3126651
2017
Earlier work this paper cites.
Shi, T., Karpathy, A., Fan, L., Hernandez, J., Liang, P.: World of bits: An open-domain platform for web-based agents. In: International Conference on Machine Learning. pp. 3135–3144. PMLR (2017)
2017
Earlier work this paper cites.
Gur, I., Rueckert, U., Faust, A., Hakkani-Tur, D.: Learning to navigate the web. In: International Conference on Learning Representations (2018)
2018
Earlier work this paper cites.
Huang, Z., Zeng, Z., Liu, B., Fu, D., Fu, J.: Pixel-bert: Aligning image pixels with text by deep multi-modal transformers (2020)
2020
Earlier work this paper cites.
Li, Y., He, J., Zhou, X., Zhang, Y., Baldridge, J.: Mapping natural language instructions to mobile ui action sequences (2020)
2020
Earlier work this paper cites.
Li, Y., Li, G., He, L., Zheng, J., Li, H., Guan, Z.: Widget captioning: Generating natural language description for mobile user interface elements (2020)
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., Zhou, M.: Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. Advances in Neural Information Processing Systems 33
2020
Earlier work this paper cites.
Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., Artzi, Y.: Bertscore: Evaluating text generation with bert (2020)
2020
Earlier work this paper cites.
Bai, C., Zang, X., Xu, Y., Sunkara, S., Rastogi, A., Chen, J., y Arcas, B.A.: Uibert: Learning generic multimodal representations for ui understanding (2021)
2021
Earlier work this paper cites.
Chen, X., Zhao, Z., Chen, L., Ji, J., Zhang, D., Luo, A., Xiong, Y., Yu, K.: Websrc: A dataset for web-based structural reading comprehension. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. pp. 4173–4185 (2021)
2021
Earlier work this paper cites.
He, Z., Sunkara, S., Zang, X., Xu, Y., Liu, L., Wichers, N., Schubiner, G., Lee, R., Chen, J., y Arcas, B.A.: Actionbert: Leveraging user actions for semantic understanding of user interfaces (2021)
2021
Earlier work this paper cites.
Li, Y., Li, G., Zhou, X., Dehghani, M., Gritsenko, A.: Vut: Versatile ui transformer for multi-modal multi-task user interface modeling (2021)
2021
Earlier work this paper cites.
Wang, B., Li, G., Zhou, X., Chen, Z., Grossman, T., Li, Y.: Screen2words: Automatic mobile ui summarization with multimodal learning. In: The 34th Annual ACM Symposium on User Interface Software and Technology. pp. 498–510 (2021)
2021
Earlier work this paper cites.
Xu, N., Masling, S., Du, M., Campagna, G., Heck, L., Landay, J., Lam, M.: Grounding open-domain instructions to automate web support tasks. In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. pp. 1022–1032 (2021)
2021
Earlier work this paper cites.
Humphreys, P.C., Raposo, D., Pohlen, T., Thornton, G., Chhaparia, R., Muldal, A., Abramson, J., Georgiev, P., Santoro, A., Lillicrap, T.: A data-driven approach for learning to control computers. In: International Conference on Machine Learning. pp. 9466–9482. PMLR (2022)
2022
Earlier work this paper cites.
LeCun, Y.: A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27 (2022)
2022
Earlier work this paper cites.
Sun, L., Chen, X., Chen, L., Dai, T., Zhu, Z., Yu, K.: Meta-gui: Towards multi-modal conversational agents on mobile gui. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. pp. 6699–6712 (2022)
2022
Earlier work this paper cites.
Yao, S., Chen, H., Yang, J., Narasimhan, K.: Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems 35
2022
Cited alongside, same era.
Pyautogui: A cross-platform gui automation python module for human beings. https://github.com/asweigart/pyautogui (2023)
2023
Cited alongside, same era.
AlShikh, W., Daaboul, M., Goddard, K., Imel, B., Kamble, K., Kulkarni, P., Russak, M.: Becoming self-instruct: introducing early stopping criteria for minimal instruct tuning (2023)
2023
Cited alongside, same era.
Banerjee, P., Mahajan, S., Arora, K., Baral, C., Riva, O.: Lexi: Self-supervised learning of the ui language (2023)
2023
Cited alongside, same era.
Chiang, W.L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J.E., Stoica, I., Xing, E.P.: Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality (March 2023), https://lmsys.org/blog/2023-03-30-vicuna/
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X.E., Adi, Y., Liu, J., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C.C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., Synnaeve, G.: Code llama: Open foundation models for code (2023)
2023
Later among the works it cites.
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X.E., Adi, Y., Liu, J., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C.C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., Synnaeve, G.: Code llama: Open foundation models for code (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Dettmers, T., Pagnoni, A., Holtzman, A., Zettlemoyer, L.: Qlora: Efficient finetuning of quantized llms (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Gupta, T., Kembhavi, A.: Visual programming: Compositional visual reasoning without training. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 14953–14962 (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.Y., Dollár, P., Girshick, R.: Segment anything (2023)
2023
Cited alongside, same era.
Li, G., Li, Y.: Spotlight: Mobile ui understanding using vision-language models with a focus (2023)
2023
Cited alongside, same era.
2023
Later among the works it cites.
Surís, D., Menon, S., Vondrick, C.: Vipergpt: Visual inference via python execution for reasoning (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
team, W.E.: InstructPalmyra-30b : Instruct tuned Palmyra-Large model. https://dev.writer.com (2023)
2023
Later among the works it cites.
team, W.E.: Palmyra-base Parameter Autoregressive Language Model. https://dev.writer.com (2023)
2023
Later among the works it cites.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., Lample, G.: Llama: Open and efficient foundation language models (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Wang, X., Chen, G., Qian, G., Gao, P., Wei, X.Y., Wang, Y., Tian, Y., Gao, W.: Large-scale multi-modal pre-trained models: A comprehensive survey (2023)
2023
Later among the works it cites.
Yang, J., Zhang, H., Li, F., Zou, X., Li, C., Gao, J.: Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., Chen, E.: A survey on multimodal large language models (2023)
2023
Later among the works it cites.
Zhang, Z., Zhang, A.: You only look at screens: Multimodal chain-of-action agents (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., Faust, A.: A real-world webagent with planning, long context understanding, and program synthesis (2024)
2024
Closest in time.
Koh, J.Y., Lo, R., Jang, L., Duvvur, V., Lim, M.C., Huang, P.Y., Neubig, G., Zhou, S., Salakhutdinov, R., Fried, D.: Visualwebarena: Evaluating multimodal agents on realistic visual web tasks (2024)
2024
Closest in time.