Fetching the paper…
Reading the bibliography…
There is a growing interest in device-control systems that can interpret human natural language instructions and execute them on a digital device by directly controlling its user interface.
Weakly supervised action learning with RNN based fine-to-coarse modeling
A. Richard, H. Kuehne, and J. Gall · 2017
Earlier work this paper cites.
World of bits: An open-domain platform for web-based agents
T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang · 2017
Earlier work this paper cites.
Attention is all you need, 2017
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
E. Z. Liu, K. Guu, P. Pasupat, and P. Liang · 2018
Earlier work this paper cites.
Learning design semantics for mobile apps
T. F. Liu, M. Craft, J. Situ, E. Yumer, R. Mech, and R. Kumar · 2018
Earlier work this paper cites.
Learning to Navigate the Web
I. Gur, U. Rueckert, A. Faust, and D. Hakkani-Tur · 2019
Earlier work this paper cites.
Dom-q-net: Grounded rl on structured language, 2019
S. Jia, J. Kiros, and J. Ba · 2019
Earlier work this paper cites.
Language models are few-shot learners, 2020
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Unblind Your Apps: Predicting Natural-Language Labels for Mobile GUI Components by Deep Learning
J. Chen, C. Chen, Z. Xing, X. Xu, L. Zhu, G. Li, and J. Wang · 2020
Earlier work this paper cites.
Object detection for graphical user interface: Old fashioned or deep learning or a combination?
J. Chen, M. Xie, Z. Xing, C. Chen, X. Xu, L. Zhu, and G. Li · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2020
Earlier work this paper cites.
Haiku: Sonnet for JAX, 2020
T. Hennigan, T. Cai, T. Norman, and I. Babuschkin · 2020
Earlier work this paper cites.
Acme: A research framework for distributed reinforcement learning
M. W. Hoffman, B. Shahriari, J. Aslanides, G. Barth-Maron, N. Momchev, D. Sinopalnikov, P. Stańczyk, S. Ramos, A. Raichuk, D. Vincent, L. Hussenot, R. Dadashi, G. Dulac-Arnold, M. Orsini, A. Jacq, J. Ferret, N. Vieillard, S. K. S. Ghasemipour, S. Girgin, O. Pietquin, F. Behbahani, T. Norman, A. Abdolmaleki, A. Cassirer, F. Yang, K. Baumli, S. Henderson, A. Friesen, R. Haroun, A. Novikov, S. G. Colmenarejo, S. Cabi, C. Gulcehre, T. L. Paine, S. Srinivasan, A. Cowie, Z. Wang, B. Piot, and N. de Freitas · 2020
Earlier work this paper cites.
Mapping natural language instructions to mobile UI action sequences
Y. Li, J. He, X. Zhou, Y. Zhang, and J. Baldridge · 2020
Cited alongside, same era.
UIBert: Learning generic multimodal representations for UI understanding
C. Bai, X. Zang, Y. Xu, S. Sunkara, A. Rastogi, J. Chen, and B. A. y Arcas · 2021
Cited alongside, same era.
A. Burns, D. Arsan, S. Agrawal, R. Kumar, K. Saenko, and B. A. Plummer · 2021
Cited alongside, same era.
ActionBert: Leveraging User Actions for Semantic Understanding of User Interfaces
Z. He, S. Sunkara, X. Zang, Y. Xu, L. Liu, N. Wichers, G. Schubiner, R. B. Lee, and J. Chen · 2021
Cited alongside, same era.
Language conditioned imitation learning over unstructured data
C. Lynch and P. Sermanet · 2021
Cited alongside, same era.
Pix2struct: Screenshot parsing as pretraining for visual language understanding, 2022
K. Lee, M. Joshi, I. Turc, H. Hu, F. Liu, J. Eisenschlos, U. Khandelwal, P. Shaw, M.-W. Chang, and K. Toutanova · 2022
Later among the works it cites.
Interactive language: Talking to robots in real time, 2022
C. Lynch, A. Wahid, J. Tompson, T. Ding, J. Betker, R. Baruch, T. Armstrong, and P. Florence · 2022
Later among the works it cites.
Towards better semantic understanding of mobile interfaces
S. Sunkara, M. Wang, L. Liu, G. Baechler, Y.-C. Hsiao, J. Chen, A. Sharma, and J. W. W. Stout · 2022
Later among the works it cites.
Ugif: Ui grounded instruction following, 2022
S. G. Venkatesh, P. Talukdar, and S. Narayanan · 2022
Later among the works it cites.
Pali: A jointly-scaled multilingual language-image model, 2023
X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, A. Kolesnikov, J. Puigcerver, N. Ding, K. Rong, H. Akbari, G. Mishra, L. Xue, A. Thapliyal, J. Bradbury, W. Kuo, M. Seyedhosseini, C. Jia, B. K. Ayan, C. Riquelme, A. Steiner, A. Angelova, X. Zhai, N. Houlsby, and R. Soricut · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Appbuddy: Learning to accomplish tasks in mobile apps via reinforcement learning, 2021
M. Shvo, Z. Hu, R. T. Icarte, I. Mohomed, A. Jepson, and S. A. McIlraith · 2021
Cited alongside, same era.
Androidenv: A reinforcement learning platform for android, 2021
D. Toyama, P. Hamel, A. Gergely, G. Comanici, A. Glaese, Z. Ahmed, T. Jackson, S. Mourad, and D. Precup · 2021
Cited alongside, same era.
Google is trying to limit what apps can use an Accessibility Service (again), 2021
XDA · 2021
Cited alongside, same era.
Screen Recognition: Creating Accessibility Metadata for Mobile Applications from Pixels
X. Zhang, L. de Greef, A. Swearngin, S. White, K. Murray, L. Yu, Q. Shan, J. Nichols, J. Wu, C. Fleizach, A. Everitt, and J. P. Bigham · 2021
Cited alongside, same era.
ACT-1: Transformer for Actions, 2022
Adept · 2022
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning, 2022
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan · 2022
Cited alongside, same era.
Lexi: Self-supervised learning of the UI language
P. Banerjee, S. Mahajan, K. Arora, C. Baral, and O. Riva · 2022
Cited alongside, same era.
Mind2Web: Towards a generalist agent for the web, 2023
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su · 2023
Closest in time.
Multimodal web navigation with instruction-finetuned foundation models, 2023
H. Furuta, O. Nachum, K.-H. Lee, Y. Matsuo, S. S. Gu, and I. Gur · 2023
Closest in time.
Understanding html with large language models, 2023
I. Gur, O. Nachum, Y. Miao, M. Safdari, A. Huang, A. Chowdhery, S. Narang, N. Fiedel, and A. Faust · 2023
Closest in time.
Language models can solve computer tasks, 2023
G. Kim, P. Baldi, and S. McAleer · 2023
Closest in time.
Gorilla: Large language model connected with massive apis, 2023
S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez · 2023
Closest in time.
Toolllm: Facilitating large language models to master 16000+ real-world apis, 2023
Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, S. Zhao, R. Tian, R. Xie, J. Zhou, M. Gerstein, D. Li, Z. Liu, and M. Sun · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools, 2023
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Closest in time.
Enabling conversational interaction with mobile ui using large language models
B. Wang, G. Li, and Y. Li · 2023
Closest in time.
Chain-of-thought prompting elicits reasoning in large language models, 2023
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou · 2023
Closest in time.