Li, Y., He, J., Zhou, X., Zhang, Y., Baldridge, J.: Mapping natural language instructions to mobile UI action sequences. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 8198–8210. Association for Computational Linguistics, Online (Jul 2020). https://doi.org/10.18653/v1/2020.acl-main.729, https://www.aclweb.org/anthology/2020.acl-main.729
2020
Later among the works it cites.
Zhu, F., Zhu, Y., Chang, X., Liang, X.: Vision-language navigation with self-supervised auxiliary reasoning tasks. In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10009–10019 (2020). https://doi.org/10.1109/CVPR42600.2020.01003
2020
Later among the works it cites.
Appalaraju, S., Jasani, B., Kota, B.U., Xie, Y., Manmatha, R.: Docformer: End-to-end transformer for document understanding. In: 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.
Blukis, V., Paxton, C., Fox, D., Garg, A., Artzi, Y.: A persistent spatial semantic representation for high-level natural language instruction execution (2021)
2021
Later among the works it cites.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An image is worth 16x16 words: Transformers for image recognition at scale (2021)
2021
Later among the works it cites.
Irshad, M.Z., Ma, C.Y., Kira, Z.: Hierarchical cross-modal agent for robotics vision-and-language navigation. In: Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) (2021), https://arxiv.org/abs/2104.10674
Original
2021
Later among the works it cites.
Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q.V., Sung, Y., Li, Z., Duerig, T.: Scaling up visual and vision-language representation learning with noisy text supervision (2021)
2021
Later among the works it cites.
Li, P., Gu, J., Kuen, J., Morariu, V.I., Zhao, H., Jain, R., Manjunatha, V., Liu, H.: Selfdoc: Self-supervised document representation learning. In: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021)
2021
Later among the works it cites.
Li, T.J.J., Mitchell, T.M., Myers, B.A.: Demonstration + Natural Language: Multimodal Interfaces for GUI-Based Interactive Task Learning Agents, pp. 495–537. Springer International Publishing, Cham (2021)
2021
Later among the works it cites.
Li, T.J.J., Popowski, L., Mitchell, T.M., Myers, B.A.: Screen2vec: Semantic embedding of gui screens and gui components. In: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. CHI ’21 (2021)
2021
Later among the works it cites.
Li, Y., Li, G., Zhou, X., Dehghani, M., Gritsenko, A.A.: VUT: versatile UI transformer for multi-modal multi-task user interface modeling. CoRR abs/2112.05692
Original
2021
Later among the works it cites.
Min, S.Y., Chaplot, D.S., Ravikumar, P., Bisk, Y., Salakhutdinov, R.: Film: Following instructions in language with modular methods (2021)
2021
Later among the works it cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. CoRR abs/2103.00020
Original
2021
Later among the works it cites.
Singh, K.P., Bhambri, S., Kim, B., Mottaghi, R., Choi, J.: Factorizing perception and policy for interactive instruction following. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.
Yamaguchi, K.: Canvasvae: Learning to generate vector graphic documents. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2021)
2021
Later among the works it cites.