Fetching the paper…
Reading the bibliography…
Visually-situated language is ubiquitous -- sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R. S., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., and Farhadi, A · 2016
Earlier work this paper cites.
Image-to-markup generation with coarse-to-fine attention
Deng, Y., Kanervisto, A., Ling, J., and Rush, A. M · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Learning design semantics for mobile apps
Liu, T. F., Craft, M., Situ, J., Yumer, E., Mech, R., and Kumar, R · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
Mishra, A., Shekhar, S., Singh, A. K., and Chakraborty, A · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Fixing the train-test resolution discrepancy
Touvron, H., Vedaldi, A., Douze, M., and Jegou, H · 2019
Earlier work this paper cites.
Unblind your apps: Predicting natural-language labels for mobile gui components by deep learning
Chen, J., Chen, C., Xing, Z., Xu, X., Zhu, L., Li, G., and Wang, J · 2020
Earlier work this paper cites.
Understanding tables with intermediate pre-training
Eisenschlos, J., Krichene, S., and Müller, T · 2020
Earlier work this paper cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2020
Earlier work this paper cites.
Mapping natural language instructions to mobile UI action sequences
Li, Y., He, J., Zhou, X., Zhang, Y., and Baldridge, J · 2020
Cited alongside, same era.
Widget captioning: Generating natural language description for mobile user interface elements
Li, Y., Li, G., He, L., Zheng, J., Li, H., and Guan, Z · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Textcaps: a dataset for image captioningwith reading comprehension
Sidorov, O., Hu, R., Rohrbach, M., and Singh, A · 2020
Cited alongside, same era.
Multimodal attention with image text spatial relationship for ocr-based image captioning
Wang, J., Tang, J., and Luo, J · 2020
Cited alongside, same era.
DocFormer: End-to-end Transformer for document understanding
Tap: Text-aware pre-training for text-vqa and text-caption
Yang, Z., Lu, Y., Wang, J., Yin, X., Florencio, D., Wang, L., Zhang, C., Zhang, L., and Luo, J · 2021
Later among the works it cites.
Screen recognition: Creating accessibility metadata for mobile applications from pixels
Zhang, X., de Greef, L., Swearngin, A., White, S., Murray, K., Yu, L., Shan, Q., Nichols, J., Wu, J., Fleizach, C., et al · 2021
Later among the works it cites.
HTLM: hyper-text pre-training and prompting of language models
Aghajanyan, A., Okhonko, D., Lewis, M., Joshi, M., Xu, H., Ghosh, G., and Zettlemoyer, L · 2022
Closest in time.
Latr: Layout-aware transformer for scene-text vqa
Biten, A. F., Litman, R., Xie, Y., Appalaraju, S., and Manmatha, R · 2022
Closest in time.
End-to-end document recognition and understanding with Dessurt
Davis, B., Morse, B., Price, B., Tensmeyer, C., Wigington, C., and Morariu, V · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Appalaraju, S., Jasani, B., Kota, B. U., Xie, Y., and Manmatha, R · 2021
Cited alongside, same era.
Uibert: Learning generic multimodal representations for ui understanding
Bai, C., Zang, X., Xu, Y., Sunkara, S., Rastogi, A., Chen, J., and Agüera y Arcas, B · 2021
Cited alongside, same era.
Due: End-to-end document understanding benchmark
Borchmann, Ł., Pietruszka, M., Stanislawek, T., Jurkiewicz, D., Turski, M., Szyndler, K., and Graliński, F · 2021
Cited alongside, same era.
Pix2seq: A language modeling framework for object detection
Chen, T., Saxena, S., Li, L., Fleet, D. J., and Hinton, G · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
Actionbert: Leveraging user actions for semantic understanding of user interfaces
He, Z., Sunkara, S., Zang, X., Xu, Y., Liu, L., Wichers, N., Schubiner, G., Lee, R., and Chen, J · 2021
Cited alongside, same era.
StructuralLM: Structural pre-training for form understanding
Li, C., Bi, B., Yan, M., Wang, W., Huang, S., Huang, F., and Si, L · 2021
Cited alongside, same era.
LayoutLMv3: Pre-training for document ai with unified text and image masking
Huang, Y., Lv, T., Cui, L., Lu, Y., and Wei, F · 2022
Closest in time.
Donut: Document understanding transformer without OCR
Kim, G., Hong, T., Yim, M., Park, J., Yim, J., Hwang, W., Yun, S., Han, D., and Park, S · 2022
Closest in time.
MarkupLM: Pre-training of text and markup language for visually rich document understanding
Li, J., Xu, Y., Cui, L., and Wei, F · 2022
Closest in time.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Masry, A., Long, D., Tan, J. Q., Joty, S., and Hoque, E · 2022
Closest in time.
Infographicvqa
Mathew, M., Bagal, V., Tito, R., Karatzas, D., Valveny, E., and Jawahar, C · 2022
Closest in time.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N., and Lewis, M · 2022
Closest in time.
Language modelling with pixels
Rust, P., Lotz, J. F., Bugliarello, E., Salesky, E., de Lhoneux, M., and Elliott, D · 2022
Closest in time.
Unifying vision, text, and layout for universal document processing
Tang, Z., Yang, Z., Wang, G., Fang, Y., Liu, Y., Zhu, C., Zeng, M., Zhang, C., and Bansal, M · 2022
Closest in time.
Webformer: The web-page transformer for structure information extraction
Wang, Q., Fang, Y., Ravula, A., Feng, F., Quan, X., and Liu, D · 2022
Closest in time.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W · 2022
Closest in time.
Spotlight: Mobile UI understanding using vision-language models with a focus
Li, G. and Li, Y · 2023
Closest in time.