Fetching the paper…
Reading the bibliography…
Alignment between image and text has shown promising improvements on patch-level pre-trained document image models.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Mixout: Effective regularization to finetune large-scale pretrained language models
Lee, C.; Cho, K.; and Kang, W. 2019 · 1909
Earlier work this paper cites.
Jiang, H.; He, P.; Chen, W.; Liu, X.; Gao, J.; and Zhao, T. 2019 · 1911
Earlier work this paper cites.
Don’t stop pretraining: adapt language models to domains and tasks
Gururangan, S.; Marasović, A.; Swayamdipta, S.; Lo, K.; Beltagy, I.; Downey, D.; and Smith, N. A. 2020 · 2004
Earlier work this paper cites.
Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval
Harley, A. W.; Ufkes, A.; and Derpanis, K. G. 2015 · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Howard, J.; and Ruder, S. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Dong, L.; Yang, N.; Wang, W.; Wei, F.; Liu, X.; Wang, Y.; Gao, J.; Zhou, M.; and Hon, H.-W. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Jaume, G.; Ekenel, H. K.; and Thiran, J.-P. 2019 · 2019
Earlier work this paper cites.
CORD: A Consolidated Receipt Dataset for Post-OCR Parsing
Park, S.; Shin, S.; Lee, B.; Lee, J.; Surh, J.; Seo, M.; and Lee, H. 2019 · 2019
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020 · 2020
Cited alongside, same era.
Layoutlm: Pre-training of text and layout for document image understanding
Xu, Y.; Li, M.; Cui, L.; Huang, S.; Wei, F.; and Zhou, M. 2020 · 2020
Cited alongside, same era.
DocFormer: End-to-End Transformer for Document Understanding
Appalaraju, S.; Jasani, B.; Kota, B. U.; Xie, Y.; and Manmatha, R. 2021 · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
An Empirical Study of Training End-to-End Vision-and-Language Transformers
Dou, Z.-Y.; Xu, Y.; Gan, Z.; Wang, J.; Wang, S.; Wang, L.; Zhu, C.; Liu, Z.; Zeng, M.; et al. 2021 · 2021
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021 · 2021
Later among the works it cites.
LAMPRET: Layout-Aware Multimodal PreTraining for Document Understanding
Wu, T.-L.; Li, C.; Zhang, M.; Chen, T.; Hombaiah, S. A.; and Bendersky, M. 2021 · 2021
Later among the works it cites.
Probing Inter-modality: Visual Parsing with Self-Attention for Vision-and-Language Pre-training
Xue, H.; Huang, Y.; Liu, B.; Peng, H.; Fu, J.; Li, H.; and Luo, J. 2021 · 2021
Later among the works it cites.
XYLayoutLM: Towards Layout-Aware Multimodal Networks For Visually-Rich Document Understanding
Gu, Z.; Meng, C.; Wang, K.; Lan, J.; Wang, W.; Gu, M.; and Zhang, L. 2022 · 2022
Closest in time.
BROS: A Pre-Trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents
Hong, T.; Kim, D.; Ji, M.; Hwang, W.; Nam, D.; and Park, S. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
LAMBERT: Layout-Aware Language Modeling for Information Extraction
Garncarek, Ł.; Powalski, R.; Stanisławek, T.; Topolski, B.; Halama, P.; Turski, M.; and Graliński, F. 2021 · 2021
Cited alongside, same era.
UniDoc: Unified Pretraining Framework for Document Understanding
Gu, J.; Kuen, J.; Morariu, V.; Zhao, H.; Jain, R.; Barmpalios, N.; Nenkova, A.; and Sun, T. 2021 · 2021
Cited alongside, same era.
Vilt: Vision-and-language transformer without convolution or region supervision
Kim, W.; Son, B.; and Kim, I. 2021 · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L.; and Liang, P. 2021 · 2021
Cited alongside, same era.
Docvqa: A dataset for vqa on document images
Mathew, M.; Karatzas, D.; and Jawahar, C. 2021 · 2021
Cited alongside, same era.
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer
Powalski, R.; Łukasz Borchmann; Jurkiewicz, D.; Dwojak, T.; Pietruszka, M.; and Pałka, G. 2021 · 2021
Cited alongside, same era.
StructuralLM: Structural Pre-training for Form Understanding
Li, C.; Bi, B.; Yan, M.; Wang, W.; Huang, S.; Huang, F.; and Si, L. 2021a
Cited in the paper.
Closest in time.
LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking
Huang, Y.; Lv, T.; Cui, L.; Lu, Y.; and Wei, F. 2022 · 2022
Closest in time.
OCR-Free Document Understanding Transformer
Kim, G.; Hong, T.; Yim, M.; Nam, J.; Park, J.; Yim, J.; Hwang, W.; Yun, S.; Han, D.; and Park, S. 2022 · 2022
Closest in time.
FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction
Lee, C.-Y.; Li, C.-L.; Dozat, T.; Perot, V.; Su, G.; Hua, N.; Ainslie, J.; Wang, R.; Fujii, Y.; and Pfister, T. 2022 · 2022
Closest in time.
DiT: Self-supervised Pre-training for Document Image Transformer
Li, J.; Xu, Y.; Lv, T.; Cui, L.; Zhang, C.; and Wei, F. 2022 · 2022
Closest in time.
LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding
Wang, J.; Jin, L.; and Ding, K. 2022 · 2022
Closest in time.
Vision-Language Pre-Training with Triple Contrastive Learning
Yang, J.; Duan, J.; Tran, S.; Xu, Y.; Chanda, S.; Chen, L.; Zeng, B.; Chilimbi, T.; and Huang, J. 2022 · 2022
Closest in time.