Fetching the paper…
Reading the bibliography…
Key information extraction (KIE) from document images requires understanding the contextual and spatial semantics of texts in two-dimensional (2D) space.
Complicated table structure recognition
Chi, Z.; Huang, H.; Xu, H.-D.; Yu, H.; Yin, W.; and Mao, X.-L. 2019 · 1908
Earlier work this paper cites.
Building a test collection for complex document information processing
Lewis, D.; Agam, G.; Argamon, S.; Frieder, O.; Grossman, D.; and Heard, J. 2006 · 2006
Earlier work this paper cites.
The significance of reading order in document recognition and its evaluation
Clausner, C.; Pletschacher, S.; and Antonacopoulos, A. 2013 · 2013
Earlier work this paper cites.
Evaluation of deep convolutional nets for document image classification and retrieval
Harley, A. W.; Ufkes, A.; and Derpanis, K. G. 2015 · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J. G.; Le, Q.; and Salakhutdinov, R. 2019 · 2019
Earlier work this paper cites.
BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding
Denk, T. I.; and Reisswig, C. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
ICDAR2019 competition on scanned receipt ocr and information extraction
Huang, Z.; Chen, K.; He, J.; Bai, X.; Karatzas, D.; Lu, S.; and Jawahar, C. 2019 · 2019
Cited alongside, same era.
Post-OCR parsing: building simple and robust parser via BIO tagging
Hwang, W.; Kim, S.; Seo, M.; Yim, J.; Park, S.; Park, S.; Lee, J.; Lee, B.; and Lee, H. 2019 · 2019
Cited alongside, same era.
FUNSD: A dataset for form understanding in noisy scanned documents
Jaume, G.; Ekenel, H. K.; and Thiran, J.-P. 2019 · 2019
Cited alongside, same era.
Graph Convolution for Multimodal Information Extraction from Visually Rich Documents
Liu, X.; Gao, F.; Zhang, Q.; and Zhao, H. 2019 · 2019
Cited alongside, same era.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Cited alongside, same era.
CORD: A Consolidated Receipt Dataset for Post-OCR Parsing
Park, S.; Shin, S.; Lee, B.; Lee, J.; Surh, J.; Seo, M.; and Lee, H. 2019 · 2019
An End-to-End OCR Text Re-organization Sequence Learning for Rich-text Detail Image Comprehension
Li, L.; Gao, F.; Bu, J.; Wang, Y.; Yu, Z.; and Zheng, Q. 2020 · 2020
Later among the works it cites.
LayoutLM: Pre-training of text and layout for document image understanding
Xu, Y.; Li, M.; Cui, L.; Huang, S.; Wei, F.; and Zhou, M. 2020 · 2020
Later among the works it cites.
DocFormer: End-to-End Transformer for Document Understanding
Appalaraju, S.; Jasani, B.; Kota, B. U.; Xie, Y.; and Manmatha, R. 2021 · 2021
Closest in time.
Spatial Dependency Parsing for Semi-Structured Document Information Extraction
Hwang, W.; Yim, J.; Park, S.; Yang, S.; and Seo, M. 2021 · 2021
Closest in time.
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer
Powalski, R.; Borchmann, Ł.; Jurkiewicz, D.; Dwojak, T.; Pietruszka, M.; and Pałka, G. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
SpanBERT: Improving pre-training by representing and predicting spans
Joshi, M.; Chen, D.; Liu, Y.; Weld, D. S.; Zettlemoyer, L.; and Levy, O. 2020 · 2020
Cited alongside, same era.
StructuralLM: Structural Pre-training for Form Understanding
Li, C.; Bi, B.; Yan, M.; Wang, W.; Huang, S.; Huang, F.; and Si, L. 2021a
Cited in the paper.
SelfDoc: Self-Supervised Document Representation Learning
Li, P.; Gu, J.; Kuen, J.; Morariu, V. I.; Zhao, H.; Jain, R.; Manjunatha, V.; and Liu, H. 2021b
Cited in the paper.
StrucTexT: Structured Text Understanding with Multi-Modal Transformers
Li, Y.; Qian, Y.; Yu, Y.; Qin, X.; Zhang, C.; Liu, Y.; Yao, K.; Han, J.; Liu, J.; and Ding, E. 2021c
Cited in the paper.
Closest in time.
LayoutReader: Pre-training of Text and Layout for Reading Order Detection
Wang, Z.; Xu, Y.; Cui, L.; Shang, J.; and Wei, F. 2021 · 2021
Closest in time.
LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding
Xu, Y.; Xu, Y.; Lv, T.; Cui, L.; Wei, F.; Wang, G.; Lu, Y.; Florencio, D.; Zhang, C.; Che, W.; et al. 2021 · 2021
Closest in time.