Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Post-ocr parsing: building simple and robust parser via bio tagging
Wonseok Hwang, Seonghyeon Kim, Jinyeong Yim, Minjoon Seo, Seunghyun Park, Sungrae Park, Junyeop Lee, Bado Lee, and Hwalsuk Lee. 2019 · 2019
Cited alongside, same era.
Funsd: A dataset for form understanding in noisy scanned documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019 · 2019
Cited alongside, same era.
Graph convolution for multimodal information extraction from visually rich documents
Xiaojing Liu, Feiyu Gao, Qiong Zhang, and Huasha Zhao. 2019 · 2019
Cited alongside, same era.
OpenCeres: When open information extraction meets the semi-structured web
Colin Lockard, Prashant Shiralkar, and Xin Luna Dong. 2019 · 2019
Cited alongside, same era.
Cord: A consolidated receipt dataset for post-ocr parsing
Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee. 2019 · 2019
Cited alongside, same era.
GraphIE: A graph-based framework for information extraction
Yujie Qian, Enrico Santus, Zhijing Jin, Jiang Guo, and Regina Barzilay. 2019 · 2019
Cited alongside, same era.
Docparser: Hierarchical structure parsing of document renderings
Original
Johannes Rausch, Octavio Martinez, Fabian Bissig, Ce Zhang, and Stefan Feuerriegel. 2019 · 2019
Cited alongside, same era.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2019 · 2019
Cited alongside, same era.
CUTIE: learning to understand documents with convolutional universal text information extractor
Xiaohui Zhao, Zhuo Wu, and Xiaoguang Wang. 2019 · 2019
Cited alongside, same era.
Image-based table recognition: data, model, and evaluation
Original
Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno-Yepes. 2019 · 2019
Cited alongside, same era.
LAMBERT: layout-aware language modeling using BERT for information extraction
Original
Lukasz Garncarek, Rafal Powalski, Tomasz Stanislawek, Bartosz Topolski, Piotr Halama, and Filip Gralinski. 2020 · 2020
Cited alongside, same era.