Fetching the paper…
Reading the bibliography…
Multimodal pre-training with text, layout, and image has made significant progress for Visually Rich Document Understanding (VRDU), especially the fixed-layout documents such as scanned document images.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke S. Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Building a test collection for complex document information processing
D. Lewis, G. Agam, S. Argamon, O. Frieder, D. Grossman, and J. Heard. 2006 · 2006
Earlier work this paper cites.
Bootstrapping information extraction from semi-structured web pages
Andrew Carlson and Charles Schafer. 2008 · 2008
Earlier work this paper cites.
From one tree to a forest: a unified solution for structured web data extraction
Qiang Hao, Rui Cai, Yanwei Pang, and Lei Zhang. 2011 · 2011
Earlier work this paper cites.
Evaluation of deep convolutional nets for document image classification and retrieval
Adam W Harley, Alex Ufkes, and Konstantinos G Derpanis. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Icdar2019 competition on scanned receipt ocr and information extraction
Z. Huang, K. Chen, J. He, X. Bai, D. Karatzas, S. Lu, and C. V. Jawahar. 2019 · 2019
Cited alongside, same era.
Funsd: A dataset for form understanding in noisy scanned documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019 · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
CORD: A consolidated receipt dataset for post-OCR parsing
Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee. 2019 · 2019
Cited alongside, same era.
Leveraging HTML in free text web named entity recognition
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Later among the works it cites.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020 · 2020
Later among the works it cites.
Docformer: End-to-end transformer for document understanding
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R. Manmatha. 2021 · 2021
Closest in time.
WebSRC: A dataset for web-based structural reading comprehension
Xingyu Chen, Zihan Zhao, Lu Chen, JiaBao Ji, Danyang Zhang, Ao Luo, Yuxuan Xiong, and Kai Yu. 2021 · 2021
Closest in time.
BROS: A pre-trained language model for understanding texts in document
Teakgyu Hong, DongHyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Colin Ashby and David Weir. 2020 · 2020
Cited alongside, same era.
Kleister: A novel task for information extraction involving long documents with complex layout
Filip Graliński, Tomasz Stanisławek, Anna Wróblewska, Dawid Lipiński, Agnieszka Kaliska, Paulina Rosalska, Bartosz Topolski, and Przemysław Biecek. 2020 · 2020
Cited alongside, same era.
Freedom: A transferable neural architecture for structured information extraction on web documents
Bill Yuchen Lin, Ying Sheng, Nguyen Vo, and Sandeep Tata. 2020 · 2020
Cited alongside, same era.
Docvqa: A dataset for vqa on document images
Minesh Mathew, Dimosthenis Karatzas, R. Manmatha, and C. V. Jawahar. 2020 · 2020
Cited alongside, same era.
Subhojeet Pramanik, Shashank Mujumdar, and Hima Patel. 2020 · 2020
Cited alongside, same era.
StructuralLM: Structural pre-training for form understanding
Chenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si. 2021a
Cited in the paper.
Selfdoc: Self-supervised document representation learning
Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, and Hongfu Liu. 2021b
Cited in the paper.
Going full-tilt boogie on document understanding with text-image-layout transformer
Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz, Tomasz Dwojak, Michał Pietruszka, and Gabriela Pałka. 2021 · 2021
Closest in time.
Lampret: Layout-aware multimodal pretraining for document understanding
Te-Lin Wu, Cheng Li, Mingyang Zhang, Tao Chen, Spurthi Amba Hombaiah, and Michael Bendersky. 2021 · 2021
Closest in time.
Webke: Knowledge extraction from semi-structured web with pre-trained markup language model
Chenhao Xie, Wenhao Huang, Jiaqing Liang, Chengsong Huang, and Yanghua Xiao. 2021 · 2021
Closest in time.
Simplified dom trees for transferable attribute extraction from the web
Yichao Zhou, Ying Sheng, Nguyen Vo, Nick Edmonds, and Sandeep Tata. 2021 · 2021
Closest in time.
Lambert: Layout-aware (language) modeling for information extraction
Łukasz Garncarek, Rafał Powalski, Tomasz Stanisławek, Bartosz Topolski, Piotr Halama, Michał Turski, and Filip Graliński. 2021 · 2021
Closest in time.