Fetching the paper…
Reading the bibliography…
Recent approaches in literature have exploited the multi-modal information in documents (text, layout, image) to serve specific downstream document tasks.
One-class svms for document classification
Larry M Manevitz and Malik Yousef. 2001 · 2001
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan. 2003 · 2003
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Evaluation of deep convolutional nets for document image classification and retrieval
Adam W Harley, Alex Ufkes, and Konstantinos G Derpanis. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
R-fcn: Object detection via region-based fully convolutional networks
Jifeng Dai, Yi Li, Kaiming He, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Cutting the error by half: Investigation of very deep cnn and advanced training strategies for document image classification
Muhammad Zeshan Afzal, Andreas Kölsch, Sheraz Ahmed, and Marcus Liwicki. 2017 · 2017
Earlier work this paper cites.
Self-supervised learning of visual features through embedding images into text topic spaces
Lluis Gomez, Yash Patel, Marçal Rusiñol, Dimosthenis Karatzas, and CV Jawahar. 2017 · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2017 · 2017
Earlier work this paper cites.
Deepdesrt: Deep learning for detection and structure recognition of tables in document images
Sebastian Schreiber, Stefan Agne, Ivo Wolf, Andreas Dengel, and Sheraz Ahmed. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Learning to extract semantic structure from documents using multimodal fully convolutional neural networks
Xiao Yang, Ersin Yumer, Paul Asente, Mike Kraley, Daniel Kifer, and C Lee Giles. 2017 · 2017
Cited alongside, same era.
Document image classification with intra-domain transfer learning and stacked generalization of deep convolutional neural networks
Arindam Das, Saikat Roy, Ujjwal Bhattacharya, and Swapan K Parui. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Chargrid: Towards understanding 2d documents
Anoop R Katti, Christian Reisswig, Cordula Guder, Sebastian Brarda, Steffen Bickel, Johannes Höhne, and Jean Baptiste Faddoul. 2018 · 2018
Cited alongside, same era.
Tablenet: Deep learning model for end-to-end table detection and tabular data extraction from scanned document images
Shubham Singh Paliwal, D Vishwanath, Rohit Rahul, Monika Sharma, and Lovekesh Vig. 2019 · 2019
Later among the works it cites.
Deterministic routing between layout abstractions for multi-scale classification of visually rich documents
Ritesh Sarkhel and Arnab Nandi. 2019 · 2019
Later among the works it cites.
Visual detection with context for document layout analysis
Carlos Soto and Shinjae Yoo. 2019 · 2019
Later among the works it cites.
Cutie: Learning to understand documents with convolutional universal text information extractor
Xiaohui Zhao, Endi Niu, Zhuo Wu, and Xiaoguang Wang. 2019 · 2019
Later among the works it cites.
Publaynet: largest dataset ever for document layout analysis
Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Cited alongside, same era.
Modular multimodal architecture for document classification
Tyler Dauphinee, Nikunj Patel, Mohammad Rashidi, and . 2019 · 2019
Cited alongside, same era.
Workshop on document intelligence at neurips 2019
DI. 2019 · 2019
Cited alongside, same era.
Icdar 2019 competition on table detection and recognition (ctdar)
Liangcai Gao, Yilun Huang, Hervé Déjean, Jean-Luc Meunier, Qinqin Yan, Yu Fang, Florian Kleber, and Eva Lang. 2019 · 2019
Cited alongside, same era.
Funsd: A dataset for form understanding in noisy scanned documents
Jean-Philippe Thiran Guillaume Jaume, Hazim Kemal Ekenel. 2019 · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Arxiv. 2020 · 2020
Closest in time.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020 · 2020
Closest in time.
Representation learning for information extraction from form-like documents
Bodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt, Qi Zhao, and Marc Najork. 2020 · 2020
Closest in time.
Cascadetabnet: An approach for end to end table detection and structure recognition from image-based documents
Devashish Prasad, Ayan Gadpal, Kshitij Kapadni, Manish Visave, and Kavita Sultanpure. 2020 · 2020
Closest in time.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020 · 2020
Closest in time.
Python wrapper for google’s tesseract-ocr engine
Python Tesseract. 2021 · 2021
Closest in time.