Fetching the paper…
Reading the bibliography…
We propose SelfDoc, a task-agnostic pre-training framework for document image understanding.
Artificial neural networks for document analysis and recognition
Simone Marinai, Marco Gori, and Giovanni Soda · 2005
Earlier work this paper cites.
Learning nongenerative grammatical models for document analysis
Michael Shilman, Percy Liang, and Paul Viola · 2005
Earlier work this paper cites.
k-means++: The advantages of careful seeding
David Arthur and Sergei Vassilvitskii · 2006
Earlier work this paper cites.
Building a test collection for complex document information processing
D. Lewis, G. Agam, S. Argamon, O. Frieder, D. Grossman, and J. Heard · 2006
Earlier work this paper cites.
Tesseract: An open-source optical character recognition engine
Anthony Kay · 2007
Earlier work this paper cites.
Evaluation of svm, mlp and gmm classifiers for layout analysis of historical documents
Hao Wei, Micheal Baechler, Fouad Slimane, and Rolf Ingold · 2013
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Evaluation of deep convolutional nets for document image classification and retrieval
Adam W Harley, Alex Ufkes, and Konstantinos G Derpanis · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Earlier work this paper cites.
Learning to extract semantic structure from documents using multimodal fully convolutional neural networks
Xiao Yang, Ersin Yumer, Paul Asente, Mike Kraley, Daniel Kifer, and C Lee Giles · 2017
Cited alongside, same era.
Chargrid: Towards understanding 2d documents
Anoop R Katti, Christian Reisswig, Cordula Guder, Sebastian Brarda, Steffen Bickel, Johannes Höhne, and Jean Baptiste Faddoul · 2018
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
https://rrc.cvc.uab.es/?ch=13&com=introduction
Icdar 2019 robust reading challenge on scanned receipts ocr and information extraction · 2019
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Later among the works it cites.
Detectron2
Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick · 2019
Later among the works it cites.
Publaynet: largest dataset ever for document layout analysis
Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes · 2019
Later among the works it cites.
https://github.com/microsoft/unilm/issues/250 , 2020
Iit-cdip test collection is unavailable? · 2020
Later among the works it cites.
arxiv bulk data access
Arxiv · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Modular multimodal architecture for document classification
Tyler Dauphinee, Nikunj Patel, and Mohammad Rashidi · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Multimodal document image classification
Rajiv Jain and Curtis Wigington · 2019
Cited alongside, same era.
Funsd: A dataset for form understanding in noisy scanned documents
G. Jaume, H. Kemal Ekenel, and J. Thiran · 2019
Cited alongside, same era.
Graph convolution for multimodal information extraction from visually rich documents
Xiaojing Liu, Feiyu Gao, Qiong Zhang, and Huasha Zhao · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Later among the works it cites.
Self-supervised representation learning on document images
Adrian Cosma, Mihai Ghidoveanu, Michael Panaitescu-Liess, and Marius Popescu · 2020
Later among the works it cites.
Self-supervised relationship probing
Jiuxiang Gu, Jason Kuen, Shafiq Joty, Jianfei Cai, Vlad Morariu, Handong Zhao, and Tong Sun · 2020
Later among the works it cites.
Subhojeet Pramanik, Shashank Mujumdar, and Hima Patel · 2020
Later among the works it cites.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou · 2020
Later among the works it cites.
Docvqa: A dataset for vqa on document images
Minesh Mathew, Dimosthenis Karatzas, R Manmatha, and CV Jawahar · 2021
Closest in time.