Fetching the paper…
Reading the bibliography…
Large pre-trained language models achieve state-of-the-art results when fine-tuned on downstream NLP tasks.
FUNSD: A dataset for form understanding in noisy scanned documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019 · 1905
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019a · 1907
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Structbert: Incorporating language structures into pre-training for deep language understanding
Wei Wang, Bin Bi, Ming Yan, Chen Wu, Zuyi Bao, Jiangnan Xia, Liwei Peng, and Luo Si. 2019 · 1908
Earlier work this paper cites.
Modular multimodal architecture for document classification
Tyler Dauphinee, Nikunj Patel, and Mohammad Rashidi. 2019 · 1912
Earlier work this paper cites.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2019 · 1912
Earlier work this paper cites.
Artificial neural networks for document analysis and recognition
S. Marinai, M. Gori, and G. Soda. 2005 · 2005
Earlier work this paper cites.
Learning nongenerative grammatical models for document analysis
M. Shilman, P. Liang, and P. Viola. 2005 · 2005
Earlier work this paper cites.
Building a test collection for complex document information processing
D. Lewis, G. Agam, S. Argamon, O. Frieder, D. Grossman, and J. Heard. 2006 · 2006
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
M. Mathew, Dimosthenis Karatzas, R. Manmatha, and C. Jawahar. 2020 · 2007
Cited alongside, same era.
Evaluation of svm, mlp and gmm classifiers for layout analysis of historical documents
H. Wei, M. Baechler, F. Slimane, and R. Ingold. 2013 · 2013
Cited alongside, same era.
Evaluation of deep convolutional nets for document image classification and retrieval
Adam W Harley, Alex Ufkes, and Konstantinos G Derpanis. 2015 · 2015
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Document image classification with intra-domain transfer learning and stacked generalization of deep convolutional neural networks
Arindam Das, Saikat Roy, Ujjwal Bhattacharya, and Swapan K Parui. 2018 · 2018
Later among the works it cites.
Chargrid: Towards understanding 2d documents
Anoop Raveendra Katti, Christian Reisswig, Cordula Guder, Sebastian Brarda, Steffen Bickel, Johannes Höhne, and Jean Baptiste Faddoul. 2018 · 2018
Later among the works it cites.
Multi-granularity hierarchical attention fusion networks for reading comprehension and question answering
Wei Wang, Ming Yan, and Chen Wu. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Deterministic routing between layout abstractions for multi-scale classification of visually rich documents
Ritesh Sarkhel and Arnab Nandi. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cutting the error by half: Investigation of very deep cnn and advanced training strategies for document image classification
Muhammad Zeshan Afzal, Andreas Kölsch, Sheraz Ahmed, and Marcus Liwicki. 2017 · 2017
Cited alongside, same era.
Fast cnn-based document layout analysis
D. A. Borges Oliveira and M. P. Viana. 2017 · 2017
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander Alemi. 2017 · 2017
Cited alongside, same era.
Learning to extract semantic structure from documents using multimodal fully convolutional neural networks
Xiao Yang, Ersin Yumer, Paul Asente, Mike Kraley, Daniel Kifer, and C Lee Giles. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Visual detection with context for document layout analysis
Carlos Soto and Shinjae Yoo. 2019 · 2019
Later among the works it cites.
Trie: End-to-end text reading and information extraction for document understanding
Peng Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu, Liang Qiao, Yi Niu, and Fei Wu. 2020 · 2020
Later among the works it cites.
{BROS}: A pre-trained language model for understanding texts in document
Teakgyu Hong, DongHyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park. 2021 · 2021
Closest in time.