Fetching the paper…
Reading the bibliography…
Document AI, or Document Intelligence, is a relatively new research topic that refers to the techniques for automatically reading, understanding, and analyzing business documents.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Image-based table recognition: data, model, and evaluation
Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno Yepes · 1911
Earlier work this paper cites.
Document analysis system
Kwan Y. Wong, Richard G. Casey, and Friedrich M. Wahl · 1982
Earlier work this paper cites.
Hierarchical representation of optically scanned documents
George Nagy and Sharad C Seth · 1984
Earlier work this paper cites.
Image segmentation by shape-directed covers
Henry S Baird, Susan E Jones, and Steven J Fortune · 1990
Earlier work this paper cites.
An experimental page layout recognition system for office document automatic classification: an integrated approach for inductive generalization
Floriana Esposito, Donato Malerba, Giovanni Semeraro, Enrico Annese, and Giovanna Scafuro · 1990
Earlier work this paper cites.
A rule-based system for document image segmentation
James L Fisher, Stuart C Hinds, and Donald P D’Amato · 1990
Earlier work this paper cites.
Multiresolution morphological approach to document image analysis
Dan S Bloomberg · 1991
Earlier work this paper cites.
The document spectrum for page layout analysis
Lawrence O’Gorman · 1993
Earlier work this paper cites.
A hybrid page segmentation method
Masayuki Okamoto and Makoto Takahashi · 1993
Earlier work this paper cites.
Document image segmentation and text area ordering
Takashi Saitoh, Michiyoshi Tachikawa, and Toshifumi Yamaai · 1993
Earlier work this paper cites.
A trainable, single-pass algorithm for column segmentation
Don Sylwester and Sharad Seth · 1995
Earlier work this paper cites.
Solving the multiple instance problem with axis-parallel rectangles
Thomas G Dietterich, Richard H Lathrop, and Tomás Lozano-Pérez · 1997
Earlier work this paper cites.
Segmentation of page images using the area voronoi diagram
Koichi Kise, Akinori Sato, and Motoi Iwata · 1998
Earlier work this paper cites.
Improvement of zone content classification by using background analysis
Yalin Wang, Robert Haralick, and Ihsin T Phillips · 2000
Earlier work this paper cites.
Automatic table ground truth generation and a background-analysis-based table structure extraction method
Yalin Wangt, Ihsin T Phillipst, and Robert Haralick · 2001
Earlier work this paper cites.
Table detection via probability optimization
Yalin Wang, Ihsin T Phillips, and Robert M Haralick · 2002
Earlier work this paper cites.
Table extraction using conditional random fields
David Pinto, Andrew McCallum, Xing Wei, and W Bruce Croft · 2003
Earlier work this paper cites.
Text region extraction in a document image based on the delaunay tessellation
Yi Xiao and Hong Yan · 2003
Earlier work this paper cites.
Line separation for complex document images using fuzzy runlength
Zhixin Shi and Venu Govindaraju · 2004
Earlier work this paper cites.
Graph convolution for multimodal information extraction from visually rich documents
Xiaojing Liu, Feiyu Gao, Qiong Zhang, and Huasha Zhao · 2005
Earlier work this paper cites.
Machine learning for digital document processing: From layout analysis to metadata extraction
Floriana Esposito, Stefano Ferilli, Teresa MA Basile, and Nicola Di Mauro · 2008
Earlier work this paper cites.
A machine-learning approach for analyzing document layout structures with two reading orders
Chung-Chih Wu, Chien-Hsing Chou, and Fu Chang · 2008
Earlier work this paper cites.
Line segmentation for degraded handwritten historical documents
Itay Bar-Yosef, Nate Hagbi, Klara Kedem, and Itshak Dinstein · 2009
Earlier work this paper cites.
Script-independent handwritten textlines segmentation using active contours
Syed Saqib Bukhari, Faisal Shafait, and Thomas M Breuel · 2009
Earlier work this paper cites.
Learning rich hidden markov models in document analysis: Table location
Ana Costa e Silva · 2009
Earlier work this paper cites.
Hybrid page layout analysis via tab-stop detection
Raymond W Smith · 2009
Earlier work this paper cites.
Document image segmentation using discriminative learning over connected components
Syed Saqib Bukhari, Mayce Ibrahim Ali Al Azawi, Faisal Shafait, and Thomas M Breuel · 2010
Earlier work this paper cites.
An open approach towards the benchmarking of table structure recognition systems
Asif Shahab, Faisal Shafait, Thomas Kieninger, and Andreas Dengel · 2010
Earlier work this paper cites.
Structure-aware pre-training for table understanding with tree-based transformers
Zhiruo Wang, Haoyu Dong, Ran Jia, Jia Li, Zhiyi Fu, Shi Han, and Dongmei Zhang · 2010
Earlier work this paper cites.
Zilong Wang, Mingjie Zhan, Xuebo Liu, and Ding Liang · 2010
Earlier work this paper cites.
Multi resolution layout analysis of medieval manuscripts using dynamic mlp
Micheal Baechler and Rolf Ingold · 2011
Earlier work this paper cites.
Table detection in noisy off-line handwritten documents
Jin Chen and Daniel Lopresti · 2011
Earlier work this paper cites.
From one tree to a forest: A unified solution for structured web data extraction
Qiang Hao, Rui Cai, Yanwei Pang, and Lei Zhang · 2011
Earlier work this paper cites.
Layout analysis for arabic historical document images using machine learning
Syed Saqib Bukhari, Thomas M Breuel, Abedelkadir Asi, and Jihad El-Sana · 2012
Earlier work this paper cites.
Dataset, ground-truth and performance metrics for table detection evaluation
Jing Fang, Xin Tao, Zhi Tang, Ruiheng Qiu, and Ying Liu · 2012
Earlier work this paper cites.
Icdar 2013 table competition
Max C. Göbel, Tamir Hassan, Ermelinda Oro, and G. Orsi · 2013
Earlier work this paper cites.
Learning to detect tables in scanned document images using line information
Thotreingam Kasar, Philippine Barlas, Sebastien Adam, Clément Chatelain, and Thierry Paquet · 2013
Cited alongside, same era.
Fixed-point model for structured labeling
Quannan Li, Jingdong Wang, David Wipf, and Zhuowen Tu · 2013
Cited alongside, same era.
Evaluation of svm, mlp and gmm classifiers for layout analysis of historical documents
Hao Wei, Micheal Baechler, Fouad Slimane, and Rolf Ingold · 2013
Cited alongside, same era.
Table extraction from document images using fixed point model
Anukriti Bansal, Gaurav Harit, and Sumantra Dutta Roy · 2014
Cited alongside, same era.
A typed and handwritten text block segmentation system for heterogeneous and complex documents
Philippine Barlas, Sébastien Adam, Clément Chatelain, and Thierry Paquet · 2014
Cited alongside, same era.
Structural similarity for document image classification and retrieval
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych · 2019
Later among the works it cites.
Table detection in invoice documents by graph neural networks
Pau Riba, Anjan Dutta, Lutz Goldmann, Alicia Fornés, Oriol Ramos, and Josep Lladós · 2019
Later among the works it cites.
Deterministic routing between layout abstractions for multi-scale classification of visually rich documents
Ritesh Sarkhel and Arnab Nandi · 2019
Later among the works it cites.
Visual detection with context for document layout analysis
Carlos Soto and Shinjae Yoo · 2019
Later among the works it cites.
Publaynet: largest dataset ever for document layout analysis
Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes · 2019
Later among the works it cites.
One-shot text field labeling using attention and belief propagation for structure information extraction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Kumar, Peng Ye, and D. Doermann · 2014
Cited alongside, same era.
Deepdocclassifier: Document classification with deep convolutional neural network
Muhammad Zeshan Afzal, Samuele Capobianco, Muhammad Imran Malik, Simone Marinai, Thomas M Breuel, Andreas Dengel, and Marcus Liwicki · 2015
Cited alongside, same era.
Evaluation of deep convolutional nets for document image classification and retrieval
Adam W Harley, Alex Ufkes, and Konstantinos G Derpanis · 2015
Cited alongside, same era.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Ssd: Single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg · 2016
Cited alongside, same era.
Faster r-cnn: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2016
Cited alongside, same era.
Cutting the error by half: Investigation of very deep cnn and advanced training strategies for document image classification
Muhammad Zeshan Afzal, Andreas Kölsch, Sheraz Ahmed, and Marcus Liwicki · 2017
Cited alongside, same era.
Mengli Cheng, Minghui Qiu, Xing Shi, Jun Huang, and Wei Lin · 2020
Later among the works it cites.
Lambert: Layout-aware (language) modeling for information extraction
Łukasz Garncarek, Rafał Powalski, Tomasz Stanisławek, Bartosz Topolski, Piotr Halama, Michał Turski, and Filip Graliński · 2020
Later among the works it cites.
Tapas: Weakly supervised table parsing via pre-training
Jonathan Herzig, Paweł Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Martin Eisenschlos · 2020
Later among the works it cites.
Bros: A pre-trained language model for understanding texts in document
Teakgyu Hong, DongHyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park · 2020
Later among the works it cites.
Spatial dependency parsing for semi-structured document information extraction
Wonseok Hwang, Jinyeong Yim, Seunghyun Park, Sohee Yang, and Minjoon Seo · 2020
Later among the works it cites.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy · 2020
Later among the works it cites.
Visualwordgrid: Information extraction from scanned documents using a multimodal approach
Mohamed Kerroumi, Othmane Sayem, and Aymen Shabou · 2020
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Later among the works it cites.
TableBank: Table benchmark for image-based table detection and recognition
Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, Ming Zhou, and Zhoujun Li · 2020
Later among the works it cites.
DocBank: A benchmark dataset for document layout analysis
Minghao Li, Yiheng Xu, Lei Cui, Shaohan Huang, Furu Wei, Zhoujun Li, and Ming Zhou · 2020
Later among the works it cites.
Representation learning for information extraction from form-like documents
Bodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt, Qi Zhao, and Marc Najork · 2020
Later among the works it cites.
Iiit-ar-13k: a new dataset for graphical object detection in documents
Ajoy Mondal, Peter Lipps, and CV Jawahar · 2020
Later among the works it cites.
Cascadetabnet: An approach for end to end table detection and structure recognition from image-based documents
Devashish Prasad, Ayan Gadpal, Kshitij Kapadni, Manish Visave, and Kavita Sultanpure · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Later among the works it cites.
Project deepform: Extract information from documents, 2020
Jonathan Stray and Stacey Svetlichnaya · 2020
Later among the works it cites.
Robust layout-aware ie for visually rich documents with pre-trained language models
Mengxi Wei, Yifan He, and Qiong Zhang · 2020
Later among the works it cites.
LayoutLM: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou · 2020
Later among the works it cites.
Trie: End-to-end text reading and information extraction for document understanding
Peng Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu, Liang Qiao, Yi Niu, and Fei Wu · 2020
Later among the works it cites.
Tncr: Table net detection and classification dataset
Abdelrahman Abdallah, Alexander Berendeyev, Islam Nuradin, and Daniyar Nurseitov · 2021
Closest in time.
Docformer: End-to-end transformer for document understanding
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R Manmatha · 2021
Closest in time.
Websrc: A dataset for web-based structural reading comprehension, 2021
Lu Chen, Xingyu Chen, Zihan Zhao, Danyang Zhang, Jiabao Ji, Ao Luo, Yuxuan Xiong, and Kai Yu · 2021
Closest in time.
Tablex: A benchmark dataset for structure and content information extraction from scientific tables, 2021
Harsh Desai, Pratik Kayal, and Mayank Singh · 2021
Closest in time.
Weihong Lin, Qifang Gao, Lei Sun, Zhuoyao Zhong, Kai Hu, Qin Ren, and Qiang Huo · 2021
Closest in time.
Going full-tilt boogie on document understanding with text-image-layout transformer
Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz, Tomasz Dwojak, Michał Pietruszka, and Gabriela Pałka · 2021
Closest in time.
Pubtables-1m: Towards a universal dataset and metrics for training and evaluating table extraction models, 2021
Brandon Smock, Rohith Pesala, and Robin Abraham · 2021
Closest in time.
Kleister: Key information extraction datasets involving long documents with complex layouts, 2021
Tomasz Stanisławek, Filip Graliński, Anna Wróblewska, Dawid Lipiński, Agnieszka Kaliska, Paulina Rosalska, Bartosz Topolski, and Przemysław Biecek · 2021
Closest in time.
Visualmrc: Machine reading comprehension on document images
Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida · 2021
Closest in time.
Lampret: Layout-aware multimodal pretraining for document understanding
Te-Lin Wu, Cheng Li, Mingyang Zhang, Tao Chen, Spurthi Amba Hombaiah, and Michael Bendersky · 2021
Closest in time.
LayoutLMv2: Multi-modal pre-training for visually-rich document understanding
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou · 2021
Closest in time.
Icdar 2021 competition on scientific literature parsing, 2021
Antonio Jimeno Yepes, Xu Zhong, and Douglas Burdick · 2021
Closest in time.
Pick: Processing key information extraction from documents using improved graph learning-convolutional networks
Wenwen Yu, Ning Lu, Xianbiao Qi, Ping Gong, and Rong Xiao · 2021
Closest in time.