Fetching the paper…
Reading the bibliography…
We propose a novel end-to-end document understanding model called SeRum (SElective Region Understanding Model) for extracting meaningful information from document images, including document analysis, retrieval, and office automation.
Simple fast algorithms for the editing distance between trees and related problems
Kaizhong Zhang and Dennis E. Shasha · 1989
Earlier work this paper cites.
Recursive x-y cut using bounding boxes of connected components
Jaekyu Ha, R.M. Haralick, and I.T. Phillips · 1995
Earlier work this paper cites.
A fast algorithm for bottom-up document layout analysis
Anikó Simon, Jean-Christophe Pret, and A. Peter Johnson · 1997
Earlier work this paper cites.
Artificial neural networks for document analysis and recognition
Simone Marinai, Marco Gori, and Giovanni Soda · 2005
Earlier work this paper cites.
Learning nongenerative grammatical models for document analysis
M. Shilman, P. Liang, and P. Viola · 2005
Earlier work this paper cites.
Building a test collection for complex document information processing
David D. Lewis, Gady Agam, Shlomo Argamon, Ophir Frieder, David A. Grossman, and Jefferson Heard · 2006
Earlier work this paper cites.
Synthetic data and artificial neural networks for natural scene text recognition
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Synthetic data for text localisation in natural images
Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Chargrid: Towards understanding 2d documents
Anoop R. Katti, Christian Reisswig, Cordula Guder, Sebastian Brarda, Steffen Bickel, Johannes Höhne, and Jean Baptiste Faddoul · 2018
Earlier work this paper cites.
Bertgrid: Contextualized embedding for 2d document representation and understanding
Timo I. Denk and Christian Reisswig · 2019
Earlier work this paper cites.
Textdragon: An end-to-end framework for arbitrary shaped text spotting
Wei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang, and Cheng-Lin Liu · 2019
Earlier work this paper cites.
EATEN: entity-aware attention for single shot visual text extraction
He Guo, Xiameng Qin, Jiaming Liu, Junyu Han, Jingtuo Liu, and Errui Ding · 2019
Earlier work this paper cites.
Post-ocr parsing: building simple and robust parser via bio tagging
Wonseok Hwang, Seonghyeon Kim, Minjoon Seo, Jinyeong Yim, Seunghyun Park, Sungrae Park, Junyeop Lee, Bado Lee, and Hwalsuk Lee · 2019
Earlier work this paper cites.
Graph convolution for multimodal information extraction from visually rich documents
Xiaojing Liu, Feiyu Gao, Qiong Zhang, and Huasha Zhao · 2019
Earlier work this paper cites.
Openceres: When open information extraction meets the semi-structured web
Colin Lockard, Prashant Shiralkar, and Xin Luna Dong · 2019
Earlier work this paper cites.
Cord: A consolidated receipt dataset for post-ocr parsing
Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee · 2019
Cited alongside, same era.
Graphie: A graph-based framework for information extraction
Yujie Qian, Enrico Santus, Zhijing Jin, Jiang Guo, and Regina Barzilay · 2019
Cited alongside, same era.
A large chinese text dataset in the wild
Tai-Ling Yuan, Zhe Zhu, Kun Xu, Cheng-Jun Li, Tai-Jiang Mu, and Shi-Min Hu · 2019
Cited alongside, same era.
Robotic process automation: An overview and comparison to other technology in industry 4.0
Bernhard Axmann and Harmoko Harmoko · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Towards augmenting lexical resources for slang and african american english
Document AI: benchmarks, models and applications
Lei Cui, Yiheng Xu, Tengchao Lv, and Furu Wei · 2021
Later among the works it cites.
ICDAR2019 competition on scanned receipt OCR and information extraction
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and C. V. Jawahar · 2021
Later among the works it cites.
Spatial dependency parsing for semi-structured document information extraction
Wonseok Hwang, Jinyeong Yim, Seunghyun Park, Sohee Yang, and Minjoon Seo · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Docvqa: A dataset for VQA on document images
Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alyssa Hwang, William R. Frey, and Kathleen R. McKeown · 2020
Cited alongside, same era.
BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Cited alongside, same era.
Abcnet: Real-time scene text spotting with adaptive bezier-curve network
Yuliang Liu, Hao Chen, Chunhua Shen, Tong He, Lianwen Jin, and Liangwei Wang · 2020
Cited alongside, same era.
Zeroshotceres: Zero-shot relation extraction from semi-structured webpages
Colin Lockard, Prashant Shiralkar, Xin Luna Dong, and Hannaneh Hajishirzi · 2020
Cited alongside, same era.
Ad-hoc document retrieval using weak-supervision with BERT and GPT2
Yosi Mass and Haggai Roitman · 2020
Cited alongside, same era.
End-to-end extraction of structured information from business documents with pointer-generator networks
Clément Sage, Alex Aussem, Véronique Eglin, Haytham Elghazel, and Jérémy Espinas · 2020
Cited alongside, same era.
Assessing the impact of OCR quality on downstream NLP tasks
Daniel van Strien, Kaspar Beelen, Mariona Coll Ardanuy, Kasra Hosseini, Barbara McGillivray, and Giovanni Colavizza · 2020
Cited alongside, same era.
Going full-tilt boogie on document understanding with text-image-layout transformer
Rafal Powalski, Lukasz Borchmann, Dawid Jurkiewicz, Tomasz Dwojak, Michal Pietruszka, and Gabriela Palka · 2021
Later among the works it cites.
Layoutxlm: Multimodal pre-training for multilingual visually-rich document understanding
Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florêncio, Cha Zhang, and Furu Wei · 2021
Later among the works it cites.
Layoutlmv2: Multi-modal pre-training for visually-rich document understanding
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou · 2021
Later among the works it cites.
Query-driven generative network for document information extraction in the wild
Haoyu Cao, Xin Li, Jiefeng Ma, Deqiang Jiang, Antai Guo, Yiqing Hu, Hao Liu, Yinsong Liu, and Bo Ren · 2022
Later among the works it cites.
GMN: generative multi-modal network for practical document information extraction
Haoyu Cao, Jiefeng Ma, Antai Guo, Yiqing Hu, Hao Liu, Deqiang Jiang, Yinsong Liu, and Bo Ren · 2022
Later among the works it cites.
End-to-end document recognition and understanding with dessurt
Brian L. Davis, Bryan S. Morse, Brian L. Price, Chris Tensmeyer, Curtis Wigington, and Vlad I. Morariu · 2022
Later among the works it cites.
BROS: A pre-trained language model focusing on text and layout for better key information extraction from documents
Teakgyu Hong, Donghyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park · 2022
Later among the works it cites.
Layoutlmv3: Pre-training for document AI with unified text and image masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei · 2022
Later among the works it cites.
Ocr-free document understanding transformer
Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park · 2022
Later among the works it cites.
Relational representation learning in visually-rich documents
Xin Li, Yan Zheng, Yiqing Hu, Haoyu Cao, Yunfei Wu, Deqiang Jiang, Yinsong Liu, and Bo Ren · 2022
Later among the works it cites.
Unifying vision, text, and layout for universal document processing
Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Cha Zhang, and Mohit Bansal · 2022
Later among the works it cites.
Spts v2: single-point scene text spotting
Yuliang Liu, Jiaxin Zhang, Dezhi Peng, Mingxin Huang, Xinyu Wang, Jingqun Tang, Can Huang, Dahua Lin, Chunhua Shen, Xiang Bai, et al · 2023
Closest in time.