Fetching the paper…
Reading the bibliography…
Structured text understanding on Visually Rich Documents (VRDs) is a crucial part of Document Intelligence.
smartFIX: A Requirements-Driven System for Document Analysis and Understanding. In DAS . Springer, 433–444
Andreas Dengel and Bertin Klein. 2002 · 2002
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database. In CVPR . IEEE, 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li. 2009 · 2009
Earlier work this paper cites.
Subhojeet Pramanik, Shashank Mujumdar, and Hima Patel. 2020 · 2009
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality. In NIPS . 3111–3119
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Accurate Scene Text Recognition Based on Recurrent Neural Network. In ACCV . Springer, 35–48
Bolan Su and Shijian Lu. 2014 · 2014
Earlier work this paper cites.
Evaluation of deep convolutional nets for document image classification and retrieval. In ICDAR . IEEE, 991–995
Adam W. Harley, Alex Ufkes, and Konstantinos G. Derpanis. 2015 · 2015
Earlier work this paper cites.
Neural Architectures for Named Entity Recognition. In ACL . ACL, 260–270
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016 · 2016
Earlier work this paper cites.
End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF. In ACL . ACL, 1064–1074
Xuezhe Ma and Eduard Hovy. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
Mask R-CNN. In ICCV . 2961–2969
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross B. Girshick. 2017 · 2017
Earlier work this paper cites.
Feature Pyramid Networks for Object Detection. In CVPR . IEEE, 936–944
Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie. 2017 · 2017
Earlier work this paper cites.
CloudScan - A Configuration-Free Invoice Analysis System Using Recurrent Neural Networks. In ICDAR . IEEE, 406–413
Rasmus Berg Palm, Ole Winther, and Florian Laws. 2017 · 2017
Earlier work this paper cites.
An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition
Baoguang Shi, Xiang Bai, and Cong Yao. 2017 · 2017
Earlier work this paper cites.
Aggregated Residual Transformations for Deep Neural Networks. In CVPR . IEEE, 5987–5995
Saining Xie, Ross B. Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. 2017 · 2017
Earlier work this paper cites.
EAST: An Efficient and Accurate Scene Text Detector. In CVPR . IEEE, 2642–2651
Xinyu Zhou, Cong Yao, He Wen, Yuzhi Wang, Shuchang Zhou, Weiran He, and Jiajun Liang. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Chargrid: Towards Understanding 2D Documents. In EMNLP . ACL, 4459–4469
Anoop R. Katti, Christian Reisswig, Cordula Guder, Sebastian Brarda, Steffen Bickel, Johannes Höhne, and Jean Baptiste Faddoul. 2018 · 2018
Cited alongside, same era.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context. In ACL . ACL, 2978–2988
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Bertgrid: Contextualized embedding for 2d document representation and understanding
Timo I Denk and Christian Reisswig. 2019 · 2019
Cited alongside, same era.
EATEN: Entity-Aware Attention for Single Shot Visual Text Extraction. In ICDAR . IEEE, 254–259
He Guo, Xiameng Qin, Jiaming Liu, Junyu Han, Jingtuo Liu, and Errui Ding. 2019 · 2019
Cited alongside, same era.
Bag of Tricks for Image Classification with Convolutional Neural Networks. In CVPR . IEEE, 558–567
Tong He, Zhi Zhang, Hang Zhang, Zhongyue Zhang, Junyuan Xie, and Mu Li. 2019 · 2019
Cutie: Learning to understand documents with convolutional universal text information extractor
Xiaohui Zhao, Endi Niu, Zhuo Wu, and Xiaoguang Wang. 2019 · 2019
Later among the works it cites.
Named Entity Recognition and Relation Extraction with Graph Neural Networks in Semi Structured Documents. In ICPR . IEEE, 9622–9627
Manuel Carbonell, Pau Riba, Mauricio Villegas, Alicia Fornés, and Josep Lladós. 2020 · 2020
Later among the works it cites.
One-shot Text Field labeling using Attention and Belief Propagation for Structure Information Extraction. In ACM Multimedia . ACM, 340–348
Mengli Cheng, Minghui Qiu, Xing Shi, Jun Huang, and Wei Lin. 2020 · 2020
Later among the works it cites.
DocBank: A Benchmark Dataset for Document Layout Analysis. In COLING . ICCL, 949–960
Minghao Li, Yiheng Xu, Lei Cui, Shaohan Huang, Furu Wei, Zhoujun Li, and Ming Zhou. 2020 · 2020
Later among the works it cites.
Mask textspotter v3: Segmentation proposal network for robust scene text spotting. In ECCV . Springer, 706–722
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Icdar2019 competition on scanned receipt ocr and information extraction. In ICDAR . IEEE, 1516–1520
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and CV Jawahar. 2019 · 2019
Cited alongside, same era.
Post-OCR parsing: building simple and robust parser via BIO tagging. In NeurIPS Workshop
Wonseok Hwang, Seonghyeon Kim, Minjoon Seo, Jinyeong Yim, Seunghyun Park, Sungrae Park, Junyeop Lee, Bado Lee, and Hwalsuk Lee. 2019 · 2019
Cited alongside, same era.
FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents. In ICDAR Workshop . IEEE, 1–6
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019 · 2019
Cited alongside, same era.
Graph Convolution for Multimodal Information Extraction from Visually Rich Documents. In NAACL-HLT . ACL, 32–39
Xiaojing Liu, Feiyu Gao, Qiong Zhang, and Huasha Zhao. 2019 · 2019
Cited alongside, same era.
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. In NeurIPS . 13–23
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
Attend, Copy, Parse End-to-end Information Extraction from Documents. In ICDAR . IEEE, 329–336
Rasmus Berg Palm, Florian Laws, and Ole Winther. 2019 · 2019
Cited alongside, same era.
GraphIE: A Graph-Based Framework for Information Extraction. In ACL . ACL, 751–761
Yujie Qian, Enrico Santus, Zhijing Jin, Jiang Guo, and Regina Barzilay. 2019 · 2019
Cited alongside, same era.
Minghui Liao, Guan Pang, Jing Huang, Tal Hassner, and Xiang Bai. 2020 · 2020
Later among the works it cites.
Representation Learning for Information Extraction from Form-like Documents. In ACL . ACL, 6495–6504
Bodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt, Qi Zhao, and Marc Najork. 2020 · 2020
Later among the works it cites.
End-to-End Extraction of Structured Information from Business Documents with Pointer-Generator Networks. In SPNLP . ACL, 43–52
Clément Sage, Alex Aussem, Véronique Eglin, Haytham Elghazel, and Jérémy Espinas. 2020 · 2020
Later among the works it cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations. In ICLR
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
ERNIE 2.0: A Continual Pre-Training Framework for Language Understanding. In AAAI . AAAI, 8968–8975
Yu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang. 2020 · 2020
Later among the works it cites.
DocStruct: A Multimodal Method to Extract Hierarchy Structure in Document for General Form Understanding. In EMNLP . ACL, 898–908
Zilong Wang, Mingjie Zhan, Xuebo Liu, and Ding Liang. 2020 · 2020
Later among the works it cites.
Robust Layout-aware IE for Visually Rich Documents with Pre-trained Language Models. In SIGIR . ACM, 2367–2376
Mengxi Wei, Yifan He, and Qiong Zhang. 2020 · 2020
Later among the works it cites.
LayoutLMv2: Multi-modal pre-training for visually-rich document understanding
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, et al · 2020
Later among the works it cites.
TRIE: End-to-End Text Reading and Information Extraction for Document Understanding. In ACM Multimedia . ACM, 1413–1422
Peng Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu, Liang Qiao, Yi Niu, and Fei Wu. 2020 · 2020
Later among the works it cites.
Spatial Dependency Parsing for Semi-Structured Document Information Extraction. In ACL-IJCNLP
Wonseok Hwang, Jinyeong Yim, Seunghyun Park, Sohee Yang, and Minjoon Seo. 2021 · 2021
Closest in time.
MatchVIE: Exploiting Match Relevancy between Entities for Visual Information Extraction. In IJCAI . ijcai.org
Guozhi Tang, Lele Xie, Lianwen Jin, Jiapeng Wang, Jingdong Chen, Zhen Xu, Qianying Wang, Yaqiang Wu, and Hui Li. 2021 · 2021
Closest in time.