Fetching the paper…
Reading the bibliography…
Recent advances in multimodal pre-trained models have significantly improved information extraction from visually-rich documents (VrDs), in which named entity recognition (NER) is treated as a sequence-labeling task of predicting the BIO entity tags for tokens, following the typical setting of NLP.
Multi-sample dropout for accelerated training and better generalization
Hiroshi Inoue. 2019 · 1905
Earlier work this paper cites.
Structext: Structured text understanding with multi-modal transformers
Yulin Li, Yuxi Qian, Yuechen Yu, Xiameng Qin, Chengquan Zhang, Yan Liu, Kun Yao, Junyu Han, Jingtuo Liu, and Errui Ding. 2021c · 1920
Earlier work this paper cites.
Text chunking using transformation-based learning
Lance A Ramshaw and Mitchell P Marcus. 1999 · 1999
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020 · 2010
Earlier work this paper cites.
Recognizing disjoint clinical concepts in clinical text using machine learning-based methods
Buzhou Tang, Qingcai Chen, Xiaolong Wang, Yonghui Wu, Yaoyun Zhang, Min Jiang, Jingqi Wang, and Hua Xu. 2015 · 2015
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Icdar2019 competition on scanned receipt ocr and information extraction
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and CV Jawahar. 2019 · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019 · 2019
Earlier work this paper cites.
Pifpaf: Composite fields for human pose estimation
Sven Kreiss, Lorenzo Bertoni, and Alexandre Alahi. 2019 · 2019
Earlier work this paper cites.
Cord: a consolidated receipt dataset for post-ocr parsing
Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee. 2019 · 2019
Earlier work this paper cites.
Named entity recognition and relation extraction with graph neural networks in semi structured documents
Manuel Carbonell, Pau Riba, Mauricio Villegas, Alicia Fornés, and Josep Lladós. 2021 · 2020
Earlier work this paper cites.
An effective transition-based model for discontinuous ner
Xiang Dai, Sarvnaz Karimi, Ben Hachey, and Cecile Paris. 2020 · 2020
Earlier work this paper cites.
End-to-end hierarchical relation extraction for generic form understanding
Tuan Anh Nguyen Dang, Duc Thanh Hoang, Quang Bach Tran, Chih-Wei Pan, and Thanh Dat Nguyen. 2021 · 2020
Cited alongside, same era.
Lambert: Layout-aware language modeling for information extraction
Lukasz Garncarek, Rafal Powalski, Tomasz Stanislawek, Bartosz Topolski, Piotr Halama, Michał Turski, and Filip Grali’nski. 2020 · 2020
Cited alongside, same era.
Project deepform: Extract information from documents
Jonathan Stray and Stacey Svetlichnaya. 2020 · 2020
Cited alongside, same era.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
End-to-end information extraction by character-level embedding and multi-stage attentional u-net
Tuan-Anh Nguyen Dang and Dat-Thanh Nguyen. 2021 · 2021
Cited alongside, same era.
Setgner: General named entity recognition as entity set generation
Yuxin He and Buzhou Tang. 2022 · 2022
Later among the works it cites.
Bros: A pre-trained language model focusing on text and layout for better key information extraction from documents
Teakgyu Hong, Donghyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park. 2022 · 2022
Later among the works it cites.
Layoutlmv3: Pre-training for document AI with unified text and image masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. 2022 · 2022
Later among the works it cites.
Dit: Self-supervised pre-training for document image transformer
Junlong Li, Yiheng Xu, Tengchao Lv, Lei Cui, Cha Zhang, and Furu Wei. 2022c · 2022
Later among the works it cites.
Toe: A grid-tagging discontinuous ner model enhanced by embedding tag/word relations and more fine-grained tags
Jiang Liu, Donghong Ji, Jingye Li, Dongdong Xie, Chong Teng, Liang Zhao, and Fei Li. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fuzzybio: a proposal for fuzzy representation of discontinuous entities
AR Dirkson, Suzan Verberne, Wessel Kraaij, E Holderness, A Jimeno Yepes, A Lavelli, AL Minard, J Pustejovsky, and F Rinaldi. 2021 · 2021
Cited alongside, same era.
Spatial dependency parsing for semi-structured document information extraction
Wonseok Hwang, Jinyeong Yim, Seunghyun Park, Sohee Yang, and Minjoon Seo. 2021 · 2021
Cited alongside, same era.
Kleister: key information extraction datasets involving long documents with complex layouts
Tomasz Stanisławek, Filip Graliński, Anna Wróblewska, Dawid Lipiński, Agnieszka Kaliska, Paulina Rosalska, Bartosz Topolski, and Przemysław Biecek. 2021 · 2021
Cited alongside, same era.
Layoutreader: Pre-training of text and layout for reading order detection
Zilong Wang, Yiheng Xu, Lei Cui, Jingbo Shang, and Furu Wei. 2021b · 2021
Cited alongside, same era.
Entity relation extraction as dependency parsing in visually rich documents
Yue Zhang, Zhang Bo, Rui Wang, Junjie Cao, Chen Li, and Zuyi Bao. 2021 · 2021
Cited alongside, same era.
Doc2graph: a task agnostic document understanding framework based on graph neural networks
Andrea Gemelli, Sanket Biswas, Enrico Civitelli, Josep Lladós, and Simone Marinai. 2022 · 2022
Cited alongside, same era.
Xylayoutlm: Towards layout-aware multimodal networks for visually-rich document understanding
Zhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan, Weiqiang Wang, Ming Gu, and Liqing Zhang. 2022 · 2022
Cited alongside, same era.
Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Teng Hu, Weichong Yin, Yongfeng Chen, Yin Zhang, et al. 2022 · 2022
Later among the works it cites.
Global pointer: Novel efficient span-based approach for named entity recognition
Jianlin Su, Ahmed Murtadha, Shengfeng Pan, Jing Hou, Jun Sun, Wanwei Huang, Bo Wen, and Yunfeng Liu. 2022 · 2022
Later among the works it cites.
Lilt: A simple yet effective language-independent layout transformer for structured document understanding
Jiapeng Wang, Lianwen Jin, and Kai Ding. 2022 · 2022
Later among the works it cites.
Layoutmask: Enhance text-layout interaction in multi-modal pre-training for document understanding
Yi Tu, Ya Guo, Huan Chen, and Jinyang Tang. 2023 · 2023
Closest in time.
Icdar 2023 competition on structured text extraction from visually-rich document images
Wenwen Yu, Chengquan Zhang, Haoyu Cao, Wei Hua, Bohan Li, Huang Chen, Mingyu Liu, Mingrui Chen, Jianfeng Kuang, Mengjun Cheng, et al. 2023 · 2023
Closest in time.
Ernie: Enhanced language representation with informative entities
Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019 · 2023
Closest in time.