Fetching the paper…
Reading the bibliography…
Visual information extraction (VIE) plays an important role in Document Intelligence.
Building a test collection for complex document information processing
David Lewis, Gady Agam, Shlomo Argamon, Ophir Frieder, David Grossman, and Jefferson Heard · 2006
Earlier work this paper cites.
Network in network
Min Lin, Qiang Chen, and Shuicheng Yan · 2014
Earlier work this paper cites.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Baoguang Shi, Xiang Bai, and Cong Yao · 2015
Earlier work this paper cites.
Scene text detection and recognition: recent advances and future trends
Yingying Zhu, Cong Yao, and Xiang Bai · 2016
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
East: An efficient and accurate scene text detector
Xinyu Zhou, Cong Yao, He Wen, Yuzhi Wang, Shuchang Zhou, Weiran He, and Jiajun Liang · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
Scene text detection and recognition: The deep learning era
Shangbang Long, Xin He, and Cong Yao · 2018
Earlier work this paper cites.
An embarrassingly simple approach for transfer learning from pretrained language models
Alexandra Chronopoulou, Christos Baziotis, and Alexandros Potamianos · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran · 2019
Earlier work this paper cites.
Graph convolution for multimodal information extraction from visually rich documents
Xiaojing Liu, Feiyu Gao, and Qiong Zhang · 2019
Earlier work this paper cites.
CORD: A consolidated receipt dataset for post-ocr parsing
Seunghyun Park, Seung Shin, Bado Lee, and Junyeop Lee · 2019
Earlier work this paper cites.
GraphIE: A graph-based framework for information extraction
Yujie Qian, Enrico Santus, and Zhijing Jin · 2019
Earlier work this paper cites.
Graphie: A graph-based framework for information extraction
Yujie Qian, Enrico Santus, Zhijing Jin, Jiang Guo, and Regina Barzilay · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Don’t stop pretraining: adapt language models to domains and tasks
Suchin Gururangan, Ana Marasovi’c, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith · 2020
Cited alongside, same era.
Real-time scene text detection with differentiable binarization
Minghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen, and Xiang Bai · 2020
Cited alongside, same era.
Merge and recognize: a geometry and 2d context aware graph model for named entity recognition from visual documents
Chuwei Luo, Yongpan Wang, Qi Zheng, Liangchen Li, Feiyu Gao, and Shiyu Zhang · 2020
Vibertgrid: a jointly trained multi-modal 2d document representation for key information extraction from documents
Weihong Lin, Qifang Gao, Lei Sun, Zhuoyao Zhong, Kai Hu, Qin Ren, and Qiang Huo · 2021
Later among the works it cites.
MatchVIE: Exploiting match relevancy between entities for visual information extraction
Guozhi Tang, Lele Xie, Lianwen Jin, and Wang · 2021
Later among the works it cites.
Layoutxlm: Multimodal pre-training for multilingual visually-rich document understanding
Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, and Furu Wei · 2021
Later among the works it cites.
Layoutlmv2: Multi-modal pre-training for visually-rich document understanding
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou · 2021
Later among the works it cites.
Entity relation extraction as dependency parsing in visually rich documents
Yue Zhang, Bo Zhang, Rui Wang, Junjie Cao, Chen Li, and Zuyi Bao · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
LayoutLM: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, and Shaohan Huang · 2020
Cited alongside, same era.
Pick: Processing key information extraction from documents using improved graph learning-convolutional networks
Wenwen Yu, Ning Lu, Xianbiao Qi, Ping Gong, and Rong Xiao · 2020
Cited alongside, same era.
DocFormer: End-to-end transformer for document understanding
Srikar Appalaraju, Bhavan Jasani, and Bhargava Urala Kota · 2021
Cited alongside, same era.
Document ai: Benchmarks, models and applications
Lei Cui, Yiheng Xu, Tengchao Lv, and Furu Wei · 2021
Cited alongside, same era.
Unified pretraining framework for document understanding
Jiuxiang Gu, Jason Kuen, Vlad I Morariu, Handong Zhao, Nikolaos Barmpalios, Rajiv Jain, Ani Nenkova, and Tong Sun · 2021
Cited alongside, same era.
Adaptive transfer learning on graph neural networks
Xueting Han, Zhenhuan Huang, Bang An, and Jing Bai · 2021
Cited alongside, same era.
Spatial dependency parsing for semi-structured document information extraction
Wonseok Hwang, Jinyeong Yim, Seunghyun Park, Sohee Yang, and Minjoon Seo · 2021
Cited alongside, same era.
Later among the works it cites.
An empirical study of training end-to-end vision-and-language transformers
Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Pengchuan Zhang, Lu Yuan, Nanyun Peng, et al · 2022
Later among the works it cites.
Xylayoutlm: Towards layout-aware multimodal networks for visually-rich document understanding
Zhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan, Weiqiang Wang, Ming Gu, and Liqing Zhang · 2022
Later among the works it cites.
Bros: A pre-trained language model focusing on text and layout for better key information extraction from documents
Teakgyu Hong, Donghyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park · 2022
Later among the works it cites.
Layoutlmv3: Pre-training for document ai with unified text and image masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei · 2022
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2022
Later among the works it cites.
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Later among the works it cites.
Bi-vldoc: Bidirectional vision-language modeling for visually-rich document understanding
Chuwei Luo, Guozhi Tang, Qi Zheng, Cong Yao, Lianwen Jin, Chenliang Li, Yang Xue, and Luo Si · 2022
Later among the works it cites.
Lilt: A simple yet effective language-independent layout transformer for structured document understanding
Jiapeng Wang, Lianwen Jin, and Kai Ding · 2022
Later among the works it cites.
Multi-granularity prediction for scene text recognition
Peng Wang, Cheng Da, and Cong Yao · 2022
Later among the works it cites.
Xfund: A benchmark dataset for multilingual visually rich form understanding
Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, and Furu Wei · 2022
Later among the works it cites.