Fetching the paper…
Reading the bibliography…
In this paper, we present a new question-answering (QA) based key-value pair extraction approach, called KVPFormer, to robustly extracting key-value relationships between entities from form-like document images.
StrucTexT: Structured Text Understanding with Multi-Modal Transformers
Li, Y.; Qian, Y.; Yu, Y.; Qin, X.; Zhang, C.; Liu, Y.; Yao, K.; Han, J.; Liu, J.; and Ding, E. 2021 · 1920
Earlier work this paper cites.
Layout Recognition of Multi-Kinds of Table-Form Documents
Watanabe, T.; Luo, Q.; and Sugie, N. 1995 · 1995
Earlier work this paper cites.
Information Management System Using Structure Analysis of Paper/Electronic Documents and Its Application
Seki, M.; Fujio, M.; Nagasaki, T.; Shinjo, H.; and Marukawa, K. 2007 · 2007
Earlier work this paper cites.
Wang, Z.; Zhan, M.; Liu, X.; and Liang, D. 2020 · 2010
Earlier work this paper cites.
Development of Template-Free form Recognition System
Hirayama, J.; Shinjo, H.; Takahashi, T.; and Nagasaki, T. 2011 · 2011
Earlier work this paper cites.
Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval
Harley, A. W.; Ufkes, A.; and Derpanis, K. G. 2015 · 2015
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Attention is All You Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
End-to-End Information Extraction by Character-Level Embedding and Multi-Stage Attentional U-Net
Dang, T.-A. N.; and Nguyen, D.-T. 2019 · 2019
Earlier work this paper cites.
Deep Visual Template-Free form Parsing
Davis, B.; Morse, B.; Cohen, S.; Price, B.; and Tensmeyer, C. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction
Huang, Z.; Chen, K.; He, J.; Bai, X.; Karatzas, D.; Lu, S.; and Jawahar, C. 2019 · 2019
Cited alongside, same era.
FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents
Jaume, G.; Ekenel, H. K.; and Thiran, J.-P. 2019 · 2019
Cited alongside, same era.
End-to-End Object Detection with Transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Cited alongside, same era.
BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents
Hong, T.; Kim, D.; Ji, M.; Hwang, W.; Nam, D.; and Park, S. 2020 · 2020
Cited alongside, same era.
LayoutLM: Pre-training of Text and Layout for Document Image Understanding
Xu, Y.; Li, M.; Cui, L.; Huang, S.; Wei, F.; and Zhou, M. 2020 · 2020
Cited alongside, same era.
ViBERTgrid: A Jointly Trained Multi-Modal 2D Document Representation for Key Information Extraction from Documents
Lin, W.; Gao, Q.; Sun, L.; Zhong, Z.; Hu, K.; Ren, Q.; and Huo, Q. 2021 · 2021
Later among the works it cites.
Docvqa: A Dataset for VQA on Document Images
Mathew, M.; Karatzas, D.; and Jawahar, C. 2021 · 2021
Later among the works it cites.
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer
Powalski, R.; Borchmann, Ł.; Jurkiewicz, D.; Dwojak, T.; Pietruszka, M.; and Pałka, G. 2021 · 2021
Later among the works it cites.
MTL-FoUn: A Multi-Task Learning Approach to Form Understanding
Prabhu, N.; Jain, H.; and Tripathi, A. 2021 · 2021
Later among the works it cites.
Text Classification Models for Form Entity Linking
Villota, M.; Dom’inguez, C.; Heras, J.; Mata, E. J.; and Pascual, V. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carbonell, M.; Riba, P.; Villegas, M.; Fornés, A.; and Lladós, J. 2021 · 2021
Cited alongside, same era.
End-to-End Hierarchical Relation Extraction for Generic Form Understanding
Dang, T. A. N.; Hoang, D. T.; Tran, Q. B.; Pan, C.-W.; and Nguyen, T. D. 2021 · 2021
Cited alongside, same era.
Visual FUDGE: Form Understanding via Dynamic Graph Editing
Davis, B.; Morse, B.; Price, B.; Tensmeyer, C.; and Wiginton, C. 2021 · 2021
Cited alongside, same era.
Value Retrieval with Arbitrary Queries for Form-like Documents
Gao, M.; Xue, L.; Ramaiah, C.; Xing, C.; Xu, R.; and Xiong, C. 2021 · 2021
Cited alongside, same era.
LAMBERT: Layout-Aware Language Modeling for Information Extraction
Garncarek, Ł.; Powalski, R.; Stanisławek, T.; Topolski, B.; Halama, P.; Turski, M.; and Graliński, F. 2021 · 2021
Cited alongside, same era.
Spatial Dependency Parsing for Semi-Structured Document Information Extraction
Hwang, W.; Yim, J.; Park, S.; Yang, S.; and Seo, M. 2021 · 2021
Cited alongside, same era.
PICK: Processing Key Information Extraction from Documents Using Improved Graph Learning-Convolutional Networks
Yu, W.; Lu, N.; Qi, X.; Gong, P.; and Xiao, R. 2021 · 2021
Later among the works it cites.
Entity Relation Extraction as Dependency Parsing in Visually Rich Documents
Zhang, Y.; Bo, Z.; Wang, R.; Cao, J.; Li, C.; and Bao, Z. 2021 · 2021
Later among the works it cites.
XYLayoutLM: Towards Layout-Aware Multimodal Networks For Visually-Rich Document Understanding
Gu, Z.; Meng, C.; Wang, K.; Lan, J.; Wang, W.; Gu, M.; and Zhang, L. 2022 · 2022
Later among the works it cites.
LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding
Wang, J.; Jin, L.; and Ding, K. 2022 · 2022
Later among the works it cites.
XFUND: A Benchmark Dataset for Multilingual Visually Rich Form Understanding
Xu, Y.; Lv, T.; Cui, L.; Wang, G.; Lu, Y.; Florencio, D.; Zhang, C.; and Wei, F. 2022 · 2022
Later among the works it cites.