Fetching the paper…
Reading the bibliography…
This paper presents a comprehensive survey of research works on the topic of form understanding in the context of scanned documents.
A knowledge-based segmentation method for document understanding
Jun’ichi Higashino, Hiromichi Fujisawa, Yasuaki Nakano, and Masakazu Ejiri · 1986
Earlier work this paper cites.
Automatic Analysis and Understanding of Documents
Yuan Y. Tang, Chang D. Yan, M. Cheriet, and Ching Y. Suen · 1993
Earlier work this paper cites.
Analysis of form images
Dacheng Wang and Sargur N Srihari · 1994
Earlier work this paper cites.
Automatic document processing: A survey
Yuan Y. Tang, Seong-Whan Lee, and Ching Y. Suen · 1996
Earlier work this paper cites.
Building digital tobacco industry document libraries at the university of california, san francisco library/center for knowledge management
Heidi Schmidt, Karen Butter, and Cynthia Rider · 2002
Earlier work this paper cites.
Camera-based analysis of text and documents: a survey
Jian Liang, David Doermann, and Huiping Li · 2005
Earlier work this paper cites.
Building a test collection for complex document information processing
David D. Lewis, Gady Agam, Shlomo Engelson Argamon, Ophir Frieder, David A. Grossman, and Jefferson Heard · 2006
Earlier work this paper cites.
Cdip dataset
David D. Lewis, Gady Agam, Shlomo Engelson Argamon, Ophir Frieder, David A. Grossman, and Jefferson Heard · 2011
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation, 2016
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Earlier work this paper cites.
Wikireading: A novel large-scale language understanding task over wikipedia
Daniel Hewlett, Alexandre Lacoste, Llion Jones, Illia Polosukhin, Andrew Fandrianto, Jay Han, Matthew Kelcey, and David Berthelot · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks, 2017
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection, 2017
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Deep visual template-free form parsing, 2019
Brian Davis, Bryan Morse, Scott Cohen, Brian Price, and Chris Tensmeyer · 2019
Earlier work this paper cites.
Publaynet: largest dataset ever for document layout analysis, 2019
Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes · 2019
Earlier work this paper cites.
Language modeling with deep transformers
Kazuki Irie, Albert Zeyer, Ralf Schlüter, and Hermann Ney · 2019
Earlier work this paper cites.
Language models with transformers, 2019
Chenguang Wang, Mu Li, and Alexander J. Smola · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Document layout analysis: A comprehensive survey
Galal M Binmakhashen and Sabri A Mahmoud · 2019
Earlier work this paper cites.
GraphRel: Modeling text as relational graphs for joint entity and relation extraction
Tsu-Jui Fu, Peng-Hsuan Li, and Wei-Yun Ma · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Kawin Ethayarajh · 2019
Earlier work this paper cites.
Enforcing encoder-decoder modularity in sequence-to-sequence models, 2019
Siddharth Dalmia, Abdelrahman Mohamed, Mike Lewis, Florian Metze, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
Evaluating sequence-to-sequence models for handwritten text recognition, 2019
Johannes Michael, Roger Labahn, Tobias Grüning, and Jochen Zöllner · 2019
Earlier work this paper cites.
Flat2layout: Flat representation for estimating layout of general room types, 2019
Chi-Wei Hsiao, Cheng Sun, Min Sun, and Hwann-Tzong Chen · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Jean-Philippe Thiran Guillaume Jaume, Hazim Kemal Ekenel · 2019
Earlier work this paper cites.
National archives forms dataset
Brian Davis, Bryan Morse, Scott Cohen, Brian Price, and Chris Tensmeyer · 2019
Earlier work this paper cites.
Pdfminer
Yusuke Shinyama et al · 2019
Earlier work this paper cites.
Icdar2019 competition on scanned receipt ocr and information extraction
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and C. V. Jawahar · 2019
Earlier work this paper cites.
Cord: A consolidated receipt dataset for post-ocr parsing
Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee · 2019
Earlier work this paper cites.
Scene text visual question answering, 2019
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marçal Rusiñol, Ernest Valveny, C. V. Jawahar, and Dimosthenis Karatzas · 2019
Cited alongside, same era.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou · 2020
Cited alongside, same era.
A survey of deep learning approaches for ocr and document understanding
Nishant Subramani, Alexandre Matton, Malcolm Greaves, and Adrian Lam · 2020
Cited alongside, same era.
Pick: Processing key information extraction from documents using improved graph learning-convolutional networks, 2020
Wenwen Yu, Ning Lu, Xianbiao Qi, Ping Gong, and Rong Xiao · 2020
Cited alongside, same era.
Trie: End-to-end text reading and information extraction for document understanding, 2020
Peng Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu, Liang Qiao, Yi Niu, and Fei Wu · 2020
Cited alongside, same era.
A survey on vision transformer
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al · 2022
Later among the works it cites.
Document image analysis and recognition: a survey
Andreeva Elena Igorevna, Bulatov Konstantin Bulatovich, Nikolaev Dmitry Petrovich, Petrova Olga Olegovna, Savelev Boris Igorevich, and Slavin Oleg Anatolevich · 2022
Later among the works it cites.
Fusion of visual representations for multimodal information extraction from unstructured transactional documents
Berke Oral and Gülşen Eryiğit · 2022
Later among the works it cites.
Multi-modal transformer for accelerated mr imaging, 2022
Chun-Mei Feng, Yunlu Yan, Geng Chen, Yong Xu, Ling Shao, and Huazhu Fu · 2022
Later among the works it cites.
Ernie-layout: Layout knowledge enhanced pre-training for visually-rich document understanding
Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Teng Hu, Weichong Yin, Yongfeng Chen, Yin Zhang, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unilmv2: Pseudo-masked language models for unified language model pre-training, 2020
Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Songhao Piao, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon · 2020
Cited alongside, same era.
Self-supervised relationship probing
Jiuxiang Gu, Jason Kuen, Shafiq Joty, Jianfei Cai, Vlad Morariu, Handong Zhao, and Tong Sun · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Transformer encoder with multi-modal multi-head attention for continuous affect recognition
Haifeng Chen, Dongmei Jiang, and Hichem Sahli · 2020
Cited alongside, same era.
Navigation-based candidate expansion and pretrained language models for citation recommendation
Rodrigo Nogueira, Zhiying Jiang, Kyunghyun Cho, and Jimmy Lin · 2020
Cited alongside, same era.
Revising funsd dataset for key-value detection in document images, 2020
Hieu M. Vu and Diep Thi-Ngoc Nguyen · 2020
Cited alongside, same era.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, and Furu Wei · 2022
Later among the works it cites.
Two-stage multimodality fusion for high-performance text-based visual question answering
Bingjia Li, Jie Wang, Minyi Zhao, and Shuigeng Zhou · 2022
Later among the works it cites.
Formnet: Structural encoding beyond sequential modeling in form document information extraction, 2022
Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, and Tomas Pfister · 2022
Later among the works it cites.
Lilt: A simple yet effective language-independent layout transformer for structured document understanding, 2022
Jiapeng Wang, Lianwen Jin, and Kai Ding · 2022
Later among the works it cites.
Context autoencoder for self-supervised representation learning
Xiaokang Chen, Mingyu Ding, Xiaodi Wang, Ying Xin, Shentong Mo, Yunhao Wang, Shumin Han, Ping Luo, Gang Zeng, and Jingdong Wang · 2022
Later among the works it cites.
Rtformer: Efficient design for real-time semantic segmentation with transformer
Jian Wang, Chenhui Gou, Qiman Wu, Haocheng Feng, Junyu Han, Errui Ding, and Jingdong Wang · 2022
Later among the works it cites.
What is optical character recognition?
Microsoft · 2022
Later among the works it cites.
Dit: Self-supervised pre-training for document image transformer
Junlong Li, Yiheng Xu, Tengchao Lv, Lei Cui, Cha Zhang, and Furu Wei · 2022
Later among the works it cites.
Abdelrahman Abdallah, Mahmoud Abdalla, Mohamed Elkasaby, Yasser Elbendary, and Adam Jatowt · 2023
Later among the works it cites.
A survey on vision transformer
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, Zhaohui Yang, Yiman Zhang, and Dacheng Tao · 2023
Later among the works it cites.
Generator-retriever-generator: A novel approach to open-domain question answering
Abdelrahman Abdallah and Adam Jatowt · 2023
Later among the works it cites.
Exploring the state of the art in legal qa systems
Abdelrahman Abdallah, Bhawna Piryani, and Adam Jatowt · 2023
Later among the works it cites.
Table understanding: Problem overview
Alexey Shigarov · 2023
Later among the works it cites.
Layoutdiffusion: Controllable diffusion model for layout-to-image generation, 2023
Guangcong Zheng, Xianpan Zhou, Xuewei Li, Zhongang Qi, Ying Shan, and Xi Li · 2023
Later among the works it cites.
Deep learning, graph-based text representation and classification: a survey, perspectives and challenges
Phu Pham, Loan TT Nguyen, Witold Pedrycz, and Bay Vo · 2023
Later among the works it cites.
Multimodal learning with transformers: A survey, 2023
Peng Xu, Xiatian Zhu, and David A. Clifton · 2023
Later among the works it cites.
Mutualformer: Multi-modality representation learning via cross-diffusion attention, 2023
Xixi Wang, Xiao Wang, Bo Jiang, Jin Tang, and Bin Luo · 2023
Later among the works it cites.
Unified pretraining framework for document understanding, May 18 2023
Jiuxiang Gu, Ani Nenkova Nenkova, Nikolaos Barmpalios, Vlad Ion Morariu, Tong Sun, Rajiv Bhawanji Jain, Jason Wen Yong Kuen, and Handong Zhao · 2023
Later among the works it cites.
Distillation of encoder-decoder transformers for sequence labelling, 2023
Marco Farina, Duccio Pappadopulo, Anant Gupta, Leslie Huang, Ozan İrsoy, and Thamar Solorio · 2023
Later among the works it cites.
Decoder-only or encoder-decoder? interpreting language model as a regularized encoder-decoder, 2023
Zihao Fu, Wai Lam, Qian Yu, Anthony Man-Cho So, Shengding Hu, Zhiyuan Liu, and Nigel Collier · 2023
Later among the works it cites.
Sequence-to-sequence pre-training with unified modality masking for visual document understanding
Shuwei Feng, Tianyang Zhan, Zhanming Jie, Trung Quoc Luong, and Xiaoran Jin · 2023
Later among the works it cites.
Docformerv2: Local features for document understanding
Srikar Appalaraju, Peng Tang, Qi Dong, Nishant Sankaran, Yichu Zhou, and R Manmatha · 2023
Later among the works it cites.
Unifying vision, text, and layout for universal document processing, 2023
Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Cha Zhang, and Mohit Bansal · 2023
Later among the works it cites.
Language independent neuro-symbolic semantic parsing for form understanding, 2023
Bhanu Prakash Voutharoja, Lizhen Qu, and Fatemeh Shiri · 2023
Later among the works it cites.
Enabling large language models to generate text with citations, 2023
Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen · 2023
Later among the works it cites.
Multi-scale cell-based layout representation for document understanding
Yuzhi Shi, Mijung Kim, and Yeongnam Chae · 2023
Later among the works it cites.
Structextv2: Masked visual-textual prediction for document image pre-training
Yuechen Yu, Yulin Li, Chengquan Zhang, Xiaoqiang Zhang, Zengyuan Guo, Xiameng Qin, Kun Yao, Junyu Han, Errui Ding, and Jingdong Wang · 2023
Later among the works it cites.
Mingliang Zhai, Yulin Li, Xiameng Qin, Chen Yi, Qunyi Xie, Chengquan Zhang, Kun Yao, Yuwei Wu, and Yunde Jia · 2023
Later among the works it cites.
Form-nlu: Dataset for the form language understanding
Yihao Ding, Siqu Long, Jiabin Huang, Kaixuan Ren, Xingxiang Luo, Hyunsuk Chung, and Soyeon Caren Han · 2023
Later among the works it cites.