Fetching the paper…
Reading the bibliography…
Compared to general document analysis tasks, form document structure understanding and retrieval are challenging.
A coefficient of agreement for nominal scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Pointer networks
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly. 2015 · 2015
Earlier work this paper cites.
The Informativeness of Substantial Shareholder Trading in the Lead up to a Takeover Bid. In Asian Finance Association (AsianFA) 2016 Conference
Millicent Chang, Raymond da Silva Rosa, and Wilson Ng. 2016 · 2016
Earlier work this paper cites.
Mask r-cnn. In Proceedings of the IEEE international conference on computer vision . 2961–2969
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Complicated table structure recognition
Zewen Chi, Heyan Huang, Heng-Da Xu, Houjin Yu, Wanxuan Yin, and Xian-Ling Mao. 2019 · 2019
Earlier work this paper cites.
Icdar 2019 robust reading challenge on scanned receipts ocr and information extraction. In International Conference on Document Analysis Recognition
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and CV Jawahar. 2019 · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents. In 2019 International Conference on Document Analysis and Recognition Workshops (ICDARW) , Vol. 2. IEEE, 1–6
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019 · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 2019
Earlier work this paper cites.
CORD: a consolidated receipt dataset for post-OCR parsing. In Workshop on Document Intelligence at NeurIPS 2019
Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee. 2019 · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes. 2019 · 2019
Cited alongside, same era.
Iterative answer prediction with pointer-augmented multimodal transformers for textvqa. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9992–10002
Ronghang Hu, Amanpreet Singh, Trevor Darrell, and Marcus Rohrbach. 2020 · 2020
Cited alongside, same era.
DocBank: A Benchmark Dataset for Document Layout Analysis. In Proceedings of the 28th International Conference on Computational Linguistics . 949–960
Minghao Li, Yiheng Xu, Lei Cui, Shaohan Huang, Furu Wei, Zhoujun Li, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
VisualMRC: Machine Reading Comprehension on Document Images. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 13878–13888
Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida. 2021 · 2021
Later among the works it cites.
Towards robust visual information extraction in real world: new dataset and novel solution. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 2738–2745
Jiapeng Wang, Chongyu Liu, Lianwen Jin, Guozhi Tang, Jiaxin Zhang, Shuaitao Zhang, Qianying Wang, Yaqiang Wu, and Mingxiang Cai. 2021 · 2021
Later among the works it cites.
Layoutxlm: Multimodal pre-training for multilingual visually-rich document understanding
Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, and Furu Wei. 2021a · 2021
Later among the works it cites.
Entity Relation Extraction as Dependency Parsing in Visually Rich Documents
Yue Zhang, Bo Zhang, Rui Wang, Junjie Cao, Chen Li, and Zuyi Bao. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Layoutlm: Pre-training of text and layout for document image understanding. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 1192–1200
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
Image-based table recognition: data, model, and evaluation. In European Conference on Computer Vision . Springer, 564–580
Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno Yepes. 2020 · 2020
Cited alongside, same era.
Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning . PMLR, 5583–5594
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021 · 2021
Cited alongside, same era.
Docvqa: A dataset for vqa on document images. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 2200–2209
Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. 2021 · 2021
Cited alongside, same era.
Docparser: Hierarchical document structure parsing from renderings. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 4328–4338
Johannes Rausch, Octavio Martinez, Fabian Bissig, Ce Zhang, and Stefan Feuerriegel. 2021 · 2021
Cited alongside, same era.
Kleister: key information extraction datasets involving long documents with complex layouts. In International Conference on Document Analysis and Recognition . Springer, 564–579
Tomasz Stanisławek, Filip Graliński, Anna Wróblewska, Dawid Lipiński, Agnieszka Kaliska, Paulina Rosalska, Bartosz Topolski, and Przemysław Biecek. 2021 · 2021
Cited alongside, same era.
LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . 2579–2591
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, et al
Cited in the paper.
Global table extractor (gte): A framework for joint table identification and cell structure recognition using visual context. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 697–706
Xinyi Zheng, Douglas Burdick, Lucian Popa, Xu Zhong, and Nancy Xin Ru Wang. 2021 · 2021
Later among the works it cites.
V-Doc: Visual questions answers with Documents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 21492–21498
Yihao Ding, Zhe Huang, Runlin Wang, YanHang Zhang, Xianru Chen, Yuzhong Ma, Hyunsuk Chung, and Soyeon Caren Han. 2022 · 2022
Later among the works it cites.
Xylayoutlm: Towards layout-aware multimodal networks for visually-rich document understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4583–4592
Zhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan, Weiqiang Wang, Ming Gu, and Liqing Zhang. 2022 · 2022
Later among the works it cites.
LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. 2022 · 2022
Later among the works it cites.
Doc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout Analysis. In Proceedings of the 29th International Conference on Computational Linguistics . 2906–2916
Siwen Luo, Yihao Ding, Siqu Long, Josiah Poon, and Soyeon Caren Han. 2022 · 2022
Later among the works it cites.
XFUND: A Benchmark Dataset for Multilingual Visually Rich Form Understanding. In Findings of the Association for Computational Linguistics: ACL 2022 . 3214–3224
Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, and Furu Wei. 2022 · 2022
Later among the works it cites.