Fetching the paper…
Reading the bibliography…
Self-supervised pre-training techniques have achieved remarkable progress in Document AI.
Building a Test Collection for Complex Document Information Processing. In SIGIR
D. Lewis, G. Agam, S. Argamon, O. Frieder, D. Grossman, and J. Heard. 2006 · 2006
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval. In ICDAR
Adam W Harley, Alex Ufkes, and Konstantinos G Derpanis. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Neural Architectures for Named Entity Recognition. In NAACL HLT
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units. In ACL
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Feature pyramid networks for object detection. In CVPR
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2017 · 2017
Earlier work this paper cites.
PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications. In ICLR
Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P. Kingma. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In NeurIPS
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Cascade r-cnn: Delving into high quality object detection. In CVPR
Zhaowei Cai and Nuno Vasconcelos. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents. In ICDARW
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In NeurIPS
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Earlier work this paper cites.
CORD: A Consolidated Receipt Dataset for Post-OCR Parsing. In Document Intelligence Workshop at Neural Information Processing Systems
Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee. 2019 · 2019
Earlier work this paper cites.
A graph-based framework for information extraction
Yujie Qian. 2019 · 2019
Earlier work this paper cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations. In ICLR
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2019 · 2019
Earlier work this paper cites.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers. In EMNLP
Hao Tan and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
Detectron2
Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. 2019 · 2019
Cited alongside, same era.
PubLayNet: largest dataset ever for document layout analysis. In ICDAR
Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes. 2019 · 2019
Cited alongside, same era.
Uniter: Universal image-text representation learning. In ECCV
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Cited alongside, same era.
X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers. In EMNLP
Jaemin Cho, Jiasen Lu, Dustin Schwenk, Hannaneh Hajishirzi, and Aniruddha Kembhavi. 2020 · 2020
Cited alongside, same era.
Unsupervised Cross-lingual Representation Learning at Scale. In ACL
LAMBERT: Layout-Aware Language Modeling for Information Extraction. In ICDAR
Łukasz Garncarek, Rafał Powalski, Tomasz Stanisławek, Bartosz Topolski, Piotr Halama, Michał Turski, and Filip Graliński. 2021 · 2021
Later among the works it cites.
UniDoc: Unified Pretraining Framework for Document Understanding. In NeurIPS
Jiuxiang Gu, Jason Kuen, Vlad Morariu, Handong Zhao, Rajiv Jain, Nikolaos Barmpalios, Ani Nenkova, and Tong Sun. 2021 · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision. In ICML
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021 · 2021
Later among the works it cites.
Benchmarking detection transfer learning with vision transformers
Yanghao Li, Saining Xie, Xinlei Chen, Piotr Dollar, Kaiming He, and Ross Girshick. 2021e · 2021
Later among the works it cites.
Docvqa: A dataset for vqa on document images. In WACV
Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Cited alongside, same era.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In ACL
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Subhojeet Pramanik, Shashank Mujumdar, and Hima Patel. 2020 · 2020
Cited alongside, same era.
Layoutlm: Pre-training of text and layout for document image understanding. In KDD
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
TRIE: end-to-end text reading and information extraction for document understanding. In ACM Multimedia
Peng Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu, Liang Qiao, Yi Niu, and Fei Wu. 2020 · 2020
Cited alongside, same era.
Xcit: Cross-covariance image transformers. In NeurIPS
Alaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek, et al · 2021
Cited alongside, same era.
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer. In ICDAR
Rafal Powalski, Łukasz Borchmann, Dawid Jurkiewicz, Tomasz Dwojak, Michal Pietruszka, and Gabriela Pałka. 2021 · 2021
Later among the works it cites.
Zero-shot text-to-image generation. In ICML
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Towards robust visual information extraction in real world: new dataset and novel solution. In AAAI
Jiapeng Wang, Chongyu Liu, Lianwen Jin, Guozhi Tang, Jiaxin Zhang, Shuaitao Zhang, Qianying Wang, Yaqiang Wu, and Mingxiang Cai. 2021 · 2021
Later among the works it cites.
LAMPRET: Layout-Aware Multimodal PreTraining for Document Understanding
Te-Lin Wu, Cheng Li, Mingyang Zhang, Tao Chen, Spurthi Amba Hombaiah, and Michael Bendersky. 2021 · 2021
Later among the works it cites.
LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding
Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, and Furu Wei. 2021a · 2021
Later among the works it cites.
Probing Inter-modality: Visual Parsing with Self-Attention for Vision-and-Language Pre-training. In NeurIPS
Hongwei Xue, Yupan Huang, Bei Liu, Houwen Peng, Jianlong Fu, Houqiang Li, and Jiebo Luo. 2021 · 2021
Later among the works it cites.
BEiT: BERT Pre-Training of Image Transformers. In ICLR
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. 2022 · 2022
Closest in time.
XYLayoutLM: Towards Layout-Aware Multimodal Networks For Visually-Rich Document Understanding. In CVPR
Zhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan, Weiqiang Wang, Ming Gu, and Liqing Zhang. 2022 · 2022
Closest in time.
BROS: A Pre-Trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents. In AAAI
Teakgyu Hong, DongHyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park. 2022 · 2022
Closest in time.
FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction. In ACL
Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, and Tomas Pfister. 2022 · 2022
Closest in time.
DiT: Self-supervised Pre-training for Document Image Transformer
Junlong Li, Yiheng Xu, Tengchao Lv, Lei Cui, Cha Zhang, and Furu Wei. 2022 · 2022
Closest in time.
LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding. In ACL
Jiapeng Wang, Lianwen Jin, and Kai Ding. 2022 · 2022
Closest in time.