Fetching the paper…
Reading the bibliography…
We propose DocFormerv2, a multi-modal transformer for Visual Document Understanding (VDU).
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
Icdar 2019 competition on scene text visual question answering
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marçal Rusinol, Minesh Mathew, CV Jawahar, Ernest Valveny, and Dimosthenis Karatzas · 2019
Earlier work this paper cites.
Scene text visual question answering
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marçal Rusinol, Ernest Valveny, CV Jawahar, and Dimosthenis Karatzas · 2019
Earlier work this paper cites.
Tabfact: A large-scale dataset for table-based fact verification
Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, SHIYANG LI, Xiyou Zhou, and William Yang Wang · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Earlier work this paper cites.
Cord: A consolidated receipt dataset for post-ocr parsing
Park Seunghyun, Shin Seung, Lee Bado, Lee Junyeop, Surh Jaeheung, Seo Minjoon, and Lee Hwalsuk · 2019
Earlier work this paper cites.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf Thomas, Debut Lysandre, Sanh Victor, Chaumond Julien, Delangue Clement, Moi Anthony, Cistac Pierric, Rault Tim, Louf Rémi, Morgan Funtowicz, et al · 2019
Earlier work this paper cites.
Towards good practices in self-supervised representation learning
Srikar Appalaraju, Yi Zhu, Yusheng Xie, and István Fehérvári · 2020
Earlier work this paper cites.
Unilmv2: Pseudo-masked language models for unified language model pre-training, 2020
Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Songhao Piao, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Finding the evidence: Localization-aware answer prediction for text visual question answering
Wei Han, Hantao Huang, and Tao Han · 2020
Earlier work this paper cites.
Bros: A pre-trained language model for understanding texts in document
Teakgyu Hong, DongHyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park · 2020
Earlier work this paper cites.
Bros: A pre-trained language model for understanding texts in document
Teakgyu Hong, DongHyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park · 2020
Earlier work this paper cites.
Iterative answer prediction with pointer-augmented multimodal transformers for textvqa
Ronghang Hu, Amanpreet Singh, Trevor Darrell, and Marcus Rohrbach · 2020
Earlier work this paper cites.
Spatial dependency parsing for semi-structured document information extraction, 2020
Wonseok Hwang, Jinyeong Yim, Seunghyun Park, Sohee Yang, and Minjoon Seo · 2020
Cited alongside, same era.
Spatially aware multimodal transformers for textvqa
Yash Kant, Dhruv Batra, Peter Anderson, Alexander Schwing, Devi Parikh, Jiasen Lu, and Harsh Agrawal · 2020
Cited alongside, same era.
Scatter: selective context attentional scene text recognizer
Ron Litman, Oron Anschel, Shahar Tsiper, Roee Litman, Shai Mazor, and R Manmatha · 2020
Cited alongside, same era.
DocVQA: A dataset for vqa on document images, 2020
Minesh Mathew, Dimosthenis Karatzas, R. Manmatha, and C. V. Jawahar · 2020
Cited alongside, same era.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou · 2020
Cited alongside, same era.
Layoutlmv2: Multi-modal pre-training for visually-rich document understanding
Latr: Layout-aware transformer for scene-text vqa
Ali Furkan Biten, Ron Litman, Yusheng Xie, Srikar Appalaraju, and R Manmatha · 2022
Later among the works it cites.
Xdoc: Unified pre-training for cross-format document understanding
Jingye Chen, Tengchao Lv, Lei Cui, Changrong Zhang, and Furu Wei · 2022
Later among the works it cites.
Pali: A jointly-scaled multilingual language-image model
Xi Chen, Xiao Wang, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel M. Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish V. Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut · 2022
Later among the works it cites.
End-to-end document recognition and understanding with dessurt
Brian L. Davis, B. Morse, Bryan Price, Chris Tensmeyer, Curtis Wigington, and Vlad I. Morariu · 2022
Later among the works it cites.
Towards escaping from language bias and ocr error: Semantics-centered text visual question answering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, et al · 2020
Cited alongside, same era.
Docformer: End-to-end transformer for document understanding
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R Manmatha · 2021
Cited alongside, same era.
Ocr-free document understanding transformer
Geewook Kim, Teakgyu Hong, Moonbin Yim, Jeongyeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park · 2021
Cited alongside, same era.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Cited alongside, same era.
Structurallm: Structural pre-training for form understanding
Chenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si · 2021
Cited alongside, same era.
Markuplm: Pre-training of text and markup language for visually rich document understanding
Junlong Li, Yiheng Xu, Lei Cui, and Furu Wei · 2021
Cited alongside, same era.
Selfdoc: Self-supervised document representation learning
Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, R. Jain, Varun Manjunatha, and Hongfu Liu · 2021
Cited alongside, same era.
Chengyang Fang, Gangyan Zeng, Yu Zhou, Daiqing Wu, Can Ma, Dayong Hu, and Weiping Wang · 2022
Later among the works it cites.
Unified pretraining framework for document understanding
Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Nikolaos Barmpalios, R. Jain, Ani Nenkova, and Tong Sun · 2022
Later among the works it cites.
Xylayoutlm: Towards layout-aware multimodal networks for visually-rich document understanding
Zhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan, Weiqiang Wang, Ming Gu, and Liqing Zhang · 2022
Later among the works it cites.
Yoro-lightweight end to end visual grounding
Chih-Hui Ho, Srikar Appalaraju, Bhavan Jasani, R Manmatha, and Nuno Vasconcelos · 2022
Later among the works it cites.
Layoutlmv3: Pre-training for document ai with unified text and image masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei · 2022
Later among the works it cites.
Prestu: Pre-training for scene-text understanding
Jihyung Kil, Soravit Changpinyo, Xi Chen, Hexiang Hu, Sebastian Goodman, Wei-Lun Chao, and Radu Soricut · 2022
Later among the works it cites.
Formnet: Structural encoding beyond sequential modeling in form document information extraction
Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, and Tomas Pfister · 2022
Later among the works it cites.
Pix2struct: Screenshot parsing as pretraining for visual language understanding
Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova · 2022
Later among the works it cites.
Two-stage multimodality fusion for high-performance text-based visual question answering
Bingjia Li, Jie Wang, Minyi Zhao, and Shuigeng Zhou · 2022
Later among the works it cites.
Seetek: Very large-scale open-set logo recognition with text-aware metric learning
Chenge Li, István Fehérvári, Xiaonan Zhao, Ives Macêdo, and Srikar Appalaraju · 2022
Later among the works it cites.
Scenegate: Scene-graph based co-attention networks for text visual question answering
Siwen Luo, Feiqi Cao, Felipe Weason Núñez, Zean Wen, Josiah Poon, and Caren Han · 2022
Later among the works it cites.
Infographicvqa
Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar · 2022
Later among the works it cites.
Ernie-layout: Layout knowledge enhanced pre-training for visually-rich document understanding
Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Teng Hu, Weichong Yin, Yongfeng Chen, Yin Zhang, et al · 2022
Later among the works it cites.
Unifying vision, text, and layout for universal document processing
Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Chao-Yue Zhang, and Mohit Bansal · 2022
Later among the works it cites.
Tag: Boosting text-vqa via text-aware visual question-answer generation
Jun Wang, Mingfei Gao, Yuqian Hu, Ramprasaath R Selvaraju, Chetan Ramaiah, Ran Xu, Joseph F JaJa, and Larry S Davis · 2022
Later among the works it cites.
Lilt: A simple yet effective language-independent layout transformer for structured document understanding
Jiapeng Wang, Lianwen Jin, and Kai Ding · 2022
Later among the works it cites.
Git: A generative image-to-text transformer for vision and language
Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang · 2022
Later among the works it cites.
A benchmark for structured extractions from complex documents
Zilong Wang, Yichao Zhou, Wei Wei, Chen-Yu Lee, and Sandeep Tata · 2022
Later among the works it cites.
Mixgen: A new multi-modal data augmentation
Xiaoshuai Hao, Yi Zhu, Srikar Appalaraju, Aston Zhang, Wanqian Zhang, Boyang Li, and Mu Li · 2023
Closest in time.