Fetching the paper…
Reading the bibliography…
Recent studies on machine reading comprehension have focused on text-level understanding but have not yet reached the level of human understanding of the visual layout and content of real-world documents.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Scao, T. L.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. M. 2019 · 1910
Earlier work this paper cites.
Bridging Text and Video: A Universal Multimodal Transformer for Video-Audio Scene-Aware Dialog
Li, Z.; Li, Z.; Zhang, J.; Feng, Y.; Niu, C.; and Zhou, J. 2020b · 2002
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
Latent Dirichlet Allocation
Blei, D. M.; Ng, A. Y.; and Jordan, M. I. 2003 · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
An Overview of the Tesseract OCR Engine
Smith, R. 2007 · 2007
Earlier work this paper cites.
Meteor Universal: Language Specific Translation Evaluation for Any Target Language
Denkowski, M. J.; and Lavie, A. 2014 · 2014
Earlier work this paper cites.
GROTOAP2 - The Methodology of Creating a Large Ground Truth Dataset of Scientific Articles
Tkaczyk, D.; Szostek, P.; and Bolikowski, L. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren; Shaoqing; He; Kaiming; Girshicka; Ross; Sun; and Jian. 2015 · 2015
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Vedantam, R.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Layer Normalization
Ba, L. J.; Kiros, R.; and Hinton, G. E. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy; Christian; Vincent, V.; Ioffe; Sergey; Shlens; Jon; Wojna; and Zbigniew. 2016 · 2016
Earlier work this paper cites.
Enriching word vectors with subword information
Bojanowski; Piotr; Grave; Edouard; Joulin; Armand; Mikolov; and Tomas. 2017 · 2017
Earlier work this paper cites.
SuperAgent: A Customer Service Chatbot for E-commerce Websites
Cui, L.; Huang, S.; Wei, F.; Tan, C.; Duan, C.; and Zhou, M. 2017 · 2017
Cited alongside, same era.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Cited alongside, same era.
Are You Smarter Than a Sixth Grader? Textbook Question Answering for Multimodal Machine Comprehension
Kembhavi, A.; Seo, M. J.; Schwenk, D.; Choi, J.; Farhadi, A.; and Hajishirzi, H. 2017 · 2017
Cited alongside, same era.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.; Shamma, D. A.; Bernstein, M. S.; and Fei-Fei, L. 2017 · 2017
Cited alongside, same era.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Natural Questions: a Benchmark for Question Answering Research
Kwiatkowski, T.; Palomaki, J.; Redfield, O.; Collins, M.; Parikh, A. P.; Alberti, C.; Epstein, D.; Polosukhin, I.; Devlin, J.; Lee, K.; Toutanova, K.; Jones, L.; Kelcey, M.; Chang, M.; Dai, A. M.; Uszkoreit, J.; Le, Q.; and Petrov, S. 2019 · 2019
Later among the works it cites.
OCR-VQA: Visual Question Answering by Reading Text in Images
Mishra, A.; Shekhar, S.; Singh, A. K.; and Chakraborty, A. 2019 · 2019
Later among the works it cites.
Towards VQA Models That Can Read
Singh, A.; Natarajan, V.; Shah, M.; Jiang, Y.; Chen, X.; Batra, D.; Parikh, D.; and Rohrbach, M. 2019 · 2019
Later among the works it cites.
Visual Detection with Context for Document Layout Analysis
Soto, C.; and Yoo, S. 2019 · 2019
Later among the works it cites.
UNITER: Learning UNiversal Image-TExt Representations
Chen, Y.; Li, L.; Yu, L.; Kholy, A. E.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to Extract Semantic Structure from Documents Using Multimodal Fully Convolutional Neural Networks
Yang, X.; Yumer, E.; Asente, P.; Kraley, M.; Kifer, D.; and Giles, C. L. 2017 · 2017
Cited alongside, same era.
VizWiz Grand Challenge: Answering Visual Questions From Blind People
Gurari, D.; Li, Q.; Stangl, A. J.; Guo, A.; Lin, C.; Grauman, K.; Luo, J.; and Bigham, J. P. 2018 · 2018
Cited alongside, same era.
DVQA: Understanding Data Visualizations via Question Answering
Kafle, K.; Price, B. L.; Cohen, S.; and Kanan, C. 2018 · 2018
Cited alongside, same era.
FigureQA: An Annotated Figure Dataset for Visual Reasoning
Kahou, S. E.; Michalski, V.; Atkinson, A.; Kádár, Á.; Trischler, A.; and Bengio, Y. 2018 · 2018
Cited alongside, same era.
Chargrid: Towards Understanding 2D Documents
Katti, A. R.; Reisswig, C.; Guder, C.; Brarda, S.; Bickel, S.; Höhne, J.; and Faddoul, J. B. 2018 · 2018
Cited alongside, same era.
The NarrativeQA Reading Comprehension Challenge
Kociský, T.; Schwarz, J.; Blunsom, P.; Dyer, C.; Hermann, K. M.; Melis, G.; and Grefenstette, E. 2018 · 2018
Cited alongside, same era.
Know What You Don’t Know: Unanswerable Questions for SQuAD
Rajpurkar, P.; Jia, R.; and Liang, P. 2018 · 2018
Cited alongside, same era.
Hu, R.; Singh, A.; Darrell, T.; and Rohrbach, M. 2020 · 2020
Later among the works it cites.
Video-Grounded Dialogues with Pretrained Generation Language Models
Le, H.; and Hoi, S. C. H. 2020 · 2020
Later among the works it cites.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; and Zettlemoyer, L. 2020 · 2020
Later among the works it cites.
RikiNet: Reading Wikipedia Pages for Natural Question Answering
Liu, D.; Gong, Y.; Fu, J.; Yan, Y.; Chen, J.; Jiang, D.; Lv, J.; and Duan, N. 2020 · 2020
Later among the works it cites.
12-in-1: Multi-Task Vision and Language Representation Learning
Lu, J.; Goswami, V.; Rohrbach, M.; Parikh, D.; and Lee, S. 2020 · 2020
Later among the works it cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Later among the works it cites.
On the General Value of Evidence, and Bilingual Scene-Text Visual Question Answering
Wang, X.; Liu, Y.; Shen, C.; Ng, C. C.; Luo, C.; Jin, L.; Chan, C. S.; van den Hengel, A.; and Wang, L. 2020 · 2020
Later among the works it cites.
LayoutLM: Pre-training of Text and Layout for Document Image Understanding
Xu, Y.; Li, M.; Cui, L.; Huang, S.; Wei, F.; and Zhou, M. 2020 · 2020
Later among the works it cites.
BERTScore: Evaluating Text Generation with BERT
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2020 · 2020
Later among the works it cites.
Unified Vision-Language Pre-Training for Image Captioning and VQA
Zhou, L.; Palangi, H.; Zhang, L.; Hu, H.; Corso, J. J.; and Gao, J. 2020 · 2020
Later among the works it cites.
DocVQA: A Dataset for VQA on Document Images
Mathew, M.; Karatzas, D.; and Jawahar, C. V. 2021 · 2021
Closest in time.