Fetching the paper…
Reading the bibliography…
Visual question answering on document images that contain textual, visual, and layout information, called document VQA, has received much attention recently.
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
Luo, H.; Ji, L.; Shi, B.; Huang, H.; Duan, N.; Li, T.; Li, J.; Bharti, T.; and Zhou, M. 2020 · 2002
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Robertson, S.; Zaragoza, H.; et al. 2009 · 2009
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren; Shaoqing; He; Kaiming; Girshicka; Ross; Sun; and Jian. 2015 · 2015
Earlier work this paper cites.
Ba, L. J.; Kiros, R.; and Hinton, G. E. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Ruder, S. 2016 · 2016
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
Tapaswi, M.; Zhu, Y.; Stiefelhagen, R.; Torralba, A.; Urtasun, R.; and Fidler, S. 2016 · 2016
Earlier work this paper cites.
SuperAgent: A Customer Service Chatbot for E-commerce Websites
Cui, L.; Huang, S.; Wei, F.; Tan, C.; Duan, C.; and Zhou, M. 2017 · 2017
Earlier work this paper cites.
Are You Smarter Than a Sixth Grader? Textbook Question Answering for Multimodal Machine Comprehension
Kembhavi, A.; Seo, M. J.; Schwenk, D.; Choi, J.; Farhadi, A.; and Hajishirzi, H. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
TVQA: Localized, Compositional Video Question Answering
Lei, J.; Yu, L.; Bansal, M.; and Berg, T. 2018 · 2018
Earlier work this paper cites.
Know What You Don’t Know: Unanswerable Questions for SQuAD
Rajpurkar, P.; Jia, R.; and Liang, P. 2018 · 2018
Earlier work this paper cites.
The Web as a Knowledge-Base for Answering Complex Questions
Talmor, A.; and Berant, J. 2018 · 2018
Earlier work this paper cites.
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W. W.; Salakhutdinov, R.; and Manning, C. D. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
Dua, D.; Wang, Y.; Dasigi, P.; Stanovsky, G.; Singh, S.; and Gardner, M. 2019 · 2019
Cited alongside, same era.
Wise—slide segmentation in the wild
Haurilet, M.; Roitberg, A.; Martinez, M.; and Stiefelhagen, R. 2019 · 2019
Cited alongside, same era.
Academic Reader: An Interactive Question Answering System on Academic Literatures
Hong, Y.; Wang, J.; Jia, Y.; Zhang, W.; and Wang, X. 2019 · 2019
Cited alongside, same era.
SPaSe - Multi-Label Page Segmentation for Presentation Slides
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering
Izacard, G.; and Grave, E. 2021 · 2021
Later among the works it cites.
DocVQA: A Dataset for VQA on Document Images
Mathew, M.; Karatzas, D.; and Jawahar, C. V. 2021 · 2021
Later among the works it cites.
Going full-tilt boogie on document understanding with text-image-layout transformer
Powalski, R.; Borchmann, Ł.; Jurkiewicz, D.; Dwojak, T.; Pietruszka, M.; and Pałka, G. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Later among the works it cites.
D2S: Document-to-Slide Generation Via Query-Based Text Summarization
Sun, E.; Hou, Y.; Wang, D.; Zhang, Y.; and Wang, N. X. R. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Monica Haurilet, Z. A.-H.; and Stiefelhagen, R. 2019 · 2019
Cited alongside, same era.
Visual question answering on image sets
Bansal, A.; Zhang, Y.; and Chellappa, R. 2020 · 2020
Cited alongside, same era.
Injecting Numerical Reasoning Skills into Language Models
Geva, M.; Gupta, A.; and Berant, J. 2020 · 2020
Cited alongside, same era.
Video-Grounded Dialogues with Pretrained Generation Language Models
Le, H.; and Hoi, S. C. H. 2020 · 2020
Cited alongside, same era.
TVQA+: Spatio-Temporal Grounding for Video Question Answering
Lei, J.; Yu, L.; Berg, T.; and Bansal, M. 2020 · 2020
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Cited alongside, same era.
Select, answer and explain: Interpretable multi-hop reading comprehension over multiple documents
Tu, M.; Huang, K.; Wang, G.; Huang, J.; He, X.; and Zhou, B. 2020 · 2020
Cited alongside, same era.
Talmor, A.; Yoran, O.; Catav, A.; Lahav, D.; Wang, Y.; Asai, A.; Ilharco, G.; Hajishirzi, H.; and Berant, J. 2021 · 2021
Later among the works it cites.
VisualMRC: Machine Reading Comprehension on Document Images
Tanaka, R.; Nishida, K.; and Yoshida, S. 2021 · 2021
Later among the works it cites.
Document Collection Visual Question Answering
Tito, R.; Karatzas, D.; and Valveny, E. 2021 · 2021
Later among the works it cites.
LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding
Xu, Y.; Xu, Y.; Lv, T.; Cui, L.; Wei, F.; Wang, G.; Lu, Y.; Florêncio, D. A. F.; Zhang, C.; Che, W.; Zhang, M.; and Zhou, L. 2021 · 2021
Later among the works it cites.
NOAHQA: Numerical Reasoning with Interpretable Graph Question Answering Dataset
Zhang, Q.; Wang, L.; Yu, S.; Wang, S.; Wang, Y.; Jiang, J.; and Lim, E.-P. 2021 · 2021
Later among the works it cites.
DOC2PPT: Automatic Presentation Slides Generation from Scientific Documents
Fu, T.; Wang, W. Y.; McDuff, D.; and Song, Y. 2022 · 2022
Later among the works it cites.
InfographicVQA
Mathew, M.; Bagal, V.; Tito, R.; Karatzas, D.; Valveny, E.; and Jawahar, C. 2022 · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Chi, E.; Le, Q.; and Zhou, D. 2022 · 2022
Later among the works it cites.
Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning Skills
Yoran, O.; Talmor, A.; and Berant, J. 2022 · 2022
Later among the works it cites.