TextCaps: a Dataset for Image Captioning with Reading Comprehension
Original
Sidorov, O.; Hu, R.; Rohrbach, M.; and Singh, A. 2020 · 2003
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Original
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Im2Text: Describing Images Using 1 Million Captioned Photographs
Ordonez, V.; Kulkarni, G.; and Berg, T. 2011 · 2011
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P.; Lai, A.; Hodosh, M.; and Hockenmaier, J. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Original
Lin, T.-Y.; Maire, M.; Belongie, S.; Bourdev, L.; Girshick, R.; Hays, J.; Perona, P.; Ramanan, D.; Zitnick, C. L.; and Dollár, P. 2015 · 2015
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Vedantam, R.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Visual Dialog
Das, A.; Kottur, S.; Gupta, K.; Singh, A.; Yadav, D.; Moura, J. M.; Parikh, D.; and Batra, D. 2017 · 2017
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Original
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Video Question Answering via Gradually Refined Attention over Appearance and Motion
Xu, D.; Zhao, Z.; Xiao, J.; Wu, F.; Zhang, H.; He, X.; and Zhuang, Y. 2017 · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D.; Li, Q.; Stangl, A. J.; Guo, A.; Lin, C.; Grauman, K.; Luo, J.; and Bigham, J. P. 2018 · 2018
Earlier work this paper cites.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Earlier work this paper cites.
Icdar2019 competition on scanned receipt ocr and information extraction
Huang, Z.; Chen, K.; He, J.; Bai, X.; Karatzas, D.; Lu, S.; and Jawahar, C. 2019 · 2019
Earlier work this paper cites.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Hudson, D. A.; and Manning, C. D. 2019 · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Jaume, G.; Ekenel, H. K.; and Thiran, J.-P. 2019 · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K.; Rastegari, M.; Farhadi, A.; and Mottaghi, R. 2019 · 2019
Earlier work this paper cites.
OCR-VQA: Visual Question Answering by Reading Text in Images
Mishra, A.; Shekhar, S.; Singh, A. K.; and Chakraborty, A. 2019 · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A.; Natarajan, V.; Shah, M.; Jiang, Y.; Chen, X.; Batra, D.; Parikh, D.; and Rohrbach, M. 2019 · 2019
Earlier work this paper cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Kiela, D.; Firooz, H.; Mohan, A.; Goswami, V.; Singh, A.; Ringshia, P.; and Testuggine, D. 2020 · 2020
Earlier work this paper cites.