D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. Ghosh, A. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu et al. , “Icdar 2015 competition on robust reading,” in ICDAR , 2015
2015
Cited alongside, same era.
M. Ren, R. Kiros, and R. Zemel, “Exploring models and data for image question answering,” in NIPS , 2015
2015
Cited alongside, same era.
Y. Patel, L. Gomez, M. Rusinol, and D. Karatzas, “Dynamic lexicon generation for natural scene images,” in ECCV , 2016
2016
Cited alongside, same era.
A. Veit, T. Matera, L. Neumann, J. Matas, and S. Belongie, “Coco-text: Dataset and benchmark for text detection and recognition in natural images,” arXiv preprint arXiv:1601.07140 , 2016
Original
2016
Cited alongside, same era.
C. Xiong, S. Merity, and R. Socher, “Dynamic memory networks for visual and textual question answering,” in ICML , 2016
2016
Cited alongside, same era.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in CVPR , 2017
2017
Cited alongside, same era.
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick, “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,” in CVPR , 2017
2017
Cited alongside, same era.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” IJCV , vol. 123, 2017
2017
Cited alongside, same era.
A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi, “Don’t just assume; look and answer: Overcoming priors for visual question answering,” in CVPR , 2018
2018
Cited alongside, same era.