Fetching the paper…
Reading the bibliography…
We address the task of detecting foiled image captions, i.e.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017 · 1997
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
VQA: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Déjà image-captions: A corpus of expressive descriptions in repetition
Jianfu Chen, Polina Kuznetsova, David Warren, and Yejin Choi. 2015 · 2015
Earlier work this paper cites.
Are you talking to a machine? Dataset and methods for multilingual image question
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu. 2015 · 2015
Earlier work this paper cites.
Visual Madlibs: Fill in the blank description generation and question answering
Licheng Yu, Eunbyung Park, Alexander C. Berg, and Tamara L. Berg. 2015 · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Generating captions without looking beyond objects
Hendrik Heuer, Christof Monz, and Arnold W. M. Smeulders. 2016 · 2016
Cited alongside, same era.
Focused evaluation for image description with binary forced-choice tasks
Micah Hodosh and Julia Hockenmaier. 2016 · 2016
Cited alongside, same era.
”Why should I trust you?” Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
Vision and language integration: Moving beyond objects
Ravi Shekhar, Sandro Pezzelle, Aurelie Herbelot, Moin Nabi, Enver Sangineto, and Raffaella Bernardi. 2017a
Oracle performance for visual captioning
Li Yao, Nicolas Ballas, Kyunghyun Cho, John R. Smith, and Bengio Yoshua. 2016 · 2016
Later among the works it cites.
Visual7W: Grounded question aswering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2016 · 2016
Later among the works it cites.
YOLO9000: Better, Faster, Stronger
Joseph Redmon and Ali Farhadi. 2017 · 2017
Later among the works it cites.
Image captioning and visual question answering based on attributes and external knowledge
Qi Wu, Chunhua Shen, Peng Wang, Anthony Dick, and Anton van den Hengel. 2018 · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited in the paper.
FOIL it! Find One mismatch between Image and Language caption
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurélie Herbelot, Moin Nabi, Enver Sangineto, and Raffaella Bernardi. 2017b
Cited in the paper.