Fetching the paper…
Reading the bibliography…
In this paper, we aim to understand whether current language and vision (LaVi) models truly grasp the interaction between the two modalities.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013 · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Comparing automatic evaluation measures for image description
Desmond Elliott and Frank Keller. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Mateusz Malinowski and Mario Fritz. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Mind’s eye: A recurrent visual representation for image caption generation
Xinlei Chen and C Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015 · 2015
Earlier work this paper cites.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh K Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C Platt, et al. 2015 · 2015
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu. 2015 · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Ask your neurons: A neural-based approach to answering questions about images
Mateusz Malinowski, Marcus Rohrbach, and Mario Fritz. 2015 · 2015
Cited alongside, same era.
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard Zemel. 2015 · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
Focused evaluation for image description with binary forced-choice tasks
Micah Hodosh and Julia Hockenmaier. 2016 · 2016
Later among the works it cites.
Revisiting visual question answering baselines
Allan Jabri, Armand Joulin, and Laurens van der Maaten. 2016 · 2016
Later among the works it cites.
CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2016 · 2016
Later among the works it cites.
Visual question answering: Datasets, algorithms, and future challenges
Kushal Kafle and Christopher Kanan. 2016 · 2016
Later among the works it cites.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016 · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bolei Zhou, Yuandong Tian, Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus. 2015 · 2015
Cited alongside, same era.
Analyzing the behavior of visual question answering models
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh. 2016 · 2016
Cited alongside, same era.
SPICE: Semantic Propositional Image Caption Evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Cited alongside, same era.
Automatic description generation from images: A survey of models, datasets, and evaluation measures
Raffaella Bernardi, Ruket Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, and Barbara Plank. 2016 · 2016
Cited alongside, same era.
Understanding image and text simultaneously: a dual vision-language machine comprehension task
Nan Ding, Sebastian Goodman, Fei Sha, and Radu Soricut. 2016 · 2016
Cited alongside, same era.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2016a
Cited in the paper.
Towards Transparent AI Systems: Interpreting Visual Question Answering Models
Yash Goyal, Akrit Mohapatra, Devi Parikh, and Dhruv Batra. 2016b
Cited in the paper.
Attentive explanations: Justifying decisions and pointing to the evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Trevor Darrell Bernt Schiele, and Marcus Rohrbach. 2016 · 2016
Later among the works it cites.
Ramprasaath R. Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra. 2016 · 2016
Later among the works it cites.
Image captioning with deep bidirectional LSTMs
Cheng Wang, Haojin Yang, Christian Bartz, and Christoph Meinel. 2016 · 2016
Later among the works it cites.
Visual question answering: A survey of methods and datasets
Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton van den Hengel. 2016 · 2016
Later among the works it cites.
Yin and yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2016 · 2016
Later among the works it cites.