Fetching the paper…
Reading the bibliography…
We introduce a new dataset for joint reasoning about natural language and images, with a focus on semantic diversity, compositionality, and visual reasoning challenges.
Binary image selection (BISON): Interpretable evaluation of visual grounding
Hexiang Hu, Ishan Misra, and Laurens van der Maaten. 2019 · 1901
Earlier work this paper cites.
Visual entailment: A novel task for fine-grained image understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019 · 1901
Earlier work this paper cites.
Learning distributions over logical forms for referring expression generation
Nicholas FitzGerald, Yoav Artzi, and Luke Zettlemoyer. 2013 · 1925
Earlier work this paper cites.
An analysis of visual question answering algorithms
Kushal Kafle and Christopher Kanan. 2017 · 1973
Earlier work this paper cites.
The measurement of observer agreement for categorical data
J. Richard Landis and Gary Koch. 1977 · 1977
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman. 1990 · 1990
Earlier work this paper cites.
WordNet: A lexical database for English
George A. Miller. 1993 · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross B. Girshick. 2017a · 1997
Earlier work this paper cites.
Walk the talk: Connecting language, knowledge, action in route instructions
Matthew MacMahon, Brian Stankiewics, and Benjamin Kuipers. 2006 · 2006
Earlier work this paper cites.
Natural reference to objects in a visual domain
Margaret Mitchell, Kees van Deemter, and Ehud Reiter. 2010 · 2010
Earlier work this paper cites.
A joint model of language and perception for grounded attribute learning
Cynthia Matuszek, Nicholas FitzGerald, Luke S. Zettlemoyer, Liefeng Bo, and Dieter Fox. 2012 · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Bringing semantics into focus using visual abstraction
C. Lawrence Zitnick and Devi Parikh. 2013 · 2013
Earlier work this paper cites.
Scalable multi-label annotation
Jia Deng, Olga Russakovsky, Jonathan Krause, Michael S. Bernstein, Alex Berg, and Li Fei-Fei. 2014 · 2014
Earlier work this paper cites.
ReferItGame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
VQA: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Effectively crowdsourcing radiology report annotations
Anne Cocos, Aaron Masino, Ting Qian, Ellie Pavlick, and Chris Callison-Burch. 2015 · 2015
Cited alongside, same era.
A survey of current datasets for vision and language research
Francis Ferraro, Nasrin Mostafazadeh, Ting-Hao Huang, Lucy Vanderwende, Jacob Devlin, Michel Galley, and Margaret Mitchell. 2015 · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016 · 2016
Cited alongside, same era.
Towards a dataset for human computer communication via grounded language acquisition
Yonatan Bisk, Daniel Marcu, and William Wong. 2016 · 2016
Cited alongside, same era.
Bootstrap, review, decode: Using out-of-domain textual data to improve image captioning
Detectron
Ross Girshick, Ilija Radosavovic, Georgia Gkioxari, Piotr Dollár, and Kaiming He. 2018 · 2018
Closest in time.
Weakly supervised semantic parsing with abstract examples
Omer Goldman, Veronica Latcinnik, Ehud Nave, Amir Globerson, and Jonathan Berant. 2018 · 2018
Closest in time.
Explainable neural computation via stack neural module networks
Ronghang Hu, Jacob Andreas, Trevor Darrell, and Kate Saenko. 2018 · 2018
Closest in time.
Compositional attention networks for machine reasoning
Drew A. Hudson and Christopher D. Manning. 2018 · 2018
Closest in time.
FigureQA: An annotated figure dataset for visual reasoning
Samira Ebrahimi Kahou, Adam Atkinson, Vincent Michalski, Ákos Kádár, Adam Trischler, and Yoshua Bengio. 2018 · 2018
Closest in time.
TVQA: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg. 2018 · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wenhu Chen, Aurélien Lucchi, and Thomas Hofmann. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan Yuille, and Kevin Murphy. 2016 · 2016
Cited alongside, same era.
Yin and yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2016 · 2016
Cited alongside, same era.
C-VQA: A compositional split of the visual question answering (VQA) v1.0 dataset
Aishwarya Agrawal, Aniruddha Kembhavi, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M.F. Moura, Devi Parikh, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
Mapping instructions to actions in 3D environments with visual goal prediction
Dipendra Misra, Andrew Bennett, Valts Blukis, Eyvind Niklasson, Max Shatkhin, and Yoav Artzi. 2018 · 2018
Closest in time.
Working memory networks: Augmenting memory networks with a relational reasoning module
Juan Pavez, Hector Alllende, and Hector Allende-Cid. 2018 · 2018
Closest in time.
FiLM: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville. 2018 · 2018
Closest in time.
DDRprog: A CLEVR differentiable dynamic reasoning programmer
Joseph Suarez, Justin Johnson, and Fei-Fei Li. 2018 · 2018
Closest in time.
Object ordering with bidirectional matchings for visual reasoning
Hao Tan and Mohit Bansal. 2018 · 2018
Closest in time.
Talk the Walk: Navigating New York City through grounded dialogue
Harm de Vries, Kurt Shuster, Dhruv Batra, Devi Parikh, Jason Weston, and Douwe Kiela. 2018 · 2018
Closest in time.
A dataset and architecture for visual reasoning with a working memory
Robert Guangyu Yang, Igor Ganichev, Xiao Jing Wang, Jonathon Shlens, and David Sussillo. 2018 · 2018
Closest in time.
Cascaded mutual modulation for visual reasoning
Yiqun Yao, Jiaming Xu, Feng Wang, and Bo Xu. 2018 · 2018
Closest in time.
Neural-symbolic VQA: Disentangling reasoning from vision and language understanding
Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum. 2018 · 2018
Closest in time.
TallyQA: Answering complex counting questions
Manoj Acharya, Kushal Kafle, and Christopher Kanan. 2019 · 2019
Closest in time.
"Caption" as a coherence relation: Evidence and implications
Malihe Alikhani and Matthew Stone. 2019 · 2019
Closest in time.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi. 2019 · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
GQA: a new dataset for compositional question answering over real-world images
Drew A. Hudson and Christopher D. Manning. 2019 · 2019
Closest in time.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Closest in time.