Fetching the paper…
Reading the bibliography…
A large amount of research about multimodal inference across text and vision has been recently developed to obtain visually grounded word and sentence representations.
Scene Graph Generation by Iterative Message Passing
Danfei Xu, Yuke Zhu, Christopher B. Choy, and Li Fei-Fei. 2017 · 1906
Earlier work this paper cites.
Applications of circumscription to formalizing common-sense knowledge
John McCarthy. 1986 · 1986
Earlier work this paper cites.
The Syntactic Process
Mark Steedman. 2000 · 2000
Earlier work this paper cites.
Representation and Inference for Natural Language: A First Course in Computational Semantics
Patrick Blackburn and Johan Bos. 2005 · 2005
Earlier work this paper cites.
DeViSE: A Deep Visual-Semantic Embedding Model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc Aurelio Ranzato, and Tomas Mikolov. 2013 · 2013
Earlier work this paper cites.
Zero-Shot Learning by Convex Combination of Semantic Embeddings
Mohammad Norouzi, Tomas Mikolov, Samy Bengio, Yoram Singer, Jonathon Shlens, Andrea Frome, Greg Corrado, and Jeffrey Dean. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Combining lexical and spatial knowledge to predict spatial relations between objects in images
Manuela Hürlimann and Johan Bos. 2016 · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei. 2017 · 2017
Cited alongside, same era.
On-demand Injection of Lexical Knowledge for Recognising Textual Entailment
Pascual Martínez-Gómez, Koji Mineshima, Yusuke Miyao, and Daisuke Bekki. 2017 · 2017
Cited alongside, same era.
Comparatives, quantifiers, proportions: a multi-task model for the learning of quantities from vision
Sandro Pezzelle, Ionut-Teodor Sorodoc, and Raffaella Bernardi. 2018 · 2018
Later among the works it cites.
Grounded textual entailment
Hoa Trong Vu, Claudio Greco, Aliia Erofeeva, Somayeh Jafaritazehjan, Guido Linders, Marc Tanti, Alberto Testoni, Raffaella Bernardi, and Albert Gatt. 2018 · 2018
Later among the works it cites.
Visual entailment task for visually-grounded language learning
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2018 · 2018
Later among the works it cites.
TallyQA: Answering complex counting questions
Manoj Acharya, Kushal Kafle, and Christopher Kanan. 2019 · 2019
Closest in time.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A. Hudson and Christopher D. Manning. 2019 · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A corpus of natural language for visual reasoning
Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi. 2017 · 2017
Cited alongside, same era.
Graph-structured representations for visual question answering
Damien Teney, Lingqiao Liu, and Anton van den Hengel. 2017 · 2017
Cited alongside, same era.
Attentive relational networks for mapping images to scene graphs
Mengshi Qi, Weijian Li, Zhengyuan Yang, Yunhong Wang, and Jiebo Luo. 2019 · 2019
Closest in time.
Higher-order logical inference with compositional semantics
Koji Mineshima, Pascual Martínez-Gómez, Yusuke Miyao, and Daisuke Bekki. 2015 · 2061
Closest in time.