Fetching the paper…
Reading the bibliography…
One of the key issues of Visual Question Answering (VQA) is to reason with semantic clues in the visual content under the guidance of the question, how to model relational semantics still remains as a great challenge.
Separating style and content
J.B. Tenebaum and W.T. Freeman · 1997
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Exploring visual relationship for image captioning
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei · 2015
Earlier work this paper cites.
Simple baseline for visual question answering
Bolei Zhou, Yuandong Tian, Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus · 2015
Earlier work this paper cites.
Knowledge graph embedding via dynamic mapping matrix
Guoliang Ji, Shizhu He, Liheng Xu, Kang Liu, and Jun Zhao · 2015
Earlier work this paper cites.
Convolutional networks on graphs for learning molecular fingerprints
David K Duvenaud, Dougal Maclaurin, Jorge Iparraguirre, Rafael Bombarell, Timothy Hirzel, Alán Aspuru-Guzik, and Ryan P Adams · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Vqa: Visual question answering
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. Lawrence Zitnick, Dhruv Batra, and Devi Parikh · 2015
Earlier work this paper cites.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, and Jianfeng Gao · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrel, and Marcus Rohrbach · 2016
Cited alongside, same era.
Image question answering using convolutional neural network with dynamic parameter prediction
Hyeonwoo Noh, Paul Hongsuck Seo, and Han Bohyung · 2016
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2016
Cited alongside, same era.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Tim Lillicrap · 2017
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Closest in time.
Compositional attention networks for machine reasoning
Drew A. Hudson and Christopher D. Manning · 2018
Closest in time.
Representation learning for scene graph completion via jointly structural and visual embedding
Hai Wan, Yonghao Luo, Bo Peng, and Wei-Shi Zheng · 2018
Closest in time.
Visual reasoning by progressive module networks
Seung Wook Kim, Makarand Tapaswi, and Sanja Fidler · 2018
Closest in time.
Iterative visual reasoning beyond convolutions
Xinlei Chen, Li-Jia Li, Li Fei-Fei, and Abhinav Gupta · 2018
Closest in time.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Graph-structured representations for visual question answering
Damien Teney, Lingqiao Liu, and Anton van den Hengel · 2017
Cited alongside, same era.
Visual translation embedding network for visual relation detection
Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang, and Tat-Seng Chua · 2017
Cited alongside, same era.
Gqa: a new dataset for compositional question answering over real-world images
Drew A. Hudson and Christopher D. Manning · 2017
Cited alongside, same era.
Hadamard product for low-rank bilinear pooling
Jin-Hwa Kim, Kyoung-Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang · 2017
Cited alongside, same era.
Neural message passing for quantum chemistry
Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl · 2017
Cited alongside, same era.
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio · 2017
Cited alongside, same era.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Cited alongside, same era.
Closest in time.
Out of the box: Reasoning with graph convolution nets for factual visual question answering
Medhini Narasimhan, Svetlana Lazebnik, and Alexander Schwing · 2018
Closest in time.
Straight to the facts: Learning knowledge base retrieval for factual visual question answering
Medhini Narasimhan and Alexander G Schwing · 2018
Closest in time.
Neural motifs: Scene graph parsing with global context
Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi · 2018
Closest in time.
Graph neural networks: A review of methods and applications
Jie Zhou, Ganqu Cui, Zhengyan Zhang, Cheng Yang, Zhiyuan Liu, and Maosong Sun · 2018
Closest in time.
Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering
Zhou Yu, Jun Yu, Chenchao Xiang, Jianping Fan, and Dacheng Tao · 2018
Closest in time.
Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks
Peng Wang, Qi Wu, Shen Chuanhua Cao, Jiewei, and Lianli Gao · 2019
Closest in time.
Large-scale visual relationship understanding
Ji Zhang, Yannis Kalantidis, Marcus Rohrbach, Manohar Paluri, Ahmed Elgammal, and Mohamed Elhoseiny · 2019
Closest in time.