Fetching the paper…
Reading the bibliography…
One of the most challenging question types in VQA is when answering the question requires outside knowledge not present in the image.
Wordnet: A lexical database for english
George A. Miller · 1995
Earlier work this paper cites.
Conceptnet—a practical commonsense reasoning tool-kit
Hugo Liu and Push Singh · 2004
Earlier work this paper cites.
One-shot learning of object categories
Li Fei-Fei, Rob Fergus, and Pietro Perona · 2006
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives · 2007
Earlier work this paper cites.
Describing objects by their attributes
A. Farhadi, I. Endres, D. Hoiem, and D.A. Forsyth · 2009
Earlier work this paper cites.
Evaluating Knowledge Transfer and Zero-Shot Learning in a Large-Scale Setting
Marcus Rohrbach, Michael Stark, and Bernt Schiele · 2011
Earlier work this paper cites.
Discovering localized attributes for fine-grained recognition
Kun Duan, Devi Parikh, David Crandall, and Kristen Grauman · 2012
Earlier work this paper cites.
Constrained semi-supervised learning using attributes and comparative attributes
Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2012
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang · 2013
Earlier work this paper cites.
Neil: Extracting visual knowledge from web data
Xinlei Chen, Abhinav Shrivastava, and Abhinav Gupta · 2013
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S. Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Transfer Learning in a Transductive Setting
Marcus Rohrbach, Sandra Ebert, and Bernt Schiele · 2013
Earlier work this paper cites.
Question answering with subgraph embeddings
Antoine Bordes, Sumit Chopra, and Jason Weston · 2014
Earlier work this paper cites.
Learning everything about anything: Webly-supervised visual concept learning
Santosh Divvala, Ali Farhadi, and Carlos Guestrin · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Attribute-based classification for zero-shot visual object categorization
Christoph H. Lampert, Hannes Nickisch, and Stefan Harmeling · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
T. Lin, M. Maire, S. J. Belongie, R. B. Girshick, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Towards a visual turing challenge
Mateusz Malinowski and Mario Fritz · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Information extraction over structured data: Question answering with freebase
Xuchen Yao and Benjamin Van Durme · 2014
Earlier work this paper cites.
Reasoning about object affordances in a knowledge base representation
Yuke Zhu, Alireza Fathi, and Li Fei-Fei · 2014
Earlier work this paper cites.
VQA: visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David A Shamma, Michael S Bernstein, and Li Fei-Fei · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
Mateusz Malinowski, Marcus Rohrbach, and Mario Fritz · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Viske: Visual knowledge extraction and question answering by visual verification of relation phrases
Fereshteh Sadeghi, Santosh K Divvala, and Ali Farhadi · 2015
Earlier work this paper cites.
Wikiqa: A challenge dataset for open-domain question answering
Yi Yang, Wen-tau Yih, and Christopher Meek · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Building a large-scale multimodal knowledge base for visual question answering
Yuke Zhu, Ce Zhang, Christopher Ré, and Li Fei-Fei · 2015
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2016
Cited alongside, same era.
Out of the box: Reasoning with graph convolution nets for factual visual question answering
Medhini Narasimhan, Svetlana Lazebnik, and Alexander G Schwing · 2018
Later among the works it cites.
Straight to the facts: Learning knowledge base retrieval for factual visual question answering
Medhini Narasimhan and Alexander G. Schwing · 2018
Later among the works it cites.
Modeling relational data with graph convolutional networks
Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne Van Den Berg, Ivan Titov, and Max Welling · 2018
Later among the works it cites.
Zero-shot recognition via semantic embeddings and knowledge graphs
Xiaolong Wang, Yufei Ye, and Abhinav Gupta · 2018
Later among the works it cites.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Satwik Kottur, Ramakrishna Vedantam, José MF Moura, and Devi Parikh · 2016
Cited alongside, same era.
Commonsense knowledge base completion
Xiang Li, Aynaz Taheri, Lifu Tu, and Kevin Gimpel · 2016
Cited alongside, same era.
Gated graph sequence neural networks
Yujia Li and Richard Zemel · 2016
Cited alongside, same era.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Cited alongside, same era.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Qi Wu, Peng Wang, Chunhua Shen, Anthony R. Dick, and Anton van den Hengel · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Cited alongside, same era.
Chris Alberti, Jeffrey Ling, Michael Collins, and David Reitter · 2019
Later among the works it cites.
Uniter: Learning universal image-text representations
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Fast graph representation learning with pytorch geometric
Matthias Fey and Jan Eric Lenssen · 2019
Later among the works it cites.
Dynamic fusion with intra-and inter-modality attention flow for visual question answering
Peng Gao, Zhengkai Jiang, Haoxuan You, Pan Lu, Steven CH Hoi, Xiaogang Wang, and Hongsheng Li · 2019
Later among the works it cites.
LVIS: A dataset for large vocabulary instance segmentation
Agrim Gupta, Piotr Dollar, and Ross Girshick · 2019
Later among the works it cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Daxin Jiang, and Ming Zhou · 2019
Later among the works it cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller · 2019
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2019
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Later among the works it cites.
End-to-end open-domain question answering with bertserini
Wei Yang, Yuqing Xie, Aileen Lin, Xingyu Li, Luchen Tan, Kun Xiong, Ming Li, and Jimmy Lin · 2019
Later among the works it cites.
Unified vision-language pre-training for image captioning and vqa
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J Corso, and Jianfeng Gao · 2019
Later among the works it cites.
Do dogs have whiskers? a new knowledge base of haspart relations
Sumithra Bhakthavatsalam, Kyle Richardson, Niket Tandon, and Peter Clark · 2020
Closest in time.
Conceptbert: Concept-aware representation for visual question answering
François Gardères, Maryam Ziaeefard, Baptiste Abeloos, and Freddy Lecue · 2020
Closest in time.
In defense of grid features for visual question answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik Learned-Miller, and Xinlei Chen · 2020
Closest in time.
How can we know what language models know?
Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig · 2020
Closest in time.
Boosting visual question answering with context-aware knowledge aggregation
Guohao Li, Xin Wang, and Wenwu Zhu · 2020
Closest in time.
12-in-1: Multi-task vision and language representation learning
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee · 2020
Closest in time.
Mmf: A multimodal framework for vision and language research
Amanpreet Singh, Vedanuj Goswami, Vivek Natarajan, Yu Jiang, Xinlei Chen, Meet Shah, Marcus Rohrbach, Dhruv Batra, and Devi Parikh · 2020
Closest in time.
Are we pretraining it right? digging deeper into visio-linguistic pretraining
Amanpreet Singh, Vedanuj Goswami, and Devi Parikh · 2020
Closest in time.
Equalization loss for long-tailed object recognition
Jingru Tan, Changbao Wang, Buyu Li, Quanquan Li, Wanli Ouyang, Changqing Yin, and Junjie Yan · 2020
Closest in time.