Fetching the paper…
Reading the bibliography…
Visual question answering (VQA) requires systems to perform concept-level reasoning by unifying unstructured (e.g., the context in question and answer; "QA context") and structured (e.g., knowledge graph for the QA context and scene; "concept graph") multimodal knowledge.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2016
Earlier work this paper cites.
Mutan: Multimodal tucker fusion for visual question answering
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
Robyn Speer, Joshua Chin, and Catherine Havasi · 2017
Earlier work this paper cites.
Graph-structured representations for visual question answering
Damien Teney, Lingqiao Liu, and Anton van Den Hengel · 2017
Earlier work this paper cites.
Bilinear attention networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang · 2018
Earlier work this paper cites.
Learning conditioned graph structures for interpretable visual question answering
Will Norcliffe-Brown, Efstathios Vafeias, and Sarah Parisot · 2018
Earlier work this paper cites.
Graph Attention Networks
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio · 2018
Earlier work this paper cites.
Relational graph attention networks
Dan Busbridge, Dane Sherburn, Pietro Cavallo, and Nils Y Hammerla · 2019
Earlier work this paper cites.
Language-conditioned graph networks for relational reasoning
Ronghang Hu, Anna Rohrbach, Trevor Darrell, and Kate Saenko · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A. Hudson and Christopher D. Manning · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Cited alongside, same era.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Cited alongside, same era.
Kvqa: Knowledge-aware visual question answering
Parth Shah, Prateek Shenoy, Shivansh Bhattad, Prathamesh Bhingardive, Abir Chakraborty, and Subbarao Kambhampati · 2019
Cited alongside, same era.
From strings to things: Knowledge-enabled vqa model that can read and reason
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, and Devi Parikh · 2019
Cited alongside, same era.
How powerful are graph neural networks?
An empirical study of gpt-3 for few-shot knowledge-based vqa
Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang · 2021
Later among the works it cites.
Qa-gnn: Reasoning with language models and knowledge graphs for question answering
Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec · 2021
Later among the works it cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2021
Later among the works it cites.
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi · 2021
Later among the works it cites.
Vinvl: Making visual representations matter in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka · 2019
Cited alongside, same era.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Scalable multi-hop relational reasoning for knowledge-aware question answering
Yanlin Feng, Xinyue Chen, Bill Yuchen Lin, Peifeng Wang, Jun Yan, and Xiang Ren · 2020
Cited alongside, same era.
ConceptBert: Concept-aware representation for visual question answering
François Gardères, Maryam Ziaeefard, Baptiste Abeloos, and Freddy Lecue · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Xiaowei Hu, Pengchuan Zhang, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao · 2020
Cited alongside, same era.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2020
Cited alongside, same era.
Explicit knowledge incorporation for visual reasoning
Yifeng Zhang, Ming Jiang, and Qi Zhao · 2021
Later among the works it cites.
Knowledge-routed visual question reasoning: Challenges for deep representation embedding
Qingxing Cao, Bailin Li, Xiaodan Liang, Keze Wang, and Liang Lin · 2022
Closest in time.
Mukea: Multimodal knowledge extraction and accumulation for knowledge-based visual question answering
Yang Ding, Jing Yu, Bang Liu, Yue Hu, Mingxin Cui, and Qi Wu · 2022
Closest in time.
Dynamic key-value memory enhanced multi-step graph reasoning for knowledge-based visual question answering
Mingxiao Li and Marie-Francine Moens · 2022
Closest in time.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Zou · 2022
Closest in time.
Unsupervised vision-language parsing: Seamlessly bridging visual scene graphs with language structures via dependency relationships
Chao Lou, Wenjuan Han, Yuhuan Lin, and Zilong Zheng · 2022
Closest in time.
Coarse-to-fine reasoning for visual question answering
Binh X Nguyen, Tuong Do, Huy Tran, Erman Tjiputra, Quang D Tran, and Anh Nguyen · 2022
Closest in time.
Vae-based adversarial multimodal domain transfer for video-level sentiment analysis
Yanan Wang, Jianming Wu, Kazuaki Furumai, Shinya Wada, and Satoshi Kurihara · 2022
Closest in time.
Sgeitl: Scene graph enhanced image-text learning for visual commonsense reasoning
Zhecan Wang, Haoxuan You, Liunian Harold Li, Alireza Zareian, Suji Park, Yiqing Liang, Kai-Wei Chang, and Shih-Fu Chang · 2022
Closest in time.
Multi-modal answer validation for knowledge-based vqa
Jialin Wu, Jiasen Lu, Ashish Sabharwal, and Roozbeh Mottaghi · 2022
Closest in time.
Deep bidirectional language-knowledge graph pretraining
Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang, Christopher D. Manning, Percy Liang, and Jure Leskovec · 2022
Closest in time.
Merlot reserve: Multimodal neural script knowledge through vision and language and sound
Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi · 2022
Closest in time.
Retrieval-augmented multimodal language modeling
Michihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Rich James, Jure Leskovec, Percy Liang, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih · 2023
Closest in time.