Fetching the paper…
Reading the bibliography…
Visual Question Answering (VQA) has emerged as one of the most challenging tasks in artificial intelligence due to its multi-modal nature.
Conceptnet—a practical commonsense reasoning tool-kit
Hugo Liu and Push Singh · 2004
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives · 2007
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Mateusz Malinowski and Mario Fritz · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Acquiring comparative commonsense knowledge from the web
Niket Tandon, Gerard De Melo, and Gerhard Weikum · 2014
Earlier work this paper cites.
Abc-cnn: An attention based convolutional neural network for visual question answering
Kan Chen, Jiang Wang, Liang-Chieh Chen, Haoyuan Gao, Wei Xu, and Ram Nevatia · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
Mateusz Malinowski, Marcus Rohrbach, and Mario Fritz · 2015
Earlier work this paper cites.
Image question answering: A visual semantic embedding model and a new dataset
Mengye Ren, Ryan Kiros, and Richard Zemel · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Learning to answer questions from image using convolutional neural network
Lin Ma, Zhengdong Lu, and Hang Li · 2016
Cited alongside, same era.
Key-value memory networks for directly reading documents
Alexander Miller, Adam Fisch, Jesse Dodge, Amir-Hossein Karimi, Antoine Bordes, and Jason Weston · 2016
Cited alongside, same era.
Dynamic memory networks for visual and textual question answering
Caiming Xiong, Stephen Merity, and Richard Socher · 2016
Cited alongside, same era.
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Cited alongside, same era.
Tips and tricks for visual question answering: Learnings from the 2017 challenge
Damien Teney, Peter Anderson, Xiaodong He, and Anton Van Den Hengel · 2018
Later among the works it cites.
Scene text visual question answering
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marçal Rusinol, Ernest Valveny, CV Jawahar, and Dimosthenis Karatzas · 2019
Later among the works it cites.
Bidirectional attentive memory networks for question answering over knowledge bases
Yu Chen, Lingfei Wu, and Mohammed J Zaki · 2019
Later among the works it cites.
Language-conditioned graph networks for relational reasoning
Ronghang Hu, Anna Rohrbach, Trevor Darrell, and Kate Saenko · 2019
Later among the works it cites.
Multi-grained attention with object-level grounding for visual question answering
Pingping Huang, Jianhui Huang, Yuqing Guo, Min Qiao, and Yong Zhu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Incorporating external knowledge to answer open-domain visual questions with dynamic memory networks
Guohao Li, Hang Su, and Wenwu Zhu · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Fvqa: Fact-based visual question answering
Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel · 2017
Cited alongside, same era.
Explicit knowledge-based reasoning for visual question answering
Peng Wang, Qi Wu, Chunhua Shen, Anton van den Hengel, and Anthony Dick · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
R-vqa: learning visual relation facts with semantic attention for visual question answering
Pan Lu, Lei Ji, Wei Zhang, Nan Duan, Ming Zhou, and Jianyong Wang · 2018
Cited alongside, same era.
Visual question answering with memory-augmented networks
Chao Ma, Chunhua Shen, Anthony Dick, Qi Wu, Peng Wang, Anton van den Hengel, and Ian Reid · 2018
Cited alongside, same era.
Visual question answering as reading comprehension
Hui Li, Peng Wang, Chunhua Shen, and Anton van den Hengel · 2019
Later among the works it cites.
Social-iq: A question answering benchmark for artificial social intelligence
Amir Zadeh, Michael Chan, Paul Pu Liang, Edmund Tong, and Louis-Philippe Morency · 2019
Later among the works it cites.
Counterfactual samples synthesizing for robust visual question answering
Long Chen, Xin Yan, Jun Xiao, Hanwang Zhang, Shiliang Pu, and Yueting Zhuang · 2020
Later among the works it cites.
Multi-modal graph neural network for joint reasoning on vision and scene text
Difei Gao, Ke Li, Ruiping Wang, Shiguang Shan, and Xilin Chen · 2020
Later among the works it cites.
In defense of grid features for visual question answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik Learned-Miller, and Xinlei Chen · 2020
Later among the works it cites.
On the general value of evidence, and bilingual scene-text visual question answering
Xinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng, Canjie Luo, Lianwen Jin, Chee Seng Chan, Anton van den Hengel, and Liangwei Wang · 2020
Later among the works it cites.
Memory augmented deep recurrent neural network for video question answering
Chengxiang Yin, Jian Tang, Zhiyuan Xu, and Yanzhi Wang · 2020
Later among the works it cites.
Mucko: Multi-layer cross-modal knowledge reasoning for fact-based visualquestion answering
Zihao Zhu, Jing Yu, Yujing Wang, Yajing Sun, Yue Hu, and Qi Wu · 2020
Later among the works it cites.
Hierarchical graph attention network for few-shot visual-semantic learning
Chengxiang Yin, Kun Wu, Zhengping Che, Bo Jiang, Zhiyuan Xu, and Jian Tang · 2021
Later among the works it cites.