Fetching the paper…
Reading the bibliography…
Visual Question Answering is a multi-modal task that aims to measure high-level visual understanding.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments
Alon Lavie and Abhaya Agarwal · 2007
Earlier work this paper cites.
Bootstrapping dialog systems with word embeddings
Gabriel Forgues, Joelle Pineau, Jean-Marie Larchevêque, and Réal Tremblay · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Generating visual explanations
Lisa Anne Hendricks, Zeynep Akata, Marcus Rohrbach, Jeff Donahue, Bernt Schiele, and Trevor Darrell · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Self-critical sequence training for image captioning
Steven J. Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel · 2016
Earlier work this paper cites.
Image captioning with semantic attention
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo · 2016
Earlier work this paper cites.
Visual7W: Grounded Question Answering in Images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes · 2017
Earlier work this paper cites.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, Jose M. F. Moura, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Learning to reason: End-to-end module networks for visual question answering
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
Inferring and executing programs for visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Judy Hoffman, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
Hadamard Product for Low-rank Bilinear Pooling
Jin-Hwa Kim, Kyoung Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang · 2017
Cited alongside, same era.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2018
Later among the works it cites.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Joshua B. Tenenbaum · 2018
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2018
Later among the works it cites.
Scene graph contextualization in visual commonsense reasoning
Florin Brad · 2019
Later among the works it cites.
Murel: Multimodal relational reasoning for visual question answering
Remi Cadene, Hedi Ben-younes, Matthieu Cord, and Nicolas Thome · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher · 2017
Cited alongside, same era.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David G. T. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter W. Battaglia, and Timothy P. Lillicrap · 2017
Cited alongside, same era.
Multi-level attention networks for visual question answering
Dongfei Yu, Jianlong Fu, Tao Mei, and Yong Rui · 2017
Cited alongside, same era.
Multi-modal factorized bilinear pooling with co-attention learning for visual question answering
Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Daniel Matthew Cer, Yinfei Yang, Sheng yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, C. Tar, Yun-Hsuan Sung, B. Strope, and R. Kurzweil · 2018
Cited alongside, same era.
Bilinear Attention Networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang · 2018
Cited alongside, same era.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2019
Later among the works it cites.
Relation-aware graph attention network for visual question answering
Linjie Li, Zhe Gan, Yu Cheng, and Jingjing Liu · 2019
Later among the works it cites.
Tab-vcr: Tags and attributes based vcr baselines
Jingxiang Lin, Unnat Jain, and Alexander G. Schwing · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, D. Parikh, and Stefan Lee · 2019
Later among the works it cites.
Generating question relevant captions to aid visual question answering
Jialin Wu, Zeyuan Hu, and Raymond J. Mooney · 2019
Later among the works it cites.
Reasoning visual dialogs with structural and partial observations
Zilong Zheng, Wenguan Wang, Siyuan Qi, and Song-Chun Zhu · 2019
Later among the works it cites.
Robust explanations for visual question answering
Badri N. Patro, Shivansh Pate, and Vinay P. Namboodiri · 2020
Closest in time.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi · 2020
Closest in time.
Meta module network for compositional visual reasoning
Wenhu Chen, Zhe Gan, Linjie Li, Yu Cheng, William Wang, and Jingjing Liu · 2021
Closest in time.