Fetching the paper…
Reading the bibliography…
Visual Question Answering (VQA) has become one of the key benchmarks of visual recognition progress.
Intentions and intentionality: Foundations of social cognition
Bertram F Malle, Louis J Moses, and Dare A Baldwin · 2001
Earlier work this paper cites.
Object recognition via recognition of finger pointing actions
Michael Hild, Motonobu Hashimoto, and Kazunobu Yoshida · 2003
Earlier work this paper cites.
Cognitive and language development in children
John Oates and Andrew Grayson · 2004
Earlier work this paper cites.
Augmenting looking, pointing and reaching gestures to enhance the searching and browsing of physical objects
David Merrill and Pattie Maes · 2007
Earlier work this paper cites.
Human pointing as a robot directive
Syed Shaukat Raza Abidi, MaryAnn Williams, and Benjamin Johnston · 2013
Earlier work this paper cites.
Training object class detectors from eye tracking data
Dim P. Papadopoulos, Alasdair D. F. Clarke, Frank Keller, and Vittorio Ferrari · 2014
Earlier work this paper cites.
Robot deictics: how gesture and context shape referential communication
Allison Sauppe and Bilge Mutlu · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Material recognition in the wild with the materials in context database
Sean Bell, Paul Upchurch, Noah Snavely, and Kavita Bala · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and J. Sun · 2015
Earlier work this paper cites.
What’s the point: Semantic segmentation with point supervision
Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Click carving: Segmenting objects in video with point clicks
Suyog Dutt Jain and Kristen Grauman · 2016
Cited alongside, same era.
Visual7w: Grounded Question Answering in Images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, Jose M. F. Moura, Devi Parikh, and Dhruv Batra · 2017
Cited alongside, same era.
VQS: Linking Segmentations to Questions and Answers for Supervised Attention in VQA and Question-Focused Semantic Segmentation
Chuang Gan, Yandong Li, Haoxiang Li, Chen Sun, and Boqing Gong · 2017
Cited alongside, same era.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2017
Iqa: Visual question answering in interactive environments
Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi · 2018
Later among the works it cites.
Multimodal Explanations: Justifying Decisions and Pointing to the Evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach · 2018
Later among the works it cites.
Pythia v0.1: the winning entry to the vqa challenge 2018
Yu Jiang*, Vivek Natarajan*, Xinlei Chen*, Marcus Rohrbach, Dhruv Batra, and Devi Parikh · 2018
Later among the works it cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Aishwarya Agrawal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2019
Later among the works it cites.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Drew A. Hudson and Christopher D. Manning · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie · 2017
Cited alongside, same era.
Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi · 2018
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Embodied Question Answering
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra · 2018
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Later among the works it cites.
Deep modular co-attention networks for visual question answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian · 2019
Later among the works it cites.
Interpretable Visual Question Answering by Visual Grounding From Attention Supervision Mining
Yundong Zhang, Juan Carlos Niebles, and Alvaro Soto · 2019
Later among the works it cites.
In defense of grid features for visual question answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik Learned-Miller, and Xinlei Chen · 2020
Closest in time.