Fetching the paper…
Reading the bibliography…
We introduce GQA, a new dataset for real-world visual reasoning and compositional question answering, seeking to address key shortcomings of previous VQA datasets.
An analysis of test-wiseness
J. Millman, C. H. Bishop, and R. Ebel · 1965
Earlier work this paper cites.
The earth mover’s distance as a metric for image retrieval
Y. Rubner, C. Tomasi, and L. J. Guibas · 2000
Earlier work this paper cites.
Asked and answered: Knowledge levels when we won’t take ‘don’t know’ for an answer
J. J. Mondak and B. C. Davis · 2001
Earlier work this paper cites.
Guess where: The position of correct answers in multiple-choice test items as a psychometric variable
Y. Attali and M. Bar-Hillel · 2003
Earlier work this paper cites.
Chi-square distribution
H. O. Lancaster and E. Seneta · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Unbiased look at dataset bias
A. Torralba and A. A. Efros · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Earlier work this paper cites.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
A. Agrawal, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Cited alongside, same era.
Generating natural questions about an image
N. Mostafazadeh, I. Misra, J. Devlin, M. Mitchell, X. He, and L. Vanderwende · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Creativity: Generating diverse questions using variational autoencoders
U. Jain, Z. Zhang, and A. G. Schwing · 2017
Later among the works it cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Later among the works it cites.
An analysis of visual question answering algorithms
K. Kafle and C. Kanan · 2017
Later among the works it cites.
Visual question answering: Datasets, algorithms, and future challenges
K. Kafle and C. Kanan · 2017
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Later among the works it cites.
The promise of premise: Harnessing question premises in visual question answering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Yin and yang: Balancing and answering binary visual questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Automatic generation of grounded visual questions
S. Zhang, L. Qu, S. You, Z. Yang, and J. Zhang · 2016
Cited alongside, same era.
Visual7W: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Cited alongside, same era.
VQA: Visual question answering
A. Agrawal, J. Lu, S. Antol, M. Mitchell, C. L. Zitnick, D. Parikh, and D. Batra · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and VQA
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2017
Cited alongside, same era.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
A. Das, H. Agrawal, L. Zitnick, D. Parikh, and D. Batra · 2017
Cited alongside, same era.
A. Mahendru, V. Prabhu, A. Mohapatra, D. Batra, and S. Lee · 2017
Later among the works it cites.
Tips and tricks for visual question answering: Learnings from the 2017 challenge
D. Teney, P. Anderson, X. He, and A. van den Hengel · 2017
Later among the works it cites.
Graph-structured representations for visual question answering
D. Teney, L. Liu, and A. van den Hengel · 2017
Later among the works it cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi · 2018
Later among the works it cites.
Compositional attention networks for machine reasoning
D. A. Hudson and C. D. Manning · 2018
Later among the works it cites.
VQA-E: Explaining, elaborating, and enhancing your answers for visual questions
Q. Li, Q. Tao, S. Joty, J. Cai, and J. Luo · 2018
Later among the works it cites.
Multimodal explanations: Justifying decisions and pointing to the evidence
D. H. Park, L. A. Hendricks, Z. Akata, A. Rohrbach, B. Schiele, T. Darrell, and M. Rohrbach · 2018
Later among the works it cites.
A corpus for reasoning about natural language grounded in photographs
A. Suhr, S. Zhou, I. Zhang, H. Bai, and Y. Artzi · 2018
Later among the works it cites.