Fetching the paper…
Reading the bibliography…
We explore blindfold (question-only) baselines for Embodied Question Answering.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Object perception and object naming in early development
B. Landau, L. Smith, and S. Jones · 1998
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin · 2003
Earlier work this paper cites.
Fostering mathematical thinking through playful learning
K. Fisher, K. Hirsh-Pasek, and R. M. Golinkoff · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Exploring nearest neighbor approaches for image captioning
J. Devlin, S. Gupta, R. Girshick, M. Mitchell, and C. L. Zitnick · 2015
Cited alongside, same era.
Video and image based emotion recognition challenges in the wild: Emotiw 2015
A. Dhall, O. Ramana Murthy, R. Goecke, J. Joshi, and T. Gedeon · 2015
Cited alongside, same era.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Cited alongside, same era.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Cited alongside, same era.
Bag of tricks for efficient text classification
A. Joulin, E. Grave, P. Bojanowski, and T. Mikolov · 2016
Cited alongside, same era.
Ai2-thor: An interactive 3d environment for visual ai
E. Kolve, R. Mottaghi, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi · 2017
Later among the works it cites.
Story cloze task: Uw nlp system
R. Schwartz, M. Sap, I. Konstas, L. Zilles, Y. Choi, and N. A. Smith · 2017
Later among the works it cites.
Home: A household multimodal environment
S. Brodeur, E. Perez, A. Anand, F. Golemo, L. Celotti, F. Strub, J. Rouat, H. Larochelle, and A. Courville · 2018
Closest in time.
Neural modular control for embodied question answering
A. Das, G. Gkioxari, S. Lee, D. Parikh, and D. Batra · 2018
Closest in time.
Annotation artifacts in natural language inference data
S. Gururangan, S. Swayamdipta, O. Levy, R. Schwartz, S. R. Bowman, and N. A. Smith · 2018
Closest in time.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Cited alongside, same era.
Joint embeddings of scene graphs and images
E. Belilovsky, M. Blaschko, J. R. Kiros, R. Urtasun, and R. Zemel · 2017
Cited alongside, same era.
Embodied question answering
A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra
Cited in the paper.
D. Kaushik and Z. C. Lipton · 2018
Closest in time.
Hypothesis only baselines in natural language inference
A. Poliak, J. Naradowsky, A. Haldar, R. Rudinger, and B. Van Durme · 2018
Closest in time.
Building generalizable agents with a realistic and rich 3d environment
Y. Wu, Y. Wu, G. Gkioxari, and Y. Tian · 2018
Closest in time.