Fetching the paper…
Reading the bibliography…
We study the problem of jointly reasoning about language and vision through a navigation and spatial reasoning task.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Walk the Talk: Connecting Language, Knowledge, Action in Route Instructions
M. MacMahon, B. Stankiewics, and B. Kuipers · 2006
Earlier work this paper cites.
Learning to interpret natural language navigation instructions from observations
D. L. Chen and R. J. Mooney · 2011
Earlier work this paper cites.
Hogwild: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
A Joint Model of Language and Perception for Grounded Attribute Learning
C. Matuszek, N. FitzGerald, L. S. Zettlemoyer, L. Bo, and D. Fox · 2012
Earlier work this paper cites.
Bringing Semantics into Focus Using Visual Abstraction
C. L. Zitnick and D. Parikh · 2013
Earlier work this paper cites.
ReferItGame: Referring to Objects in Photographs of Natural Scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
T.-Y. Lin, M. Maire, S. J. Belongie, L. D. Bourdev, R. B. Girshick, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
VQA: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Microsoft COCO Captions: Data Collection and Evaluation Server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Cited alongside, same era.
Towards a Dataset for Human Computer Communication via Grounded Language Acquisition
Y. Bisk, D. Marcu, and W. Wong · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Generation and Comprehension of Unambiguous Object Descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
Cited alongside, same era.
Learning Deep Representations of Fine-Grained Visual Descriptions
S. Reed, Z. Akata, H. Lee, and B. Schiele · 2016
Cited alongside, same era.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. D. Reid, S. Gould, and A. van den Hengel · 2018
Closest in time.
Mapping Navigation Instructions to Continuous Control Actions with Position Visitation Prediction
V. Blukis, D. Misra, R. A. Knepper, and Y. Artzi · 2018
Closest in time.
Gated-Attention Architectures for Task-Oriented Language Grounding
D. S. Chaplot, K. M. Sathyendra, R. K. Pasumarthi, D. Rajagopal, and R. Salakhutdinov · 2018
Closest in time.
Embodied Question Answering
A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra · 2018
Closest in time.
Talk the Walk: Navigating New York City through Grounded Dialogue
H. de Vries, K. Shuster, D. Batra, D. Parikh, J. Weston, and D. Kiela · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Cited alongside, same era.
Where is Misty? Interpreting Spatial Descriptors by Modeling Regions in Space
N. Kitaev and D. Klein · 2017
Cited alongside, same era.
AI2-THOR: An Interactive 3D Environment for Visual AI
E. Kolve, R. Mottaghi, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi · 2017
Cited alongside, same era.
Mapping Instructions and Visual Observations to Actions with Reinforcement Learning
D. Misra, J. Langford, and Y. Artzi · 2017
Cited alongside, same era.
A Corpus of Natural Language for Visual Reasoning
A. Suhr, M. Lewis, J. Yeh, and Y. Artzi · 2017
Cited alongside, same era.
On evaluation of embodied navigation agents
P. Anderson, A. X. Chang, D. S. Chaplot, A. Dosovitskiy, S. Gupta, V. Koltun, J. Kosecka, J. Malik, R. Mottaghi, M. Savva, and A. R. Zamir · 2018
Cited alongside, same era.
Weakly supervised semantic parsing with abstract examples
O. Goldman, V. Latcinnik, E. Nave, A. Globerson, and J. Berant · 2018
Closest in time.
IQA: Visual Question Answering in Interactive Environments
D. Gordon, A. Kembhavi, M. Rastegari, J. Redmon, D. Fox, and A. Farhadi · 2018
Closest in time.
Learning to Navigate in Cities without a Map
P. Mirowski, M. K. Grimes, M. Malinowski, K. M. Hermann, K. Anderson, D. Teplyashin, K. Simonyan, K. Kavukcuoglu, A. Zisserman, and R. Hadsell · 2018
Closest in time.
Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction
D. Misra, A. Bennett, V. Blukis, E. Niklasson, M. Shatkhin, and Y. Artzi · 2018
Closest in time.
A corpus for reasoning about natural language grounded in photographs
A. Suhr, S. Zhou, I. F. Zhang, H. Bai, and Y. Artzi · 2018
Closest in time.
CHALET: Cornell House Agent Learning Environment
C. Yan, D. Misra, A. Bennnett, A. Walsman, Y. Bisk, and Y. Artzi · 2018
Closest in time.