Fetching the paper…
Reading the bibliography…
Vision-and-Language Navigation (VLN) is a challenging task in which an agent needs to follow a language-specified path to reach a target destination.
Effective and general evaluation for instruction conditioned navigation using dynamic time warping
Magalhaes, G., Jain, V., Ku, A., Ie, E., Baldridge, J., 2019 · 1907
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., Salakhutdinov, R., 2014 · 1958
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R.J., 1992 · 1992
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., Frasconi, P., et al., 1994 · 1994
Earlier work this paper cites.
Using dynamic time warping to find patterns in time series, in: Proceedings of the International Conference on Knowledge Discovery and Data Mining
Berndt, D.J., Clifford, J., 1994 · 1994
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L., 2009 · 2009
Earlier work this paper cites.
The Ecological Approach to Visual Perception: Classic Edition
Gibson, J.J., 2014 · 2014
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation, in: Proceedings of the Conference on Empirical Methods in Natural Language Processing
Pennington, J., Socher, R., Manning, C.D., 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering, in: Proceedings of the International Conference on Computer Vision
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D., 2015 · 2015
Earlier work this paper cites.
Adam: a method for stochastic optimization, in: Proceedings of the International Conference on Learning Representations
Kingma, D., Ba, J., 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Vinyals, O., Toshev, A., Bengio, S., Erhan, D., 2015 · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention, in: Proceedings of the International Conference on Machine Learning
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., Bengio, Y., 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
He, K., Zhang, X., Ren, S., Sun, J., 2016 · 2016
Earlier work this paper cites.
Matterport3D: Learning from RGB-D Data in Indoor Environments, in: Proceedings of the International Conference on 3D Vision
Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., Song, S., Zeng, A., Zhang, Y., 2017 · 2017
Cited alongside, same era.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D., 2017 · 2017
Cited alongside, same era.
Cognitive mapping and planning for visual navigation, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Gupta, S., Davidson, J., Levine, S., Sukthankar, R., Malik, J., 2017 · 2017
Cited alongside, same era.
Attention is all you need, in: Advances in Neural Information Processing Systems
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I., 2017 · 2017
Cited alongside, same era.
Embodied Vision-and-Language Navigation with Dynamic Convolutional Filters, in: Proceedings of the British Machine Vision Conference
Landi, F., Baraldi, L., Corsini, M., Cucchiara, R., 2019 · 2019
Closest in time.
Robust Navigation with Language Pretraining and Stochastic Sampling, in: Proceedings of the Conference on Empirical Methods in Natural Language Processing
Li, X., Li, C., Xia, Q., Bisk, Y., Celikyilmaz, A., Gao, J., Smith, N., Choi, Y., 2019 · 2019
Closest in time.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks, in: Advances in Neural Information Processing Systems
Lu, J., Batra, D., Parikh, D., Lee, S., 2019 · 2019
Closest in time.
Habitat: A Platform for Embodied AI Research, in: Proceedings of the International Conference on Computer Vision
Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Parikh, D., Batra, D., 2019 · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fried, D., Hu, R., Cirik, V., Rohrbach, A., Andreas, J., Morency, L.P., Berg-Kirkpatrick, T., Saenko, K., Klein, D., Darrell, T., 2018 · 2018
Cited alongside, same era.
Shifting the Baseline: Single Modality Performance on Visual Navigation & QA, in: Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics
Thomason, J., Gordon, D., Bisk, Y., 2018 · 2018
Cited alongside, same era.
Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation, in: Proceedings of the European Conference on Computer Vision
Wang, X., Xiong, W., Wang, H., Yang Wang, W., 2018 · 2018
Cited alongside, same era.
Gibson env: Real-world perception for embodied agents, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Xia, F., Zamir, A.R., He, Z., Sax, A., Malik, J., Savarese, S., 2018 · 2018
Cited alongside, same era.
Touchdown: Natural language navigation and spatial reasoning in visual street environments, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Chen, H., Suhr, A., Misra, D., Snavely, N., Artzi, Y., 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, in: Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2019 · 2019
Cited alongside, same era.
From Language to Goals: Inverse Reinforcement Learning for Vision-Based Instruction Following, in: Proceedings of the International Conference on Learning Representations
Fu, J., Korattikara, A., Levine, S., Guadarrama, S., 2019 · 2019
Cited alongside, same era.
Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation, in: Proceedings of Annual Meeting of the Association for Computational Linguistics
Jain, V., Magalhaes, G., Ku, A., Vaswani, A., Ie, E., Baldridge, J., 2019 · 2019
Cited alongside, same era.
Shen, W.B., Xu, D., Zhu, Y., Guibas, L.J., Fei-Fei, L., Savarese, S., 2019 · 2019
Closest in time.
Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout, in: Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics
Tan, H., Yu, L., Bansal, M., 2019 · 2019
Closest in time.
Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Wang, X., Huang, Q., Celikyilmaz, A., Gao, J., Shen, D., Wang, Y.F., Wang, W.Y., Zhang, L., 2019 · 2019
Closest in time.
Embodied Amodal Recognition: Learning to Move to Perceive Objects, in: Proceedings of the International Conference on Computer Vision
Yang, J., Ren, Z., Xu, M., Chen, X., Crandall, D., Parikh, D., Batra, D., 2019 · 2019
Closest in time.
Evolving graphical planner: Contextual global planning for vision-and-language navigation, in: Advances in Neural Information Processing Systems
Deng, Z., Narasimhan, K., Russakovsky, O., 2020 · 2020
Closest in time.
Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-training, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Hao, W., Li, C., Li, X., Carin, L., Gao, J., 2020 · 2020
Closest in time.
Language and visual entity relationship graph for agent navigation, in: Advances in Neural Information Processing Systems
Hong, Y., Rodriguez-Opazo, C., Qi, Y., Wu, Q., Gould, S., 2020 · 2020
Closest in time.
Language-guided Navigation via Cross-Modal Grounding and Alternate Adversarial Learning
Zhang, W., Ma, C., Wu, Q., Yang, X., 2020 · 2020
Closest in time.
Vision-language navigation with self-supervised auxiliary reasoning tasks, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Zhu, F., Zhu, Y., Chang, X., Liang, X., 2020 · 2020
Closest in time.