Fetching the paper…
Reading the bibliography…
Vision-and-language navigation (VLN) simulates a visual agent that follows natural-language navigation instructions in real-world scenes.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
The development of embodied cognition: Six lessons from babies
L. Smith and M. Gasser · 2005
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Earlier work this paper cites.
Matterport3d: Learning from rgb-d data in indoor environments
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang · 2017
Earlier work this paper cites.
Cognitive mapping and planning for visual navigation
S. Gupta, J. Davidson, S. Levine, R. Sukthankar, and J. Malik · 2017
Earlier work this paper cites.
On evaluation of embodied navigation agents
P. Anderson, A. Chang, D. S. Chaplot, A. Dosovitskiy, S. Gupta, V. Koltun, J. Kosecka, J. Malik, R. Mottaghi, M. Savva, et al · 2018
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel · 2018
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell · 2018
Earlier work this paper cites.
Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation
X. Wang, W. Xiong, H. Wang, and W. Y. Wang · 2018
Earlier work this paper cites.
Perceive, transform, and act: Multi-modal attention networks for vision-and-language navigation
F. L. L. B. M. Cornia and M. C. R. Cucchiara · 2019
Earlier work this paper cites.
Stay on the path: Instruction fidelity in vision-and-language navigation
V. Jain, G. Magalhaes, A. Ku, A. Vaswani, E. Ie, and J. Baldridge · 2019
Earlier work this paper cites.
Tactical rewind: Self-correction via backtracking in vision-and-language navigation
L. Ke, X. Li, Y. Bisk, A. Holtzman, Z. Gan, J. Liu, J. Gao, Y. Choi, and S. Srinivasa · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. D. M.-W. C. Kenton and L. K. Toutanova · 2019
Earlier work this paper cites.
Robust navigation with language pretraining and stochastic sampling
X. Li, C. Li, Q. Xia, Y. Bisk, A. Celikyilmaz, J. Gao, N. Smith, and Y. Choi · 2019
Earlier work this paper cites.
Self-monitoring navigation agent via auxiliary progress estimation
C.-Y. Ma, J. Lu, Z. Wu, G. AlRegib, Z. Kira, R. Socher, and C. Xiong · 2019
Earlier work this paper cites.
The regretful agent: Heuristic-aided navigation through progress estimation
C.-Y. Ma, Z. Wu, G. AlRegib, C. Xiong, and Z. Kira · 2019
Earlier work this paper cites.
Provably powerful graph networks
H. Maron, H. Ben-Hamu, H. Serviansky, and Y. Lipman · 2019
Earlier work this paper cites.
Learning to navigate unseen environments: Back translation with environmental dropout
H. Tan, L. Yu, and M. Bansal · 2019
Earlier work this paper cites.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y.-F. Wang, W. Y. Wang, and L. Zhang · 2019
Earlier work this paper cites.
Neural topological slam for visual navigation
D. S. Chaplot, R. Salakhutdinov, A. Gupta, and S. Gupta · 2020
Cited alongside, same era.
Evolving graphical planner: Contextual global planning for vision-and-language navigation
Z. Deng, K. Narasimhan, and O. Russakovsky · 2020
Cited alongside, same era.
Benchmarking graph neural networks
V. P. Dwivedi, C. K. Joshi, T. Laurent, Y. Bengio, and X. Bresson · 2020
Cited alongside, same era.
Towards learning a generic agent for vision-and-language navigation via pre-training
W. Hao, C. Li, X. Li, L. Carin, and J. Gao · 2020
Cited alongside, same era.
Language and visual entity relationship graph for agent navigation
Y. Hong, C. Rodriguez, Y. Qi, Q. Wu, and S. Gould · 2020
Cited alongside, same era.
Structured scene memory for vision-language navigation
H. Wang, W. Wang, W. Liang, C. Xiong, and J. Shen · 2021
Later among the works it cites.
Self-motivated communication agent for real-world vision-dialog navigation
Y. Zhu, Y. Weng, F. Zhu, X. Liang, Q. Ye, Y. Lu, and J. Jiao · 2021
Later among the works it cites.
Reinforced structured state-evolution for vision-language navigation
J. Chen, C. Gao, E. Meng, Q. Zhang, and S. Liu · 2022
Later among the works it cites.
Think global, act local: Dual-scale graph transformer for vision-and-language navigation
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev · 2022
Later among the works it cites.
Episodic memory question answering
S. Datta, S. Dharur, V. Cartillier, R. Desai, M. Khanna, D. Batra, and D. Parikh · 2022
Later among the works it cites.
Cross-modal map learning for vision and language navigation
G. Georgakis, K. Schmeckpeper, K. Wanchoo, S. Dan, E. Miltsakaki, D. Roth, and K. Daniilidis · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge · 2020
Cited alongside, same era.
Room-Across-Room: Multilingual vision-and-language navigation with dense spatiotemporal grounding
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge · 2020
Cited alongside, same era.
Improving vision-and-language navigation with image-text pairs from the web
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra · 2020
Cited alongside, same era.
Object-and-action aware model for visual language navigation
Y. Qi, Z. Pan, S. Zhang, A. v. d. Hengel, and Q. Wu · 2020
Cited alongside, same era.
Reverie: Remote embodied visual referring expression in real indoor environments
Y. Qi, Q. Wu, P. Anderson, X. Wang, W. Y. Wang, C. Shen, and A. v. d. Hengel · 2020
Cited alongside, same era.
Vision-and-dialog navigation
J. Thomason, M. Murray, M. Cakmak, and L. Zettlemoyer · 2020
Cited alongside, same era.
Active visual information gathering for vision-language navigation
H. Wang, W. Wang, T. Shu, W. Liang, and J. Shen · 2020
Cited alongside, same era.
Later among the works it cites.
Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation
Y. Hong, Z. Wang, Q. Wu, and S. Gould · 2022
Later among the works it cites.
Envedit: Environment editing for vision-and-language navigation
J. Li, H. Tan, and M. Bansal · 2022
Later among the works it cites.
Layout-aware dreamer for embodied referring expression grounding
M. Li, Z. Wang, T. Tuytelaars, and M.-F. Moens · 2022
Later among the works it cites.
Adapt: Vision-language navigation with modality-aligned action prompts
B. Lin, Y. Zhu, Z. Chen, X. Liang, J. Liu, and X. Liang · 2022
Later among the works it cites.
Multimodal transformer with variable-length memory for vision-and-language navigation
C. Lin, Y. Jiang, J. Cai, L. Qu, G. Haffari, and Z. Yuan · 2022
Later among the works it cites.
Hop: History-and-order aware pre-training for vision-and-language navigation
Y. Qiao, Y. Qi, Y. Hong, Z. Yu, P. Wang, and Q. Wu · 2022
Later among the works it cites.
Rapid exploration for open-world navigation with latent goal models
D. Shah, B. Eysenbach, N. Rhinehart, and S. Levine · 2022
Later among the works it cites.
Visitron: Visual semantics-aligned interactively trained object-navigator
A. Shrivastava, K. Gopalakrishnan, Y. Liu, R. Piramuthu, G. Tür, D. Parikh, and D. Hakkani-Tur · 2022
Later among the works it cites.
Cross-modal semantic alignment pre-training for vision-and-language navigation
S. Wu, X. Fu, F. Wu, and Z.-J. Zha · 2022
Later among the works it cites.
Target-driven structured transformer planner for vision-language navigation
Y. Zhao, J. Chen, C. Gao, W. Wang, L. Yang, H. Ren, H. Xia, and S. Liu · 2022
Later among the works it cites.
Etpnav: Evolving topological planning for vision-language navigation in continuous environments
D. An, H. Wang, W. Wang, Z. Wang, Y. Huang, K. He, and L. Wang · 2023
Closest in time.
Iterative vision-and-language navigation
J. Krantz, S. Banerjee, W. Zhu, J. Corso, P. Anderson, S. Lee, and J. Thomason · 2023
Closest in time.
Improving vision-and-language navigation by generating future-view image semantics
J. Li and M. Bansal · 2023
Closest in time.
Kerm: Knowledge enhanced reasoning for vision-and-language navigation
X. Li, Z. Wang, J. Yang, Y. Wang, and S. Jiang · 2023
Closest in time.
Gridmm: Grid memory map for vision-and-language navigation
Z. Wang, X. Li, J. Yang, Y. Liu, and S. Jiang · 2023
Closest in time.