Fetching the paper…
Reading the bibliography…
Vision-language navigation is a task that requires an agent to follow instructions to navigate in environments.
S. B. Udin and J. W. Fawcett, “Formation of topographic maps,” Annual review of neuroscience , vol. 11, no. 1, pp. 289–327, 1988
1988
Earlier work this paper cites.
A. Elfes, “Using occupancy grids for mobile robot perception and navigation,” Computer , vol. 22, no. 6, pp. 46–57, 1989
1989
Earlier work this paper cites.
B. Kuipers and Y.-T. Byun, “A robot exploration and mapping strategy based on a semantic hierarchy of spatial representations,” Robotics and autonomous systems , vol. 8, no. 1-2, pp. 47–63, 1991
1991
Earlier work this paper cites.
J. A. Sethian, “A fast marching level set method for monotonically advancing fronts.” proceedings of the National Academy of Sciences , vol. 93, no. 4, pp. 1591–1595, 1996
1996
Earlier work this paper cites.
F. Dellaert, D. Fox, W. Burgard, and S. Thrun, “Monte carlo localization for mobile robots,” in international Conference on Robotics and Automation , vol. 2. IEEE, 1999, pp. 1322–1328
1999
Earlier work this paper cites.
H. Choset and K. Nagatani, “Topological simultaneous localization and mapping (slam): toward exact localization without explicit localization,” IEEE Transactions on robotics and automation , vol. 17, no. 2, pp. 125–137, 2001
2001
Earlier work this paper cites.
R. A. Newcombe, S. J. Lovegrove, and A. J. Davison, “Dtam: Dense tracking and mapping in real-time,” in Proceedings of the IEEE/CVF International Conference on Computer Vision . IEEE, 2011, pp. 2320–2327
2011
Earlier work this paper cites.
S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in Proceedings of the fourteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 2011, pp. 627–635
2011
Earlier work this paper cites.
R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos, “Orb-slam: a versatile and accurate monocular slam system,” IEEE Transactions on Robotics , vol. 31, no. 5, pp. 1147–1163, 2015
2015
Earlier work this paper cites.
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” Advances in Neural Information Processing Systems , vol. 28, pp. 1171–1179, 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, pp. 5998–6008, 2017
2017
Earlier work this paper cites.
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niebner, M. Savva, S. Song, A. Zeng, and Y. Zhang, “Matterport3d: Learning from rgb-d data in indoor environments,” in International Conference on 3D Vision . IEEE, 2017, pp. 667–676
2017
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2018, pp. 3674–3683
2018
Earlier work this paper cites.
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell, “Speaker-follower models for vision-and-language navigation,” Advances in Neural Information Processing Systems , vol. 31, pp. 3318–3329, 2018
2018
Earlier work this paper cites.
X. Wang, W. Xiong, H. Wang, and W. Y. Wang, “Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation,” in Proceedings of the European Conference on Computer Vision , 2018, pp. 37–53
2018
Earlier work this paper cites.
J. F. Henriques and A. Vedaldi, “Mapnet: An allocentric spatial memory for mapping environments,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 8476–8484
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y.-F. Wang, W. Y. Wang, and L. Zhang, “Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 6629–6638
2019
Earlier work this paper cites.
E. Wijmans, A. Kadian, A. Morcos, S. Lee, I. Essa, D. Parikh, M. Savva, and D. Batra, “Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
H. Chen, A. Suhr, D. Misra, N. Snavely, and Y. Artzi, “Touchdown: Natural language navigation and spatial reasoning in visual street environments,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 12 538–12 547
2019
Earlier work this paper cites.
K. Nguyen and H. Daumé III, “Help, anna! visual navigation with natural multimodal assistance via retrospective curiosity-encouraging imitation learning,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing , 2019, pp. 684–695
2019
Earlier work this paper cites.
C.-Y. Ma, J. Lu, Z. Wu, G. AlRegib, Z. Kira, R. Socher, and C. Xiong, “Self-monitoring navigation agent via auxiliary progress estimation,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
H. Tan, L. Yu, and M. Bansal, “Learning to navigate unseen environments: Back translation with environmental dropout,” in Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics , 2019, pp. 2610–2621
2019
Earlier work this paper cites.
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik et al. , “Habitat: A platform for embodied ai research,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 9339–9347
2019
Earlier work this paper cites.
D. S. Chaplot, D. Gandhi, S. Gupta, A. Gupta, and R. Salakhutdinov, “Learning to explore using active neural slam,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
H. Tan and M. Bansal, “Lxmert: Learning cross-modality encoder representations from transformers,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing , 2019, pp. 5100–5111
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in Neural Information Processing Systems , vol. 32, pp. 8024–8035, 2019
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
J. Krantz, E. Wijmans, A. Majumdar, D. Batra, and S. Lee, “Beyond the nav-graph: Vision-and-language navigation in continuous environments,” in Proceedings of the European Conference on Computer Vision . Springer, 2020, pp. 104–120
2020
Cited alongside, same era.
Z. Deng, K. Narasimhan, and O. Russakovsky, “Evolving graphical planner: Contextual global planning for vision-and-language navigation,” Advances in Neural Information Processing Systems , vol. 33, pp. 20 660–20 672, 2020
2020
Cited alongside, same era.
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge, “Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing , 2020, pp. 4392–4412
2020
Cited alongside, same era.
J. Thomason, M. Murray, M. Cakmak, and L. Zettlemoyer, “Vision-and-dialog navigation,” in Conference on Robot Learning . PMLR, 2020, pp. 394–406
2020
Cited alongside, same era.
M. Hahn, D. S. Chaplot, S. Tulsiani, M. Mukadam, J. M. Rehg, and A. Gupta, “No rl, no simulation: Learning to navigate without navigating,” Advances in Neural Information Processing Systems , vol. 34, pp. 26 661–26 673, 2021
2021
Later among the works it cites.
X. Zhao, H. Agrawal, D. Batra, and A. G. Schwing, “The surprising effectiveness of visual odometry techniques for embodied pointgoal navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 16 127–16 136
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 8748–8763
2021
Later among the works it cites.
Y. Hong, Z. Wang, Q. Wu, and S. Gould, “Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 439–15 449
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Qi, Q. Wu, P. Anderson, X. Wang, W. Y. Wang, C. Shen, and A. v. d. Hengel, “Reverie: Remote embodied visual referring expression in real indoor environments,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 9982–9991
2020
Cited alongside, same era.
Y. Qi, Z. Pan, S. Zhang, A. v. d. Hengel, and Q. Wu, “Object-and-action aware model for visual language navigation,” in Proceedings of the European Conference on Computer Vision . Springer, 2020, pp. 303–317
2020
Cited alongside, same era.
Y. Hong, C. Rodriguez, Y. Qi, Q. Wu, and S. Gould, “Language and visual entity relationship graph for agent navigation,” Advances in Neural Information Processing Systems , vol. 33, pp. 7685–7696, 2020
2020
Cited alongside, same era.
H. Wang, Q. Wu, and C. Shen, “Soft expert reward learning for vision-and-language navigation,” in Proceedings of the European Conference on Computer Vision . Springer, 2020, pp. 126–141
2020
Cited alongside, same era.
H. Wang, W. Wang, T. Shu, W. Liang, and J. Shen, “Active visual information gathering for vision-language navigation,” in Proceedings of the European Conference on Computer Vision . Springer, 2020, pp. 307–322
2020
Cited alongside, same era.
A. Parvaneh, E. Abbasnejad, D. Teney, J. Q. Shi, and A. van den Hengel, “Counterfactual vision-and-language navigation: Unravelling the unseen,” Advances in Neural Information Processing Systems , vol. 33, pp. 5296–5307, 2020
2020
Cited alongside, same era.
W. Hao, C. Li, X. Li, L. Carin, and J. Gao, “Towards learning a generic agent for vision-and-language navigation via pre-training,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 13 137–13 146
2020
Cited alongside, same era.
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra, “Improving vision-and-language navigation with image-text pairs from the web,” in Proceedings of the European Conference on Computer Vision . Springer, 2020, pp. 259–274
2020
Cited alongside, same era.
2022
Later among the works it cites.
J. Krantz and S. Lee, “Sim-2-sim transfer for vision-and-language navigation in continuous environments,” in Proceedings of the European Conference on Computer Vision . Springer, 2022, pp. 588–603
2022
Later among the works it cites.
2022
Later among the works it cites.
G. Georgakis, K. Schmeckpeper, K. Wanchoo, S. Dan, E. Miltsakaki, D. Roth, and K. Daniilidis, “Cross-modal map learning for vision and language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 460–15 470
2022
Later among the works it cites.
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev, “Think global, act local: Dual-scale graph transformer for vision-and-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 537–16 547
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Li, H. Tan, and M. Bansal, “Envedit: Environment editing for vision-and-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 407–15 417
2022
Later among the works it cites.
H. Wang, W. Liang, J. Shen, L. Van Gool, and W. Wang, “Counterfactual cycle-consistent learning for instruction following and generation in vision-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 471–15 481
2022
Later among the works it cites.
S. Wang, C. Montgomery, J. Orbay, V. Birodkar, A. Faust, I. Gur, N. Jaques, A. Waters, J. Baldridge, and P. Anderson, “Less is more: Generating grounded navigation instructions from landmarks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 428–15 438
2022
Later among the works it cites.
C. Lin, Y. Jiang, J. Cai, L. Qu, G. Haffari, and Z. Yuan, “Multimodal transformer with variable-length memory for vision-and-language navigation,” in Proceedings of the European Conference on Computer Vision . Springer, 2022, pp. 380–397
2022
Later among the works it cites.
Y. Zhao, J. Chen, C. Gao, W. Wang, L. Yang, H. Ren, H. Xia, and S. Liu, “Target-driven structured transformer planner for vision-language navigation,” in Proceedings of the ACM International Conference on Multimedia , 2022, pp. 4194–4203
2022
Later among the works it cites.
B. Lin, Y. Zhu, Z. Chen, X. Liang, J. Liu, and X. Liang, “Adapt: Vision-language navigation with modality-aligned action prompts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 396–15 406
2022
Later among the works it cites.
X. Liang, F. Zhu, L. Lingling, H. Xu, and X. Liang, “Visual-language navigation pretraining via prompt-based environmental self-exploration,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics , 2022, pp. 4837–4851
2022
Later among the works it cites.
M. Z. Irshad, N. C. Mithun, Z. Seymour, H.-P. Chiu, S. Samarasekera, and R. Kumar, “Sasra: Semantically-aware spatio-temporal reasoning agent for vision-and-language navigation in continuous environments,” International Conference on Pattern Recognition , pp. 4065–4071, 2022
2022
Later among the works it cites.
P. Chen, D. Ji, K. Lin, R. Zeng, T. H. Li, M. Tan, and C. Gan, “Weakly-supervised multi-granularity map learning for vision-and-language navigation,” in Advances in Neural Information Processing Systems , 2022, pp. 38 149–38 161
2022
Later among the works it cites.
H. Luo, A. Yue, Z.-W. Hong, and P. Agrawal, “Stubborn: A strong baseline for indoor object navigation,” in IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 2022, pp. 3287–3293
2022
Later among the works it cites.
P. Xu, X. Zhu, and D. A. Clifton, “Multimodal learning with transformers: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 10, pp. 12 113–12 132, 2023
2023
Closest in time.
H. Wang, W. Wang, W. Liang, S. C. Hoi, J. Shen, and L. V. Gool, “Active perception for visual-language navigation,” International Journal of Computer Vision , vol. 131, no. 3, pp. 607–625, 2023
2023
Closest in time.
A. Kamath, P. Anderson, S. Wang, J. Y. Koh, A. Ku, A. Waters, Y. Yang, J. Baldridge, and Z. Parekh, “A new path: Scaling vision-and-language navigation with synthetic instructions and imitation learning,” pp. 10 813–10 823, 2023
2023
Closest in time.
X. Wang, W. Wang, J. Shao, and Y. Yang, “Lana: A language-capable navigator for instruction dreaming and generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 19 048–19 058
2023
Closest in time.
X. Wang, W. Wang, and J. Shao, “Learning to follow and generate instructions for language-capable navigation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , pp. 1–17, 2023
2023
Closest in time.
Y. Qiao, Y. Qi, Y. Hong, Z. Yu, P. Wang, and Q. Wu, “Hop+: History-enhanced and order-aware pre-training for vision-and-language navigation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 7, pp. 8524–8537, 2023
2023
Closest in time.
M. Hwang, J. Jeong, M. Kim, Y. Oh, and S. Oh, “Meta-explore: Exploratory hierarchical vision-and-language navigation using scene object spectrum grounding,” pp. 6683–6693, 2023
2023
Closest in time.
R. Liu, X. Wang, W. Wang, and Y. Yang, “Bird’s-eye-view scene graph for vision-language navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 10 968–10 980
2023
Closest in time.
D. An, Y. Qi, Y. Li, Y. Huang, L. Wang, T. Tan, and J. Shao, “Bevbert: Multimodal map pre-training for language-guided navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2737–2748
2023
Closest in time.
H. Wang, W. Liang, L. V. Gool, and W. Wang, “Dreamwalker: Mental planning for continuous vision-language navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 10 873–10 883
2023
Closest in time.
T. Gervet, S. Chintala, D. Batra, J. Malik, and D. S. Chaplot, “Navigating to objects in the real world,” Science Robotics , vol. 8, no. 79, p. eadf6991, 2023
2023
Closest in time.
J. Chen, W. Wang, S. Liu, H. Li, and Y. Yang, “Omnidirectional information gathering for knowledge transfer-based audio-visual navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 10 993–11 003
2023
Closest in time.