Fetching the paper…
Reading the bibliography…
We address the task of Vision-Language Navigation in Continuous Environments (VLN-CE) under the zero-shot setting.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in Advances in neural information processing systems , 2020, pp. 1877–1901
1901
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
B. Yamauchi, “A frontier-based approach for autonomous exploration,” in Proceedings IEEE International Symposium on Computational Intelligence in Robotics and Automation , 1997, pp. 146–151
1997
Earlier work this paper cites.
J. A. Sethian, “Fast marching methods,” SIAM review , vol. 41, no. 2, pp. 199–235, 1999
1999
Earlier work this paper cites.
R. Achanta, A. Shaji, K. Smith, A. Lucchi, P. Fua, and S. Süsstrunk, “Slic superpixels compared to state-of-the-art superpixel methods,” IEEE transactions on pattern analysis and machine intelligence , vol. 34, no. 11, pp. 2274–2282, 2012
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang, “Matterport3d: Learning from rgb-d data in indoor environments,” in International Conference on 3D Vision , 2017, pp. 667–676
2017
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2018, pp. 3674–3683
2018
Earlier work this paper cites.
X. Wang, W. Xiong, H. Wang, and W. Y. Wang, “Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation,” in Proceedings of the European Conference on Computer Vision , 2018, pp. 4392–4412
2018
Earlier work this paper cites.
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell, “Speaker-follower models for vision-and-language navigation,” in Advances in Neural Information Processing Systems , 2018, pp. 3318–3329
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
C.-Y. Ma, J. Lu, Z. Wu, G. AlRegib, Z. Kira, R. Socher, and C. Xiong, “Self-monitoring navigation agent via auxiliary progress estimation,” in International Conference on Learning Representations , 2019, pp. 5613–5630
2019
Earlier work this paper cites.
C.-Y. Ma, Z. Wu, G. AlRegib, C. Xiong, and Z. Kira, “The regretful agent: Heuristic-aided navigation through progress estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 6732–6740
2019
Earlier work this paper cites.
L. Ke, X. Li, Y. Bisk, A. Holtzman, Z. Gan, J. Liu, J. Gao, Y. Choi, and S. Srinivasa, “Tactical rewind: Self-correction via backtracking in vision-and-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 6741–6749
2019
Earlier work this paper cites.
X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y.-F. Wang, W. Y. Wang, and L. Zhang, “Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 6629–6638
2019
Earlier work this paper cites.
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik et al. , “Habitat: A platform for embodied ai research,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 9338–9346
2019
Earlier work this paper cites.
J. Krantz, E. Wijmans, A. Majundar, D. Batra, and S. Lee, “Beyond the nav-graph: Vision and language navigation in continuous environments,” in Proceedings of the European Conference on Computer Vision , 2020, pp. 104–120
2020
Earlier work this paper cites.
Y. Qi, Z. Pan, S. Zhang, A. v. d. Hengel, and Q. Wu, “Object-and-action aware model for visual language navigation,” in Proceedings of the European Conference on Computer Vision , 2020, pp. 303–317
2020
Earlier work this paper cites.
W. Hao, C. Li, X. Li, L. Carin, and J. Gao, “Towards learning a generic agent for vision-and-language navigation via pre-training,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 13 134–13 143
2020
Earlier work this paper cites.
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra, “Improving vision-and-language navigation with image-text pairs from the web,” in Proceedings of the European Conference on Computer Vision , 2020, pp. 259–274
2020
Earlier work this paper cites.
D. S. Chaplot, D. Gandhi, A. Gupta, and R. Salakhutdinov, “Object goal navigation using goal-oriented semantic exploration,” in Advances in Neural Information Processing Systems , 2020, pp. 4247–4258
2020
Earlier work this paper cites.
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge, “Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing , 2020, pp. 4392–4412
2020
Earlier work this paper cites.
A. Wunderlich and K. Gramann, “Landmark-based navigation instructions improve incidental spatial knowledge acquisition in real-world environments,” Journal of Environmental Psychology , vol. 77, pp. 101 645–101 677, 2021
2021
Earlier work this paper cites.
C. Gao, J. Chen, S. Liu, L. Wang, Q. Zhang, and Q. Wu, “Room-and-object aware knowledge reasoning for remote embodied referring expression,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 3064–3073
2021
Earlier work this paper cites.
Y. Qi, Z. Pan, Y. Hong, M.-H. Yang, A. van den Hengel, and Q. Wu, “The road to know-where: An object-and-room informed sequential bert for indoor vision-language navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1635–1644
2021
Earlier work this paper cites.
A. Moudgil, A. Majumdar, H. Agrawal, S. Lee, and D. Batra, “Soat: A scene-and object-aware transformer for vision-and-language navigation,” in Advances in Neural Information Processing Systems , 2021, pp. 7357–7367
2021
Earlier work this paper cites.
X. Lin, G. Li, and Y. Yu, “Scene-intuitive agent for remote embodied visual grounding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 7036–7045
2021
Earlier work this paper cites.
S. Chen, P.-L. Guhur, C. Schmid, and I. Laptev, “History aware multimodal transformer for vision-and-language navigation,” in Advances in Neural Information Processing Systems , 2021, pp. 5834–5847
2021
Earlier work this paper cites.
D. An, Y. Qi, Y. Huang, Q. Wu, L. Wang, and T. Tan, “Neighbor-view enhanced model for vision and language navigation,” in Proceedings of the ACM International Conference on Multimedia , 2021, pp. 5101–5109
2021
Cited alongside, same era.
Y. Hong, Q. Wu, Y. Qi, C. Rodriguez-Opazo, and S. Gould, “Vln bert: A recurrent vision-and-language bert for navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 1643–1653
2021
Cited alongside, same era.
P.-L. Guhur, M. Tapaswi, S. Chen, I. Laptev, and C. Schmid, “Airbert: In-domain pretraining for vision-and-language navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1614–1623
2021
Cited alongside, same era.
K. He, Y. Huang, Q. Wu, J. Yang, D. An, S. Sima, and L. Wang, “Landmark-rxr: Solving vision-and-language navigation with fine-grained alignment supervision,” in Advances in Neural Information Processing Systems , 2021, pp. 652–663
2021
2023
Later among the works it cites.
2023
Later among the works it cites.
K. Zhou, K. Zheng, C. Pryor, Y. Shen, H. Jin, L. Getoor, and X. E. Wang, “Esc: Exploration with soft commonsense constraints for zero-shot object navigation,” in International Conference on Machine Learning , 2023, pp. 42 829–42 842
2023
Later among the works it cites.
D. Shah, M. R. Equi, B. Osiński, F. Xia, B. Ichter, and S. Levine, “Navigation with large language models: Semantic guesswork as a heuristic for planning,” in Conference on Robot Learning , 2023, pp. 2683–2699
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
P. Anderson, A. Shrivastava, J. Truong, A. Majumdar, D. Parikh, D. Batra, and S. Lee, “Sim-to-real transfer for vision-and-language navigation,” in Conference on Robot Learning , 2021, pp. 671–681
2021
Cited alongside, same era.
J. Krantz, A. Gokaslan, D. Batra, S. Lee, and O. Maksymets, “Waypoint models for instruction-guided navigation in continuous environments,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 15 142–15 151
2021
Cited alongside, same era.
S. Raychaudhuri, S. Wani, S. Patel, U. Jain, and A. X. Chang, “Language-aligned waypoint (law) supervision for vision-and-language navigation in continuous environments,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing , 2021, pp. 4018–4028
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning , 2021, pp. 8748–8763
2021
Cited alongside, same era.
2021
Cited alongside, same era.
K. Chen, J. K. Chen, J. Chuang, M. Vázquez, and S. Savarese, “Topological planning with transformers for vision-and-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 11 276–11 286
2021
Cited alongside, same era.
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev, “Think global, act local: Dual-scale graph transformer for vision-and-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 516–16 526
2022
Cited alongside, same era.
Y. Qiao, Y. Qi, Y. Hong, Z. Yu, P. Wang, and Q. Wu, “Hop: History-and-order aware pre-training for vision-and-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 418–15 427
2022
Cited alongside, same era.
2023
Later among the works it cites.
S. Y. Gadre, M. Wortsman, G. Ilharco, L. Schmidt, and S. Song, “Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 23 171–23 181
2023
Later among the works it cites.
2023
Later among the works it cites.
D. An, H. Wang, W. Wang, Z. Wang, Y. Huang, K. He, and L. Wang, “Etpnav: Evolving topological planning for vision-language navigation in continuous environments,” IEEE Transactions on Pattern Analysis and Machine Intelligence , pp. 1–16, 2024
2024
Closest in time.
OpenAI, “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2024
2024
Closest in time.
G. Zhou, Y. Hong, and Q. Wu, “Navgpt: Explicit reasoning in vision-and-language navigation with large language models,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2024, pp. 7641–7649
2024
Closest in time.
J. Chen, B. Lin, R. Xu, Z. Chai, X. Liang, and K.-Y. K. Wong, “Mapgpt: Map-guided prompting for unified vision-and-language navigation,” in Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics , 2024, pp. 9796–9810
2024
Closest in time.
2024
Closest in time.
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su et al. , “Grounding dino: Marrying dino with grounded pre-training for open-set object detection,” in Proceedings of the European Conference on Computer Vision , 2024, pp. 38–55
2024
Closest in time.
H. Hong, S. Wang, Z. Huang, Q. Wu, and J. Liu, “Why only text: Empowering vision-and-language navigation with multi-modal prompts,” 2024, pp. 839–847
2024
Closest in time.
K. He, C. Si, Z. Lu, Y. Huang, L. Wang, and X. Wang, “Frequency-enhanced data augmentation for vision-and-language navigation,” in Advances in Neural Information Processing Systems , 2024, pp. 1–14
2024
Closest in time.
J. Li and M. Bansal, “Panogen: Text-conditioned panoramic environment generation for vision-and-language navigation,” in Advances in Neural Information Processing Systems , 2024, pp. 21 878–21 894
2024
Closest in time.
Z. Wang, X. Li, J. Yang, S. Jiang et al. , “Sim-to-real transfer via 3d feature fields for vision-and-language navigation,” in Conference on Robot Learning , 2024, pp. 1–14
2024
Closest in time.
J. Zhang, K. Wang, R. Xu, G. Zhou, Y. Hong, X. Fang, Q. Wu, Z. Zhang, and W. He, “Navid: Video-based vlm plans the next step for vision-and-language navigation,” Robotics: Science and Systems , pp. 1–17, 2024
2024
Closest in time.
D. Zheng, S. Huang, L. Zhao, Y. Zhong, and L. Wang, “Towards learning a generalist model for embodied navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 624–13 634
2024
Closest in time.
G. Zhou, Y. Hong, Z. Wang, X. E. Wang, and Q. Wu, “Navgpt-2: Unleashing navigational reasoning capability for large vision-language models,” in Proceedings of the European Conference on Computer Vision , 2024, pp. 260–278
2024
Closest in time.
2024
Closest in time.
Y. Long, X. Li, W. Cai, and H. Dong, “Discuss before moving: Visual language navigation via multi-expert discussions,” in Proceedings of the IEEE International Conference on Robotics and Automation , 2024, pp. 17 380–17 387
2024
Closest in time.
2024
Closest in time.
Y. Long, W. Cai, H. Wang, G. Zhan, and H. Dong, “Instructnav: Zero-shot system for generic instruction navigation in unexplored environment,” in Conference on Robot Learning , 2024, pp. 1–12
2024
Closest in time.
N. Yokoyama, S. Ha, D. Batra, J. Wang, and B. Bucher, “Vlfm: Vision-language frontier maps for zero-shot semantic navigation,” in Proceedings of the IEEE International Conference on Robotics and Automation , 2024, pp. 42–48
2024
Closest in time.
A. Wang, H. Chen, Z. Lin, J. Han, and G. Ding, “Repvit: Revisiting mobile cnn from vit perspective,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 15 909–15 920
2024
Closest in time.
L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” in Advances in Neural Information Processing Systems , 2024, pp. 21 875–21 911
2024
Closest in time.
J. Chen, B. Lin, X. Liu, X. Liang, and K.-Y. K. Wong, “Affordances-oriented planning using foundation models for continuous vision-language navigation,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2025, pp. 1–9
2025
Closest in time.