Fetching the paper…
Reading the bibliography…
The aspiration of the Vision-and-Language Navigation (VLN) task has long been to develop an embodied agent with robust adaptability, capable of seamlessly transferring its navigation capabilities across various tasks.
J. BORENSTEIN and Y. KOREN, “Real-time obstacle avoidance for fast mobile robots,” IEEE Transactions on systems, Man, and Cybernetics , vol. 19, no. 5, pp. 1179–1187, 1989
1989
Earlier work this paper cites.
J. Borenstein, Y. Koren et al. , “The vector field histogram-fast obstacle avoidance for mobile robots,” IEEE transactions on robotics and automation , vol. 7, no. 3, pp. 278–288, 1991
1991
Earlier work this paper cites.
M. G. Dissanayake, P. Newman, S. Clark, H. F. Durrant-Whyte, and M. Csorba, “A solution to the simultaneous localization and map building (slam) problem,” IEEE Transactions on robotics and automation , vol. 17, no. 3, pp. 229–241, 2001
2001
Earlier work this paper cites.
S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster R-CNN: towards real-time object detection with region proposal networks,” in Adv. Neural Inform. Process. Syst. , 2015, pp. 91–99
2015
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. D. Reid, S. Gould, and A. van den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 3674–3683
2018
Earlier work this paper cites.
F. Xia, A. R. Zamir, Z. He, A. Sax, J. Malik, and S. Savarese, “Gibson env: Real-world perception for embodied agents,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 9068–9079
2018
Earlier work this paper cites.
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niebner, M. Savva, S. Song, A. Zeng, and Y. Zhang, “Matterport3d: Learning from rgb-d data in indoor environments,” in 7th IEEE Int. Conf. on 3D Vision, 3DV 2017 , 2018, pp. 667–676
2018
Earlier work this paper cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 6077–6086
2018
Earlier work this paper cites.
Y. Qiao, Z. Yu, and Q. Wu, “Vln-petl: Parameter-efficient transfer learning for vision-and-language navigation,” in Int. Conf. Comput. Vis. , 2023, pp. 15 443–15 452
2018
Earlier work this paper cites.
J.-C. Piao and S.-D. Kim, “Real-time visual–inertial slam based on adaptive keyframe selection for mobile ar applications,” IEEE Transactions on Multimedia , vol. 21, no. 11, pp. 2827–2836, 2019
2019
Earlier work this paper cites.
V. Jain, G. Magalhaes, A. Ku, A. Vaswani, E. Ie, and J. Baldridge, “Stay on the path: Instruction fidelity in vision-and-language navigation,” in Proc. Conf. Assoc. Comput. Linguistics , Florence, Italy, 2019, pp. 1862–1872
2019
Earlier work this paper cites.
X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y.-F. Wang, W. Y. Wang, and L. Zhang, “Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 6629–6638
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge, “Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding,” in Proc. Conf. Empirical Methods Natural Lang. Process. Online: Association for Computational Linguistics, 2020, pp. 4392–4412
2020
Earlier work this paper cites.
J. Thomason, M. Murray, M. Cakmak, and L. Zettlemoyer, “Vision-and-dialog navigation,” in Conf. on Robot Learn. , 2020, pp. 394–406
2020
Earlier work this paper cites.
Y. Qi, Q. Wu, P. Anderson, X. Wang, W. Y. Wang, C. Shen, and A. van den Hengel, “REVERIE: remote embodied visual referring expression in real indoor environments,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2020, pp. 9979–9988
2020
Earlier work this paper cites.
W. Hao, C. Li, X. Li, L. Carin, and J. Gao, “Towards learning a generic agent for vision-and-language navigation via pre-training,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2020, pp. 13 134–13 143
2020
Earlier work this paper cites.
J. Krantz, E. Wijmans, A. Majumdar, D. Batra, and S. Lee, “Beyond the nav-graph: Vision-and-language navigation in continuous environments,” in Eur. Conf. Comput. Vis. Springer, 2020, pp. 104–120
2020
Earlier work this paper cites.
P.-J. Duh, Y.-C. Sung, L.-Y. F. Chiang, Y.-J. Chang, and K.-W. Chen, “V-eye: A vision-based navigation system for the visually impaired,” IEEE Transactions on Multimedia , vol. 23, pp. 1567–1580, 2020
2020
Earlier work this paper cites.
Y. Hong, C. Rodriguez, Y. Qi, Q. Wu, and S. Gould, “Language and visual entity relationship graph for agent navigation,” Adv. Neural Inform. Process. Syst. , vol. 33, pp. 7685–7696, 2020
2020
Earlier work this paper cites.
Z. Deng, K. Narasimhan, and O. Russakovsky, “Evolving graphical planner: Contextual global planning for vision-and-language navigation,” Advances in Neural Information Processing Systems , vol. 33, pp. 20 660–20 672, 2020
2020
Earlier work this paper cites.
W. Zhu, H. Hu, J. Chen, Z. Deng, V. Jain, E. Ie, and F. Sha, “BabyWalk: Going farther in vision-and-language navigation by taking baby steps,” in Proc. Conf. Assoc. Comput. Linguistics , 2020, pp. 2539–2556
2020
Earlier work this paper cites.
Y. Hong, C. Rodriguez, Q. Wu, and S. Gould, “Sub-instruction aware vision-and-language navigation,” in Proc. Conf. Empirical Methods Natural Lang. Process. , 2020, pp. 3360–3376
2020
Earlier work this paper cites.
T.-J. Fu, X. E. Wang, M. F. Peterson, S. T. Grafton, M. P. Eckstein, and W. Y. Wang, “Counterfactual vision-and-language navigation via adversarial path sampler,” in Eur. Conf. Comput. Vis. , 2020, pp. 71–86
2020
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in Adv. Neural Inform. Process. Syst. , 2020
2020
Earlier work this paper cites.
F. Zhu, X. Liang, Y. Zhu, Q. Yu, X. Chang, and X. Liang, “SOON: scenario oriented object navigation with graph-based exploration,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2021, pp. 12 689–12 699
2021
Earlier work this paper cites.
S. Chen, P. Guhur, C. Schmid, and I. Laptev, “History aware multimodal transformer for vision-and-language navigation,” in Adv. Neural Inform. Process. Syst. , 2021, pp. 5834–5847
2021
Earlier work this paper cites.
J. Li, H. Tan, and M. Bansal, “Improving cross-modal alignment in vision language navigation via syntactic information,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics , 2021, pp. 1041–1050
2021
Earlier work this paper cites.
K. Chen, J. K. Chen, J. Chuang, M. Vázquez, and S. Savarese, “Topological planning with transformers for vision-and-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 11 276–11 286
2021
Earlier work this paper cites.
Y. Hong, Q. Wu, Y. Qi, C. R. Opazo, and S. Gould, “A recurrent vision-and-language BERT for navigation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2021
2021
Earlier work this paper cites.
P.-L. Guhur, M. Tapaswi, S. Chen, I. Laptev, and C. Schmid, “Airbert: In-domain pretraining for vision-and-language navigation,” in Int. Conf. Comput. Vis. , 2021, pp. 1634–1643
2021
Cited alongside, same era.
K. He, Y. Huang, Q. Wu, J. Yang, D. An, S. Sima, and L. Wang, “Landmark-rxr: Solving vision-and-language navigation with fine-grained alignment supervision,” in Adv. Neural Inform. Process. Syst. , 2021, pp. 652–663
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in Int. Conf. Mach. Learn. , 2021, pp. 8748–8763
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
M. Hwang, J. Jeong, M. Kim, Y. Oh, and S. Oh, “Meta-explore: Exploratory hierarchical vision-and-language navigation using scene object spectrum grounding,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2023, pp. 6683–6693
2023
Later among the works it cites.
C. Gao, X. Peng, M. Yan, H. Wang, L. Yang, H. Ren, H. Li, and S. Liu, “Adaptive zone-aware hierarchical planner for vision-language navigation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2023, pp. 14 911–14 920
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
J. Gu, E. Stefani, Q. Wu, J. Thomason, and X. Wang, “Vision-and-language navigation: A survey of tasks, methods, and future directions,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , 2022, pp. 7606–7623
2022
Cited alongside, same era.
W. Zhu, Y. Qi, P. Narayana, K. Sone, S. Basu, X. Wang, Q. Wu, M. Eckstein, and W. Y. Wang, “Diagnosing vision-and-language navigation: What really matters,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics , 2022, pp. 5981–5993
2022
Cited alongside, same era.
S. Chen, P. Guhur, M. Tapaswi, C. Schmid, and I. Laptev, “Think global, act local: Dual-scale graph transformer for vision-and-language navigation,” in IEEE Conf. Comput. Vis. Pattern Recog. IEEE, 2022, pp. 16 516–16 526
2022
Cited alongside, same era.
G. Georgakis, K. Schmeckpeper, K. Wanchoo, S. Dan, E. Miltsakaki, D. Roth, and K. Daniilidis, “Cross-modal map learning for vision and language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 460–15 470
2022
Cited alongside, same era.
M. Z. Irshad, N. C. Mithun, Z. Seymour, H.-P. Chiu, S. Samarasekera, and R. Kumar, “Semantically-aware spatio-temporal reasoning agent for vision-and-language navigation in continuous environments,” in 2022 26th International Conference on Pattern Recognition (ICPR) . IEEE, 2022, pp. 4065–4071
2022
Cited alongside, same era.
P. Chen, D. Ji, K. Lin, R. Zeng, T. Li, M. Tan, and C. Gan, “Weakly-supervised multi-granularity map learning for vision-and-language navigation,” Advances in Neural Information Processing Systems , vol. 35, pp. 38 149–38 161, 2022
2022
Cited alongside, same era.
Y. Zhang and P. Kordjamshidi, “LOViS: Learning orientation and visual signals for vision and language navigation,” in Proceedings of the 29th International Conference on Computational Linguistics . Gyeongju, Republic of Korea: International Committee on Computational Linguistics, 2022, pp. 5745–5754
2022
Cited alongside, same era.
Later among the works it cites.
Y. Hong, Y. Zhou, R. Zhang, F. Dernoncourt, T. Bui, S. Gould, and H. Tan, “Learning navigational visual representations with semantic map supervision,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3055–3067
2023
Later among the works it cites.
Y. Cui, L. Xie, Y. Zhang, M. Zhang, Y. Yan, and E. Yin, “Grounded entity-landmark adaptive pre-training for vision-and-language navigation,” in Int. Conf. Comput. Vis. , 2023, pp. 12 043–12 053
2023
Later among the works it cites.
X. Li, Z. Wang, J. Yang, Y. Wang, and S. Jiang, “Kerm: Knowledge enhanced reasoning for vision-and-language navigation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2023, pp. 2583–2592
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Zhang and P. Kordjamshidi, “Vln-trans: Translator for the vision and language navigation agent,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , 2023, pp. 13 219–13 233
2023
Later among the works it cites.
Q. Qi, A. Zhang, Y. Liao, W. Sun, Y. Wang, X. Li, and S. Liu, “Simultaneously training and compressing vision-and-language pre-training model,” IEEE Transactions on Multimedia , vol. 25, pp. 8194–8203, 2023
2023
Later among the works it cites.
D. Shah, B. Osiński, S. Levine et al. , “Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,” in Conf. on Robot Learn. , 2023, pp. 492–504
2023
Later among the works it cites.
Y. Qiao, Y. Qi, Z. Yu, J. Liu, and Q. Wu, “March in chat: Interactive prompting for remote embodied referring expression,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 15 758–15 767
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
K. Zhou, K. Zheng, C. Pryor, Y. Shen, H. Jin, L. Getoor, and X. E. Wang, “Esc: Exploration with soft commonsense constraints for zero-shot object navigation,” in Int. Conf. Mach. Learn. , 2023, pp. 42 829–42 842
2023
Later among the works it cites.
2023
Later among the works it cites.
C. H. Song, J. Wu, C. Washington, B. M. Sadler, W.-L. Chao, and Y. Su, “Llm-planner: Few-shot grounded planning for embodied agents with large language models,” in Int. Conf. Comput. Vis. , 2023, pp. 2998–3009
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International conference on machine learning . PMLR, 2023, pp. 19 730–19 742
2023
Later among the works it cites.
X. Wang, W. Wang, J. Shao, and Y. Yang, “Lana: A language-capable navigator for instruction following and generation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2023, pp. 19 048–19 058
2023
Later among the works it cites.
J. Li and M. Bansal, “Panogen: Text-conditioned panoramic environment generation for vision-and-language navigation,” Adv. Neural Inform. Process. Syst. , vol. 36, 2024
2024
Later among the works it cites.
G. Zhou, Y. Hong, and Q. Wu, “Navgpt: Explicit reasoning in vision-and-language navigation with large language models,” in AAAI , vol. 38, no. 7, 2024, pp. 7641–7649
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Wu, X. Fu, F. Wu, and Z.-J. Zha, “Vision-and-language navigation via latent semantic alignment learning,” IEEE Trans. Multimedia , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Li, B. Hu, X. Chen, L. Ma, Y. Xu, and M. Zhang, “Lmeye: An interactive perception network for large language models,” IEEE Transactions on Multimedia , pp. 1–13, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu et al. , “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 24 185–24 198
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Qiao, Q. Liu, J. Liu, J. Liu, and Q. Wu, “Llm as copilot for coarse-grained vision-and-language navigation,” in European Conference on Computer Vision . Springer, 2025, pp. 459–476
2025
Closest in time.