Fetching the paper…
Reading the bibliography…
LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) task.
Self-monitoring navigation agent via auxiliary progress estimation
Ma, C.-Y.; Lu, J.; Wu, Z.; AlRegib, G.; Kira, Z.; Socher, R.; and Xiong, C. 2019 · 1901
Earlier work this paper cites.
Learning to move with affordance maps
Qi, W.; Mullapudi, R. T.; Gupta, S.; and Ramanan, D. 2020a · 2001
Earlier work this paper cites.
The ecological approach to visual perception: classic edition
Gibson, J. J. 2014 · 2014
Earlier work this paper cites.
Matterport3D: Learning from RGB-D Data in Indoor Environments
Chang, A.; Dai, A.; Funkhouser, T.; Halber, M.; Niebner, M.; Savva, M.; Song, S.; Zeng, A.; and Zhang, Y. 2017 · 2017
Earlier work this paper cites.
Learning to segment affordances
Luddecke, T.; and Worgotter, F. 2017 · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sünderhauf, N.; Reid, I.; Gould, S.; and Van Den Hengel, A. 2018 · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2018 · 2018
Earlier work this paper cites.
Habitat: A platform for embodied ai research
Savva, M.; Kadian, A.; Maksymets, O.; Zhao, Y.; Wijmans, E.; Jain, B.; Straub, J.; Liu, J.; Koltun, V.; Malik, J.; et al. 2019 · 2019
Earlier work this paper cites.
Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout
Tan, H.; Yu, L.; and Bansal, M. 2019 · 2019
Earlier work this paper cites.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
Wang, X.; Huang, Q.; Celikyilmaz, A.; Gao, J.; Shen, D.; Wang, Y.-F.; Wang, W. Y.; and Zhang, L. 2019 · 2019
Earlier work this paper cites.
Evolving graphical planner: Contextual global planning for vision-and-language navigation
Deng, Z.; Narasimhan, K.; and Russakovsky, O. 2020 · 2020
Earlier work this paper cites.
Beyond the nav-graph: Vision-and-language navigation in continuous environments
Krantz, J.; Wijmans, E.; Majumdar, A.; Batra, D.; and Lee, S. 2020 · 2020
Earlier work this paper cites.
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding
Ku, A.; Anderson, P.; Patel, R.; Ie, E.; and Baldridge, J. 2020 · 2020
Earlier work this paper cites.
History aware multimodal transformer for vision-and-language navigation
Chen, S.; Guhur, P.-L.; Schmid, C.; and Laptev, I. 2021 · 2021
Earlier work this paper cites.
Airbert: In-domain pretraining for vision-and-language navigation
Guhur, P.-L.; Tapaswi, M.; Chen, S.; Laptev, I.; and Schmid, C. 2021 · 2021
Earlier work this paper cites.
vlnbert: A recurrent vision-and-language bert for navigation
Hong, Y.; Wu, Q.; Qi, Y.; Rodriguez-Opazo, C.; and Gould, S. 2021 · 2021
Earlier work this paper cites.
Hierarchical cross-modal agent for robotics vision-and-language navigation
Irshad, M. Z.; Ma, C.-Y.; and Kira, Z. 2021 · 2021
Earlier work this paper cites.
Waypoint models for instruction-guided navigation in continuous environments
Krantz, J.; Gokaslan, A.; Batra, D.; Lee, S.; and Maksymets, O. 2021 · 2021
Cited alongside, same era.
Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments
Raychaudhuri, S.; Wani, S.; Patel, S.; Jain, U.; and Chang, A. 2021 · 2021
Cited alongside, same era.
SASRA: Semantically-aware Spatio-temporal Reasoning Agent for Vision-and-Language Navigation in Continuous Environments
Zubair Irshad, M.; Chowdhury Mithun, N.; Seymour, Z.; Chiu, H.-P.; Samarasekera, S.; and Kumar, R. 2021 · 2021
Cited alongside, same era.
BEVBert: Topo-Metric Map Pre-training for Language-guided Navigation
An, D.; Qi, Y.; Li, Y.; Huang, Y.; Wang, L.; Tan, T.; and Shao, J. 2022 · 2022
Cited alongside, same era.
Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation
Chen, S.; Guhur, P.-L.; Tapaswi, M.; Schmid, C.; and Laptev, I. 2022 · 2022
LangNav: Language as a Perceptual Representation for Navigation
Pan, B.; Panda, R.; Jin, S.; Feris, R.; Oliva, A.; Isola, P.; and Kim, Y. 2023 · 2023
Later among the works it cites.
March in Chat: Interactive Prompting for Remote Embodied Referring Expression
Qiao, Y.; Qi, Y.; Yu, Z.; Liu, J.; and Wu, Q. 2023 · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Later among the works it cites.
NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Clip-nav: Using clip for zero-shot vision-and-language navigation
Dorbala, V. S.; Sigurdsson, G.; Piramuthu, R.; Thomason, J.; and Sukhatme, G. S. 2022 · 2022
Cited alongside, same era.
CLIP on Wheels: Zero-Shot Object Navigation as Object Localization and Exploration
Gadre, S. Y.; Wortsman, M.; Ilharco, G.; Schmidt, L.; and Song, S. 2022 · 2022
Cited alongside, same era.
Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation
Hong, Y.; Wang, Z.; Wu, Q.; and Gould, S. 2022 · 2022
Cited alongside, same era.
Sim-2-sim transfer for vision-and-language navigation in continuous environments
Krantz, J.; and Lee, S. 2022 · 2022
Cited alongside, same era.
Zson: Zero-shot object-goal navigation using multimodal goal embeddings
Majumdar, A.; Aggarwal, G.; Devnani, B.; Hoffman, J.; and Batra, D. 2022 · 2022
Cited alongside, same era.
HOP: History-and-Order Aware Pre-training for Vision-and-Language Navigation
Qiao, Y.; Qi, Y.; Hong, Y.; Yu, Z.; Wang, P.; and Wu, Q. 2022 · 2022
Cited alongside, same era.
Anil, R.; Dai, A. M.; Firat, O.; Johnson, M.; Lepikhin, D.; Passos, A.; Shakeri, S.; Taropa, E.; Bailey, P.; Chen, Z.; et al. 2023 · 2023
Cited alongside, same era.
Zhou, G.; Hong, Y.; and Wu, Q. 2023 · 2023
Later among the works it cites.
Etpnav: Evolving topological planning for vision-language navigation in continuous environments
An, D.; Wang, H.; Wang, W.; Wang, Z.; Huang, Y.; He, K.; and Wang, L. 2024 · 2024
Closest in time.
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
Chen, J.; Lin, B.; Xu, R.; Chai, Z.; Liang, X.; and Wong, K.-Y. K. 2024 · 2024
Closest in time.
Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models
Lei, X.; Yang, Z.; Chen, X.; Li, P.; and Liu, Y. 2024 · 2024
Closest in time.
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
Lin, B.; Nie, Y.; Wei, Z.; Chen, J.; Ma, S.; Han, J.; Xu, H.; Chang, X.; and Liang, X. 2024 · 2024
Closest in time.
Openeqa: Embodied question answering in the era of foundation models
Majumdar, A.; Ajay, A.; Zhang, X.; Putta, P.; Yenamandra, S.; Henaff, M.; Silwal, S.; Mcvay, P.; Maksymets, O.; Arnaud, S.; et al. 2024 · 2024
Closest in time.
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
Nasiriany, S.; Xia, F.; Yu, W.; Xiao, T.; Liang, J.; Dasgupta, I.; Xie, A.; Driess, D.; Wahid, A.; Xu, Z.; et al. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M.; Savinov, N.; Teplyashin, D.; Lepikhin, D.; Lillicrap, T.; Alayrac, J.-b.; Soricut, R.; Lazaridou, A.; Firat, O.; Schrittwieser, J.; et al. 2024 · 2024
Closest in time.
Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
Ren, T.; Liu, S.; Zeng, A.; Lin, J.; Li, K.; Cao, H.; Chen, J.; Huang, X.; Chen, Y.; Yan, F.; Zeng, Z.; Zhang, H.; Li, F.; Yang, J.; Li, H.; Jiang, Q.; and Zhang, L. 2024 · 2024
Closest in time.
Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language Navigation
Wang, Z.; Li, X.; Yang, J.; Liu, Y.; Hu, J.; Jiang, M.; and Jiang, S. 2024 · 2024
Closest in time.
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
Zhang, J.; Wang, K.; Xu, R.; Zhou, G.; Hong, Y.; Fang, X.; Wu, Q.; Zhang, Z.; and He, W. 2024 · 2024
Closest in time.
Towards learning a generalist model for embodied navigation
Zheng, D.; Huang, S.; Zhao, L.; Zhong, Y.; and Wang, L. 2024 · 2024
Closest in time.