Fetching the paper…
Reading the bibliography…
Aerial Vision-and-Language Navigation (VLN) is a novel task enabling Unmanned Aerial Vehicles (UAVs) to navigate in outdoor environments through natural language instructions and visual cues.
Learning to navigate unseen environments: Back translation with environmental dropout
Tan, H.; Yu, L.; and Bansal, M. 2019 · 1904
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sünderhauf, N.; Reid, I.; Gould, S.; and Van Den Hengel, A. 2018 · 2018
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
Fried, D.; Hu, R.; Cirik, V.; Rohrbach, A.; Andreas, J.; Morency, L.-P.; Berg-Kirkpatrick, T.; Saenko, K.; Klein, D.; and Darrell, T. 2018 · 2018
Earlier work this paper cites.
Mapping instructions to actions in 3d environments with visual goal prediction
Misra, D.; Bennett, A.; Blukis, V.; Niklasson, E.; Shatkhin, M.; and Artzi, Y. 2018 · 2018
Earlier work this paper cites.
Text mining: use of TF-IDF to examine the relevance of words to documents
Qaiser, S.; and Ali, R. 2018 · 2018
Earlier work this paper cites.
Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation
Wang, X.; Xiong, W.; Wang, H.; and Wang, W. Y. 2018 · 2018
Earlier work this paper cites.
Towards learning a generic agent for vision-and-language navigation via pre-training
Hao, W.; Li, C.; Li, X.; Carin, L.; and Gao, J. 2020 · 2020
Earlier work this paper cites.
Language and visual entity relationship graph for agent navigation
Hong, Y.; Rodriguez, C.; Qi, Y.; Wu, Q.; and Gould, S. 2020 · 2020
Earlier work this paper cites.
Improving vision-and-language navigation with image-text pairs from the web
Majumdar, A.; Shrivastava, A.; Lee, S.; Anderson, P.; Parikh, D.; and Batra, D. 2020 · 2020
Earlier work this paper cites.
Object-and-action aware model for visual language navigation
Qi, Y.; Pan, Z.; Zhang, S.; van den Hengel, A.; and Wu, Q. 2020 · 2020
Earlier work this paper cites.
Soft expert reward learning for vision-and-language navigation
Wang, H.; Wu, Q.; and Shen, C. 2020 · 2020
Earlier work this paper cites.
Neighbor-view enhanced model for vision and language navigation
An, D.; Qi, Y.; Huang, Y.; Wu, Q.; Wang, L.; and Tan, T. 2021 · 2021
Earlier work this paper cites.
Do as I can, not as I say: Grounding language in robotic affordances
Ahn, M.; Brohan, A.; Brown, N.; Chebotar, Y.; Cortes, O.; David, B.; Finn, C.; Gopalakrishnan, K.; Hausman, K.; Herzog, A.; et al. 2022 · 2022
Earlier work this paper cites.
Can foundation models perform zero-shot task specification for robot manipulation?
Cui, Y.; Niekum, S.; Gupta, A.; Kumar, V.; and Rajeswaran, A. 2022 · 2022
Cited alongside, same era.
Clip-nav: Using clip for zero-shot vision-and-language navigation
Dorbala, V. S.; Sigurdsson, G.; Piramuthu, R.; Thomason, J.; and Sukhatme, G. S. 2022 · 2022
Cited alongside, same era.
Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation
Hong, Y.; Wang, Z.; Wu, Q.; and Gould, S. 2022 · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; and Zhou, D. 2022 · 2022
Cited alongside, same era.
Socratic models: Composing zero-shot multimodal reasoning with language
Language to rewards for robotic skill synthesis
Yu, W.; Gileadi, N.; Fu, C.; Kirmani, S.; Lee, K.-H.; Arenas, M. G.; Chiang, H.-T. L.; Erez, T.; Hasenclever, L.; Humplik, J.; et al. 2023 · 2023
Later among the works it cites.
Citynav: Language-goal aerial navigation dataset with geographic information
Lee, J.; Miyanishi, T.; Kurita, S.; Sakamoto, K.; Azuma, D.; Matsuo, Y.; and Inoue, N. 2024 · 2024
Closest in time.
TINA: Think, Interaction, and Action Framework for Zero-Shot Vision Language Navigation
Li, D.; Chen, W.; and Lin, X. 2024 · 2024
Closest in time.
Embodied agent interface: Benchmarking llms for embodied decision making
Li, M.; Zhao, S.; Wang, Q.; Wang, K.; Zhou, Y.; Srivastava, S.; Gokmen, C.; Lee, T.; Li, E. L.; Zhang, R.; et al. 2024 · 2024
Closest in time.
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zeng, A.; Attarian, M.; Ichter, B.; Choromanski, K.; Wong, A.; Welker, S.; Tombari, F.; Purohit, A.; Ryoo, M.; Sindhwani, V.; Lee, J.; Vanhoucke, V.; and Florence, P. 2022 · 2022
Cited alongside, same era.
Open-vocabulary queryable scene representations for real world planning
Chen, B.; Xia, F.; Ichter, B.; Rao, K.; Gopalakrishnan, K.; Ryoo, M. S.; Stone, A.; and Kappler, D. 2023 · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D.; Xia, F.; Sajjadi, M. S.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. 2023 · 2023
Cited alongside, same era.
Voxposer: Composable 3d value maps for robotic manipulation with language models
Huang, W.; Wang, C.; Zhang, R.; Li, Y.; Wu, J.; and Fei-Fei, L. 2023 · 2023
Cited alongside, same era.
Code as policies: Language model programs for embodied control
Liang, J.; Huang, W.; Xia, F.; Xu, P.; Hausman, K.; Ichter, B.; Florence, P.; and Zeng, A. 2023 · 2023
Cited alongside, same era.
Tokenize Anything via Prompting
Pan, T.; Tang, L.; Wang, X.; and Shan, S. 2023 · 2023
Cited alongside, same era.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
Shah, D.; Osiński, B.; Levine, S.; et al. 2023 · 2023
Cited alongside, same era.
Progprompt: Generating situated robot task plans using large language models
Singh, I.; Blukis, V.; Mousavian, A.; Goyal, A.; Xu, D.; Tremblay, J.; Fox, D.; Thomason, J.; and Garg, A. 2023 · 2023
Cited alongside, same era.
Lin, B.; Nie, Y.; Wei, Z.; Chen, J.; Ma, S.; Han, J.; Xu, H.; Chang, X.; and Liang, X. 2024 · 2024
Closest in time.
Improved baselines with visual instruction tuning
Liu, H.; Li, C.; Li, Y.; and Lee, Y. J. 2024 · 2024
Closest in time.
Towards realistic uav vision-language navigation: Platform, benchmark, and methodology
Wang, X.; Yang, D.; Wang, Z.; Kwan, H.; Chen, J.; W000000u, W.; Li, H.; Liao, Y.; and Liu, S. 2024 · 2024
Closest in time.
Navgpt: Explicit reasoning in vision-and-language navigation with large language models
Zhou, G.; Hong, Y.; and Wu, Q. 2024 · 2024
Closest in time.
Visual large language models for generalized and specialized applications
Li, Y.; Lai, Z.; Bao, W.; Tan, Z.; Dao, A.; Sui, K.; Shen, J.; Liu, D.; Liu, H.; and Kong, Y. 2025 · 2025
Closest in time.
UAV-VLA: Vision-language-action system for large scale aerial mission generation
Sautenkov, O.; Yaqoot, Y.; Lykov, A.; Mustafa, M. A.; Tadevosyan, G.; Akhmetkazy, A.; Cabrera, M. A.; Martynov, M.; Karaf, S.; and Tsetserukou, D. 2025 · 2025
Closest in time.
Vision-language models for edge networks: A comprehensive survey
Sharshar, A.; Khan, L. U.; Ullah, W.; and Guizani, M. 2025 · 2025
Closest in time.
Mind the gap: Benchmarking spatial reasoning in vision-language models
Stogiannidis, I.; McDonagh, S.; and Tsaftaris, S. A. 2025 · 2025
Closest in time.
UAVs meet LLMs: Overviews and perspectives towards agentic low-altitude mobility
Tian, Y.; Lin, F.; Li, Y.; Zhang, T.; Zhang, Q.; Fu, X.; Huang, J.; Dai, X.; Wang, Y.; Tian, C.; et al. 2025 · 2025
Closest in time.