Fetching the paper…
Reading the bibliography…
Embodied navigation is a fundamental capability of embodied intelligence, enabling robots to move and interact within physical environments.
Robustnav: Towards benchmarking robustness in embodied navigation
Chattopadhyay, P.; Hoffman, J.; Mottaghi, R.; and Kembhavi, A. 2021 · 2021
Earlier work this paper cites.
Vln bert: A recurrent vision-and-language bert for navigation
Hong, Y.; Wu, Q.; Qi, Y.; Rodriguez-Opazo, C.; and Gould, S. 2021 · 2021
Earlier work this paper cites.
Embodied visual navigation with automatic curriculum learning in real environments
Morad, S. D.; Mecca, R.; Poudel, R. P.; Liwicki, S.; and Cipolla, R. 2021 · 2021
Earlier work this paper cites.
Bi-directional domain adaptation for sim2real transfer of embodied navigation agents
Truong, J.; Chernova, S.; and Batra, D. 2021 · 2021
Earlier work this paper cites.
Qilin-med-vl: Towards chinese large vision-language model for general healthcare
Liu, J.; Wang, Z.; Ye, Q.; Chong, D.; Zhou, P.; and Hua, Y. 2023 · 2023
Earlier work this paper cites.
Cheap and quick: Efficient vision-language instruction tuning for large language models
Luo, G.; Zhou, Y.; Ren, T.; Chen, S.; Sun, X.; and Ji, R. 2023 · 2023
Earlier work this paper cites.
Pirlnav: Pretraining with imitation and rl finetuning for objectnav
Ramrakhya, R.; Batra, D.; Wijmans, E.; and Das, A. 2023 · 2023
Earlier work this paper cites.
Visionllm: Large language model is also an open-ended decoder for vision-centric tasks
Wang, W.; Chen, Z.; Chen, X.; Wu, J.; Zhu, X.; Zeng, G.; Luo, P.; Lu, T.; Zhou, J.; Qiao, Y.; et al. 2023 · 2023
Earlier work this paper cites.
L3mvn: Leveraging large language models for visual target navigation
Yu, B.; Kasaei, H.; and Cao, M. 2023 · 2023
Earlier work this paper cites.
The claude 3 model family: Opus, sonnet, haiku
Anthropic. 2024 · 2024
Earlier work this paper cites.
Paligemma: A versatile 3b vlm for transfer
Beyer, L.; Steiner, A.; Pinto, A. S.; Kolesnikov, A.; Wang, X.; Salz, D.; Neumann, M.; Alabdulmohsin, I.; Tschannen, M.; Bugliarello, E.; et al. 2024 · 2024
Earlier work this paper cites.
Spatialrgpt: Grounded spatial reasoning in vision-language models
Cheng, A.-C.; Yin, H.; Fu, Y.; Guo, Q.; Yang, R.; Kautz, J.; Wang, X.; and Liu, S. 2024 · 2024
Earlier work this paper cites.
Towards Multimodal In-context Learning for Vision and Language Models
Doveh, S.; Perek, S.; Mirza, M. J.; Lin, W.; Alfassy, A.; Arbelle, A.; Ullman, S.; and Karlinsky, L. 2024 · 2024
Cited alongside, same era.
Hurst, A.; Lerer, A.; Goucher, A. P.; Perelman, A.; Ramesh, A.; Clark, A.; Ostrow, A.; Welihinda, A.; Hayes, A.; Radford, A.; et al. 2024 · 2024
Cited alongside, same era.
Vision-language model-driven scene understanding and robotic object manipulation
Liu, S.; Zhang, J.; Gao, R. X.; Wang, X. V.; and Wang, L. 2024 · 2024
Cited alongside, same era.
Discuss before moving: Visual language navigation via multi-expert discussions
Long, Y.; Li, X.; Cai, W.; and Dong, H. 2024 · 2024
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
O’Neill, A.; Rehman, A.; Maddukuri, A.; Gupta, A.; Padalkar, A.; Lee, A.; Pooley, A.; Gupta, A.; Mandlekar, A.; Jain, A.; et al. 2024 · 2024
Cited alongside, same era.
Gong, Z.; Li, R.; Hu, T.; Qiu, R.; Kong, L.; Zhang, L.; Ding, Y.; Zhang, L.; and Liang, J. 2025 · 2025
Closest in time.
Robobrain: A unified brain model for robotic manipulation from abstract to concrete
Ji, Y.; Tan, H.; Shi, J.; Hao, X.; Zhang, Y.; Zhang, H.; Wang, P.; Zhao, M.; Mu, Y.; An, P.; et al. 2025 · 2025
Closest in time.
Navcot: Boosting llm-based vision-and-language navigation via learning disentangled reasoning
Lin, B.; Nie, Y.; Wei, Z.; Chen, J.; Ma, S.; Han, J.; Xu, H.; Chang, X.; and Liang, X. 2025 · 2025
Closest in time.
Liu, Y.; Chi, D.; Wu, S.; Zhang, Z.; Hu, Y.; Zhang, L.; Zhang, Y.; Wu, S.; Cao, T.; Huang, G.; et al. 2025 · 2025
Closest in time.
VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploitation-guided exploration for semantic embodied navigation
Wasserman, J.; Chowdhary, G.; Gupta, A.; and Jain, U. 2024 · 2024
Cited alongside, same era.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning
Zhai, S.; Bai, H.; Lin, Z.; Pan, J.; Tong, P.; Zhou, Y.; Suhr, A.; Xie, S.; LeCun, Y.; Ma, Y.; et al. 2024 · 2024
Cited alongside, same era.
Navgpt: Explicit reasoning in vision-and-language navigation with large language models
Zhou, G.; Hong, Y.; and Wu, Q. 2024 · 2024
Cited alongside, same era.
Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; et al. 2025 · 2025
Cited alongside, same era.
Janus-pro: Unified multimodal understanding and generation with data and model scaling
Chen, X.; Wu, Z.; Liu, X.; Pan, Z.; Liu, W.; Xie, Z.; Yu, X.; and Ruan, C. 2025 · 2025
Cited alongside, same era.
Navila: Legged robot vision-language-action model for navigation
Cheng, A.-C.; Ji, Y.; Yang, Z.; Gongye, Z.; Zou, X.; Kautz, J.; Bıyık, E.; Yin, H.; Liu, S.; and Wang, X. 2025 · 2025
Cited alongside, same era.
OctoNav: Towards Generalist Embodied Navigation
Gao, C.; Jin, L.; Peng, X.; Zhang, J.; Deng, Y.; Li, A.; Wang, H.; and Liu, S. 2025 · 2025
Cited alongside, same era.
Qi, Z.; Zhang, Z.; Yu, Y.; Wang, J.; and Zhao, H. 2025 · 2025
Closest in time.
Reason-rft: Reinforcement fine-tuning for visual reasoning
Tan, H.; Ji, Y.; Hao, X.; Lin, M.; Wang, P.; Wang, Z.; and Zhang, S. 2025 · 2025
Closest in time.
Affordgrasp: In-context affordance reasoning for open-vocabulary task-oriented grasping in clutter
Tang, Y.; Zhang, S.; Hao, X.; Wang, P.; Wu, J.; Wang, Z.; and Zhang, S. 2025 · 2025
Closest in time.
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
Wu, J.; Guan, J.; Feng, K.; Liu, Q.; Wu, S.; Wang, L.; Wu, W.; and Tan, T. 2025 · 2025
Closest in time.
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction in Robotics
Yuan, W.; Duan, J.; Blukis, V.; Pumacay, W.; Krishna, R.; Murali, A.; Mousavian, A.; and Fox, D. 2025 · 2025
Closest in time.
Training-free Generation of Temporally Consistent Rewards from VLMs
Zhao, Y.; Yuan, J.; Xu, Z.; Hao, X.; Zhang, X.; Wu, K.; Che, Z.; Liu, C. H.; and Tang, J. 2025 · 2025
Closest in time.
Railway side slope hazard detection system based on generative models
Zheng, X.; He, Y.; Luo, Y.; Zhang, L.; Wang, J.; Shi, T.; and Bai, Y. 2025 · 2025
Closest in time.
InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
Long, Y.; Cai, W.; Wang, H.; Zhan, G.; and Dong, H. 2025 · 2060
Closest in time.