Fetching the paper…
Reading the bibliography…
Enabling robots to navigate following diverse language instructions in unexplored environments is an attractive goal for human-robot interaction.
AI2-THOR: An Interactive 3D Environment for Visual AI
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi · 2017
Earlier work this paper cites.
Gibson env: Real-world perception for embodied agents
F. Xia, A. R. Zamir, Z. He, A. Sax, J. Malik, and S. Savarese · 2018
Earlier work this paper cites.
Beyond the nav-graph: Vision and language navigation in continuous environments
J. Krantz, E. Wijmans, A. Majundar, D. Batra, and S. Lee · 2020
Earlier work this paper cites.
Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge · 2020
Earlier work this paper cites.
ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects
D. Batra, A. Gokaslan, A. Kembhavi, O. Maksymets, R. Mottaghi, M. Savva, A. Toshev, and E. Wijmans · 2020
Earlier work this paper cites.
Improving vision-and-language navigation with image-text pairs from the web
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Object goal navigation using goal-oriented semantic exploration
D. S. Chaplot, D. Gandhi, A. Gupta, and R. Salakhutdinov · 2020
Earlier work this paper cites.
The marathon 2: A navigation system
S. Macenski, F. Martín, R. White, and J. Ginés Clavero · 2020
Earlier work this paper cites.
Soat: A scene-and object-aware transformer for vision-and-language navigation
A. Moudgil, A. Majumdar, H. Agrawal, S. Lee, and D. Batra · 2021
Earlier work this paper cites.
Airbert: In-domain Pretraining for Vision-and-Language Navigation, 2021
P.-L. Guhur, M. Tapaswi, S. Chen, I. Laptev, and C. Schmid · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Habitat-matterport 3d dataset (HM3d): 1000 large-scale 3d environments for embodied AI
S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. M. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang, M. Savva, Y. Zhao, and D. Batra · 2021
Earlier work this paper cites.
Slam toolbox: Slam for the dynamic world
S. Macenski and I. Jambrecic · 2021
Earlier work this paper cites.
Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation
Y. Hong, Z. Wang, Q. Wu, and S. Gould · 2022
Earlier work this paper cites.
Predicting dense and context-aware cost maps for semantic robot navigation
Y. Goel, N. Vaskevicius, L. Palmieri, N. Chebrolu, and C. Stachniss · 2022
Earlier work this paper cites.
Introducing chatgpt, 2022
OpenAI · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Habitat-web: Learning embodied object-search strategies from human demonstrations at scale
R. Ramrakhya, E. Undersander, D. Batra, and A. Das · 2022
Cited alongside, same era.
Semantically-aware spatio-temporal reasoning agent for vision-and-language navigation in continuous environments
M. Z. Irshad, N. C. Mithun, Z. Seymour, H.-P. Chiu, S. Samarasekera, and R. Kumar · 2022
Cited alongside, same era.
ProcTHOR: Large-Scale Embodied AI Using Procedural Generation
M. Deitke, E. VanderBilt, A. Herrasti, L. Weihs, J. Salvador, K. Ehsani, W. Han, E. Kolve, A. Farhadi, A. Kembhavi, and R. Mottaghi · 2022
Cited alongside, same era.
A tri-layer plugin to improve occluded detection
G. Zhan, W. Xie, and A. Zisserman · 2022
Cited alongside, same era.
Habitat challenge 2023, 2023
K. Yadav, J. Krantz, R. Ramrakhya, S. K. Ramakrishnan, J. Yang, A. Wang, J. Turner, A. Gokaslan, V.-P. Berges, R. Mootaghi, O. Maksymets, A. X. Chang, M. Savva, A. Clegg, D. S. Chaplot, and D. Batra · 2023
Cited alongside, same era.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
D. Shah, B. Osiński, S. Levine, et al · 2023
Later among the works it cites.
Navgpt: Explicit reasoning in vision-and-language navigation with large language models, 2023
G. Zhou, Y. Hong, and Q. Wu · 2023
Later among the works it cites.
Sayplan: Grounding large language models using 3d scene graphs for scalable task planning
K. Rana, J. Haviland, S. Garg, J. Abou-Chakra, I. Reid, and N. Suenderhauf · 2023
Later among the works it cites.
Discuss before moving: Visual language navigation via multi-expert discussions, 2023
Y. Long, X. Li, W. Cai, and H. Dong · 2023
Later among the works it cites.
Saynav: Grounding large language models for dynamic planning to navigation in new environments, 2023
A. Rajvanshi, K. Sikka, X. Lin, B. Lee, H.-P. Chiu, and A. Velasquez · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Find what you want: Learning demand-conditioned object attribute space for demand-driven navigation
H. Wang, A. G. H. Chen, X. Li, M. Wu, and H. Dong · 2023
Cited alongside, same era.
Esc: Exploration with soft commonsense constraints for zero-shot object navigation
K. Zhou, K. Zheng, C. Pryor, Y. Shen, H. Jin, L. Getoor, and X. E. Wang · 2023
Cited alongside, same era.
Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill, 2023
W. Cai, S. Huang, G. Cheng, Y. Long, P. Gao, C. Sun, and H. Dong · 2023
Cited alongside, same era.
Learning navigational visual representations with semantic map supervision, 2023
Y. Hong, Y. Zhou, R. Zhang, F. Dernoncourt, T. Bui, S. Gould, and H. Tan · 2023
Cited alongside, same era.
The dawn of lmms: Preliminary explorations with gpt-4v(ision), 2023
Z. Yang, L. Li, K. Lin, J. Wang, C.-C. Lin, Z. Liu, and L. Wang · 2023
Cited alongside, same era.
Vlfm: Vision-language frontier maps for zero-shot semantic navigation
N. H. Yokoyama, S. Ha, D. Batra, J. Wang, and B. Bucher · 2023
Cited alongside, same era.
Visual language maps for robot navigation
C. Huang, O. Mees, A. Zeng, and W. Burgard · 2023
Cited alongside, same era.
General object foundation model for images and videos at scale, 2023
J. Wu, Y. Jiang, Q. Liu, Z. Yuan, X. Bai, and S. Bai · 2023
Later among the works it cites.
Etpnav: Evolving topological planning for vision-language navigation in continuous environments
D. An, H. Wang, W. Wang, Z. Wang, Y. Huang, K. He, and L. Wang · 2023
Later among the works it cites.
Offline visual representation learning for embodied navigation
K. Yadav, R. Ramrakhya, A. Majumdar, V.-P. Berges, S. Kuhar, D. Batra, A. Baevski, and O. Maksymets · 2023
Later among the works it cites.
Zson: Zero-shot object-goal navigation using multimodal goal embeddings
A. Majumdar, G. Aggarwal, B. Devnani, J. Hoffman, and D. Batra · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Later among the works it cites.
Bevbert: Multimodal map pre-training for language-guided navigation
D. An, Y. Qi, Y. Li, Y. Huang, L. Wang, T. Tan, and J. Shao · 2023
Later among the works it cites.
From the desks of ros maintainers: A survey of modern and capable mobile robotics algorithms in the robot operating system 2
S. Macenski, T. Moore, D. Lu, A. Merzlyakov, and M. Ferguson · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
Voronav: Voronoi-based zero-shot object navigation with large language model, 2024
P. Wu, Y. Mu, B. Wu, Y. Hou, J. Ma, S. Zhang, and C. Liu · 2024
Closest in time.
Navid: Video-based vlm plans the next step for vision-and-language navigation
J. Zhang, K. Wang, R. Xu, G. Zhou, Y. Hong, X. Fang, Q. Wu, Z. Zhang, and W. He · 2024
Closest in time.
Amodal ground truth and completion in the wild
G. Zhan, C. Zheng, W. Xie, and A. Zisserman · 2024
Closest in time.