Fetching the paper…
Reading the bibliography…
Object-goal navigation requires mobile robots to efficiently locate targets with visual and spatial information, yet existing methods struggle with generalization in unseen environments.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in Proceedings of the International Conference on Machine Learning , 2016, pp. 1928–1937
1937
Earlier work this paper cites.
B. Yamauchi, “A frontier-based approach for autonomous exploration,” in Proceedings 1997 IEEE International Symposium on Computational Intelligence in Robotics and Automation CIRA’97. ’Towards New Computational Principles for Robotics and Automation’ , 1997, pp. 146–151
1997
Earlier work this paper cites.
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi, “AI2-THOR: An Interactive 3D Environment for Visual AI,” arXiv , 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, D. Parikh, and D. Batra, “Habitat: A Platform for Embodied AI Research,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2019
2019
Earlier work this paper cites.
M. Wortsman, K. Ehsani, M. Rastegari, A. Farhadi, and R. Mottaghi, “Learning to learn how to learn: Self-adaptive visual navigation using meta-learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 6750–6759
2019
Earlier work this paper cites.
N. Reimers and I. Gurevych, “Sentence-bert: Sentence embeddings using siamese bert-networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 11 2019
2019
Earlier work this paper cites.
W. Yang, X. Wang, A. Farhadi, A. Gupta, and R. Mottaghi, “Visual semantic navigation using scene priors,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
H. Li, Q. Zhang, and D. Zhao, “Deep reinforcement learning-based automatic exploration for navigation in unknown environment,” IEEE Transactions on Neural Networks and Learning Systems , vol. 31, no. 6, pp. 2064–2076, 2020
2020
Earlier work this paper cites.
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox, “Alfred: A benchmark for interpreting grounded instructions for everyday tasks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 10 740–10 749
2020
Earlier work this paper cites.
H. Du, X. Yu, and L. Zheng, “Learning object relation graph and tentative policy for visual navigation,” in Computer Vision – ECCV 2020 , A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm, Eds. Cham: Springer International Publishing, 2020, pp. 19–34
2020
Earlier work this paper cites.
E. Wijmans, A. Kadian, A. Morcos, S. Lee, I. Essa, D. Parikh, M. Savva, and D. Batra, “DD-PPO: learning near-perfect pointgoal navigators from 2.5 billion frames,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
S. Zhang, X. Song, Y. Bai, W. Li, Y. Chu, and S. Jiang, “Hierarchical object-to-zone graph for object navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 15 130–15 140
2021
Earlier work this paper cites.
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch, “Decision transformer: Reinforcement learning via sequence modeling,” in Advances in Neural Information Processing Systems , vol. 34. Curran Associates, Inc., 2021, pp. 15 084–15 097
2021
Earlier work this paper cites.
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 9650–9660
2021
Cited alongside, same era.
C. Lin, Y. Jiang, J. Cai, L. Qu, G. Haffari, and Z. Yuan, “Multimodal transformer with variable-length memory for vision-and-language navigation,” in Computer Vision – ECCV 2022 , 2022, pp. 380–397
2022
Cited alongside, same era.
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev, “Think global, act local: Dual-scale graph transformer for vision-and-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 537–16 547
2022
Cited alongside, same era.
S. Y. Min, D. S. Chaplot, P. K. Ravikumar, Y. Bisk, and R. Salakhutdinov, “FILM: following instructions in language with modular methods,” in International Conference on Learning Representations , 2022
2022
A. J. Zhai and S. Wang, “Peanut: predicting and navigating to unseen targets,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 10 926–10 935
2023
Later among the works it cites.
S. Y. Gadre, M. Wortsman, G. Ilharco, L. Schmidt, and S. Song, “Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2023, pp. 23 171–23 181
2023
Later among the works it cites.
B. Yu, H. Kasaei, and M. Cao, “L3mvn: Leveraging large language models for visual target navigation,” in 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2023
2023
Later among the works it cites.
C. H. Song, J. Wu, C. Washington, B. M. Sadler, W.-L. Chao, and Y. Su, “Llm-planner: Few-shot grounded planning for embodied agents with large language models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2998–3009
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
S. K. Ramakrishnan, D. S. Chaplot, Z. Al-Halah, J. Malik, and K. Grauman, “Poni: Potential functions for objectgoal navigation with interaction-free learning,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022, pp. 18 868–18 878
2022
Cited alongside, same era.
2022
Cited alongside, same era.
K. Yadav, S. K. Ramakrishnan, A. Gokaslan, O. Maksymets, R. Jain, R. Ramrakhya, A. X. Chang, A. Clegg, M. Savva, E. Undersander, D. S. Chaplot, and D. Batra, “Habitat challenge 2022,” https://aihabitat.org/challenge/2022/ , 2022
2022
Cited alongside, same era.
V. Blukis, C. Paxton, D. Fox, A. Garg, and Y. Artzi, “A persistent spatial semantic representation for high-level natural language instruction execution,” in Conference on Robot Learning . PMLR, 2022, pp. 706–717
2022
Cited alongside, same era.
M. Murray and M. Cakmak, “Following natural language instructions for household tasks with landmark guided search and reinforced pose adjustment,” IEEE Robotics and Automation Letters , vol. 7, no. 3, pp. 6870–6877, 2022
2022
Cited alongside, same era.
X. Liu, H. Palacios, and C. Muise, “A planning based neural-symbolic approach for embodied instruction following,” Interactions , vol. 9, no. 8, p. 17, 2022
2022
Cited alongside, same era.
R. Ramrakhya, E. Undersander, D. Batra, and A. Das, “Habitat-web: Learning embodied object-search strategies from human demonstrations at scale,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 5173–5183
2022
Cited alongside, same era.
M. Deitke, E. VanderBilt, A. Herrasti, L. Weihs, K. Ehsani, J. Salvador, W. Han, E. Kolve, A. Kembhavi, and R. Mottaghi, “Procthor: Large-scale embodied ai using procedural generation,” Advances in Neural Information Processing Systems , vol. 35, pp. 5982–5994, 2022
2022
Cited alongside, same era.
Later among the works it cites.
B. Kim, J. Kim, Y. Kim, C. Min, and J. Choi, “Context-aware planning and environment-aware memory for instruction following embodied agents,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 10 936–10 946
2023
Later among the works it cites.
R. Ramrakhya, D. Batra, E. Wijmans, and A. Das, “Pirlnav: Pretraining with imitation and rl finetuning for objectnav,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 17 896–17 906
2023
Later among the works it cites.
A. Rajvanshi, K. Sikka, X. Lin, B. Lee, H.-P. Chiu, and A. Velasquez, “Saynav: Grounding large language models for dynamic planning to navigation in new environments,” in Proceedings of the International Conference on Automated Planning and Scheduling , vol. 34, 2024, pp. 464–474
2024
Closest in time.
G. Zhou, Y. Hong, and Q. Wu, “Navgpt: Explicit reasoning in vision-and-language navigation with large language models,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 7, 2024, pp. 7641–7649
2024
Closest in time.
B. Li, H. Li, Y. Zhu, and D. Zhao, “Mat: Morphological adaptive transformer for universal morphology policy learning,” IEEE Transactions on Cognitive and Developmental Systems , vol. 16, no. 4, pp. 1611–1621, 2024
2024
Closest in time.
K.-H. Zeng, Z. Zhang, K. Ehsani, R. Hendrix, J. Salvador, A. Herrasti, R. Girshick, A. Kembhavi, and L. Weihs, “Poliformer: Scaling on-policy rl with transformers results in masterful navigators,” in Proceedings of the Conference on Robot Learning (CoRL) , 2024
2024
Closest in time.
Y. Chen, X. Zhang, Y. Chen, D. Zhao, Y. Zhao, Z. Zhao, and P. Hu, “Common sense language-guided exploration and hierarchical dense perception for instruction following embodied agents,” in 2024 IEEE International Conference on Multimedia and Expo (ICME) , 2024, pp. 1–6
2024
Closest in time.
J. Sun, J. Wu, Z. Ji, and Y.-K. Lai, “A survey of object goal navigation,” IEEE Transactions on Automation Science and Engineering , vol. 22, pp. 2292–2308, 2025
2025
Closest in time.
Y. Chen, H. Li, Y. Chen, and D. Zhao, “Leaffordnav: Enhancing open-vocabulary mobile manipulation with llm-guided exploration and affordance-aware navigation,” in 2025 IEEE International Conference on Multimedia and Expo (ICME) , 2025
2025
Closest in time.
Y. Chen, W. Cui, Y. Chen, M. Tan, X. Zhang, J. Liu, H. Li, D. Zhao, and H. Wang, “Robogpt: an llm-based long-term decision-making embodied agent for instruction following tasks,” IEEE Transactions on Cognitive and Developmental Systems , pp. 1–11, 2025
2025
Closest in time.