Fetching the paper…
Reading the bibliography…
Large scale scenes such as multifloor homes can be robustly and efficiently mapped with a 3D graph of landmarks estimated jointly with robot poses in a factor graph, a technique commonly used in commercial robots such as drones and robot vacuums.
B. Yamauchi, “A frontier-based approach for autonomous exploration,” in Proceedings 1997 IEEE International Symposium on Computational Intelligence in Robotics and Automation CIRA’97.’Towards New Computational Principles for Robotics and Automation’ . IEEE, 1997, pp. 146–151
1997
Earlier work this paper cites.
F. Dellaert, S. Seitz, S. Thrun, and C. Thorpe, “Feature correspondence: A markov chain monte carlo approach,” Advances in Neural Information Processing Systems , vol. 13, 2000
2000
Earlier work this paper cites.
S. Thrun and M. Montemerlo, “The graph slam algorithm with applications to large-scale mapping of urban structures,” The International Journal of Robotics Research , vol. 25, no. 5-6, pp. 403–429, 2006
2006
Earlier work this paper cites.
E. Eade, P. Fong, and M. E. Munich, “Monocular graph slam with complexity reduction,” in 2010 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 2010, pp. 3017–3024
2010
Earlier work this paper cites.
W. Knight, “With a roomba capable of navigation, irobot eyes advanced home robots,” MIT Technology Review , 2015
2015
Earlier work this paper cites.
R. Mur-Artal and J. D. Tardós, “Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,” IEEE transactions on robotics , vol. 33, no. 5, pp. 1255–1262, 2017
2017
Earlier work this paper cites.
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang, “Matterport3d: Learning from rgb-d data in indoor environments,” Intl. Conf. on 3D Comput. Vis. , 2017
2017
Earlier work this paper cites.
F. Dellaert, M. Kaess, et al. , “Factor graphs for robot perception,” Foundations and Trends® in Robotics , vol. 6, no. 1-2, pp. 1–139, 2017
2017
Earlier work this paper cites.
S. L. Bowman, N. Atanasov, K. Daniilidis, and G. J. Pappas, “Probabilistic data association for semantic slam,” in ICRA . IEEE, 2017, pp. 1722–1729
2017
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in CVPR , 2018, pp. 3674–3683
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Labbé and F. Michaud, “Rtab-map as an open-source lidar and visual simultaneous localization and mapping library for large-scale and long-term online operation,” Journal of field robotics , vol. 36, no. 2, pp. 416–446, 2019
2019
Earlier work this paper cites.
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, et al. , “Habitat: A platform for embodied ai research,” in ICCV , 2019, pp. 9339–9347
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Rosinol, M. Abate, Y. Chang, and L. Carlone, “Kimera: an open-source library for real-time metric-semantic localization and mapping,” in ICRA . IEEE, 2020, pp. 1689–1696
2020
Earlier work this paper cites.
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge, “Room-Across-Room: Multilingual vision-and-language navigation with dense spatiotemporal grounding,” 2020
2020
Earlier work this paper cites.
J. Krantz, E. Wijmans, A. Majumdar, D. Batra, and S. Lee, “Beyond the nav-graph: Vision-and-language navigation in continuous environments,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVIII 16 . Springer, 2020, pp. 104–120
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Qi, Q. Wu, P. Anderson, X. Wang, W. Y. Wang, C. Shen, and A. v. d. Hengel, “Reverie: Remote embodied visual referring expression in real indoor environments,” in CVPR , 2020, pp. 9982–9991
2020
Earlier work this paper cites.
S. Wani, S. Patel, U. Jain, A. Chang, and M. Savva, “MultiON: Benchmarking Semantic Map Memory using Multi-Object Navigation,” NeurIPS , vol. 33, pp. 9700–9712, 2020
2020
Earlier work this paper cites.
D. S. Chaplot, D. P. Gandhi, A. Gupta, and R. R. Salakhutdinov, “Object goal navigation using goal-oriented semantic exploration,” Advances in Neural Information Processing Systems , vol. 33, pp. 4247–4258, 2020
2020
Earlier work this paper cites.
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra, “Improving vision-and-language navigation with image-text pairs from the web,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16 . Springer, 2020, pp. 259–274
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
C. Campos, R. Elvira, J. J. G. Rodríguez, J. M. Montiel, and J. D. Tardós, “Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam,” IEEE Transactions on Robotics , vol. 37, no. 6, pp. 1874–1890, 2021
2021
Earlier work this paper cites.
A. Rosinol, A. Violette, M. Abate, N. Hughes, Y. Chang, J. Shi, A. Gupta, and L. Carlone, “Kimera: From SLAM to spatial perception with 3D dynamic scene graphs,” The International Journal of Robotics Research , vol. 40, no. 12-14, pp. 1510–1546, 2021
2021
Earlier work this paper cites.
E. Ackerman, “Exyn brings level 4 autonomy to drones,” IEEE Spectrum , 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. , “Learning transferable visual models from natural language supervision,” in ICML . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
P. Anderson, A. Shrivastava, J. Truong, A. Majumdar, D. Parikh, D. Batra, and S. Lee, “Sim-to-real transfer for vision-and-language navigation,” in Conference on Robot Learning . PMLR, 2021, pp. 671–681
2021
Cited alongside, same era.
F. Zhu, X. Liang, Y. Zhu, Q. Yu, X. Chang, and X. Liang, “Soon: Scenario oriented object navigation with graph-based exploration,” in CVPR , 2021, pp. 12 689–12 699
2021
Cited alongside, same era.
S. Raychaudhuri, S. Wani, S. Patel, U. Jain, and A. Chang, “Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 4018–4028
2021
Cited alongside, same era.
J. Krantz, A. Gokaslan, D. Batra, S. Lee, and O. Maksymets, “Waypoint models for instruction-guided navigation in continuous environments,” in ICCV , 2021, pp. 15 162–15 171
2021
Cited alongside, same era.
K. Mazur, E. Sucar, and A. J. Davison, “Feature-realistic neural fusion for real-time, open set scene understanding,” in ICRA . IEEE, 2023, pp. 8201–8207
2023
Later among the works it cites.
Z. Wang, X. Li, J. Yang, Y. Liu, and S. Jiang, “Gridmm: Grid memory map for vision-and-language navigation,” in ICCV , 2023, pp. 15 625–15 636
2023
Later among the works it cites.
D. An, Y. Qi, Y. Li, Y. Huang, L. Wang, T. Tan, and J. Shao, “Bevbert: Multimodal map pre-training for language-guided navigation,” in ICCV , 2023, pp. 2737–2748
2023
Later among the works it cites.
C. Huang, O. Mees, A. Zeng, and W. Burgard, “Visual language maps for robot navigation,” in ICRA . IEEE, 2023, pp. 10 608–10 615
2023
Later among the works it cites.
C. Xu, H. T. Nguyen, C. Amato, and L. Wong, “Vision and language navigation in the real world via online visual language mapping,” in 2nd Workshop on Language and Robot Learning: Language as Grounding , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. M. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang, M. Savva, Y. Zhao, and D. Batra, “Habitat-matterport 3D dataset (HM3d): 1000 large-scale 3D environments for embodied AI,” in NeurIPS Datasets and Benchmarks Track (Round 2) , 2021
2021
Cited alongside, same era.
V. Tschernezki, I. Laina, D. Larlus, and A. Vedaldi, “Neural feature fusion fields: 3d distillation of self-supervised 2d image representations,” in 2022 International Conference on 3D Vision (3DV) . IEEE, 2022, pp. 443–453
2022
Cited alongside, same era.
S. Kobayashi, E. Matsumoto, and V. Sitzmann, “Decomposing nerf for editing via feature field distillation,” Advances in Neural Information Processing Systems , vol. 35, pp. 23 311–23 330, 2022
2022
Cited alongside, same era.
N. M. M. Shafiullah, C. Paxton, L. Pinto, S. Chintala, and A. Szlam, “Clip-fields: Weakly supervised semantic fields for robotic memory,” in Conference on Robot Learning Workshops , 2022
2022
Cited alongside, same era.
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev, “Think global, act local: Dual-scale graph transformer for vision-and-language navigation,” in CVPR , 2022, pp. 16 537–16 547
2022
Cited alongside, same era.
G. Georgakis, K. Schmeckpeper, K. Wanchoo, S. Dan, E. Miltsakaki, D. Roth, and K. Daniilidis, “Cross-modal map learning for vision and language navigation,” in CVPR , 2022, pp. 15 460–15 470
2022
Cited alongside, same era.
V. S. Dorbala, G. A. Sigurdsson, J. Thomason, R. Piramuthu, and G. S. Sukhatme, “CLIP-nav: Using CLIP for zero-shot vision-and-language navigation,” 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
K. Yadav, R. Ramrakhya, S. K. Ramakrishnan, T. Gervet, J. Turner, A. Gokaslan, N. Maestre, A. X. Chang, D. Batra, M. Savva, et al. , “Habitat-matterport 3d semantics dataset,” in CVPR , 2023, pp. 4927–4936
2023
Later among the works it cites.
Z. Wu, Y. Wang, J. Ye, and L. Kong, “Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering,” in The 61st Annual Meeting of the Association for Computational Linguistics (09/07/2023-14/07/2023, Toronto, Canada) , 2023
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in ICML . PMLR, 2023, pp. 19 730–19 742
2023
Later among the works it cites.
C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, “Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,” in CVPR , 2023, pp. 7464–7475
2023
Later among the works it cites.
F. Engelmann, F. Manhardt, M. Niemeyer, K. Tateno, M. Pollefeys, and F. Tombari, “OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views,” in International Conference on Learning Representations , 2024
2024
Closest in time.
2024
Closest in time.
J.-C. Shi, M. Wang, H.-B. Duan, and S.-H. Guan, “Language embedded 3D Gaussians for open-vocabulary scene understanding,” in CVPR , 2024, pp. 5333–5343
2024
Closest in time.
K. Yamazaki, T. Hanyu, K. Vo, T. Pham, M. Tran, G. Doretto, A. Nguyen, and N. Le, “Open-fusion: Real-time open-vocabulary 3D mapping and queryable scene representation,” in ICRA . IEEE, 2024, pp. 9411–9417
2024
Closest in time.
Q. Gu, A. Kuwajerwala, S. Morin, K. M. Jatavallabhula, B. Sen, A. Agarwal, C. Rivera, W. Paul, K. Ellis, R. Chellappa, et al. , “Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,” in ICRA . IEEE, 2024, pp. 5021–5028
2024
Closest in time.
S. Koch, N. Vaskevicius, M. Colosi, P. Hermosilla, and T. Ropinski, “Open3DSG: Open-vocabulary 3D scene graphs from point clouds with queryable objects and open-set relationships,” in CVPR , June 2024
2024
Closest in time.
S. Raychaudhuri, T. Campari, U. Jain, M. Savva, and A. X. Chang, “Mopa: Modular object navigation with pointgoal agents,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 5763–5773
2024
Closest in time.
M. Chang, T. Gervet, M. Khanna, S. Yenamandra, D. Shah, S. Y. Min, K. Shah, C. Paxton, S. Gupta, D. Batra, et al. , “Goat: Go to any thing,” Robotics: Science and Systems (RSS) , 2024
2024
Closest in time.
N. Yokoyama, S. Ha, D. Batra, J. Wang, and B. Bucher, “Vlfm: Vision-language frontier maps for zero-shot semantic navigation,” ICRA , 2024
2024
Closest in time.
R. Liu, W. Wang, and Y. Yang, “Volumetric environment representation for vision-language navigation,” in CVPR , 2024, pp. 16 317–16 328
2024
Closest in time.
G. Zhou, Y. Hong, and Q. Wu, “Navgpt: Explicit reasoning in vision-and-language navigation with large language models,” vol. 38, no. 7, pp. 7641–7649, 2024
2024
Closest in time.
Y. Long, X. Li, W. Cai, and H. Dong, “Discuss before moving: Visual language navigation via multi-expert discussions,” pp. 17 380–17 387, 2024
2024
Closest in time.
J. Chen, B. Lin, R. Xu, Z. Chai, X. Liang, and K.-Y. K. Wong, “Mapgpt: Map-guided prompting with adaptive path planning for vision-and-language navigation,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , 2024
2024
Closest in time.
2024
Closest in time.
M. Khanna, R. Ramrakhya, G. Chhablani, S. Yenamandra, T. Gervet, M. Chang, Z. Kira, D. S. Chaplot, D. Batra, and R. Mottaghi, “Goat-bench: A benchmark for multi-modal lifelong navigation,” in CVPR , 2024, pp. 16 373–16 383
2024
Closest in time.
Y. Zhang, X. Huang, J. Ma, Z. Li, Z. Luo, Y. Xie, Y. Qin, T. Luo, Y. Li, S. Liu, et al. , “Recognize anything: A strong image tagging model,” in CVPR , 2024, pp. 1724–1732
2024
Closest in time.
Y. Long, W. Cai, H. Wang, G. Zhan, and H. Dong, “Instructnav: Zero-shot system for generic instruction navigation in unexplored environment,” CoRL , 2024
2024
Closest in time.
J. Zhang, R. Dong, and K. Ma, “Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip,” in ICCV , 2023, pp. 2048–2059
2059
Closest in time.