Fetching the paper…
Reading the bibliography…
How can we build robots for open-world semantic navigation tasks, like searching for target objects in novel scenes? While foundation models have the rich knowledge and generalisation needed for these tasks, a suitable scene representation is needed to connect them into a complete robot system.
A fast marching level set method for monotonically advancing fronts
J. A. Sethian · 1996
Earlier work this paper cites.
A frontier-based approach for autonomous exploration
B. Yamauchi · 1997
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
C. Diuk, A. Cohen, and M. L. Littman · 2008
Earlier work this paper cites.
Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age
C. Cadena, L. Carlone, H. Carrillo, Y. Latif, D. Scaramuzza, J. Neira, I. D. Reid, and J. J. Leonard · 2016
Earlier work this paper cites.
Gibson env: real-world perception for embodied agents
F. Xia, A. R. Zamir, Z.-Y. He, A. Sax, J. Malik, and S. Savarese · 2018
Earlier work this paper cites.
3d scene graph: A structure for unified semantics, 3d space, and camera
I. Armeni, Z.-Y. He, J. Gwak, A. R. Zamir, M. Fischer, J. Malik, and S. Savarese · 2019
Earlier work this paper cites.
Multi-object search using object-oriented pomdps
A. Wandzel, Y. Oh, M. Fishman, N. Kumar, L. L. Wong, and S. Tellex · 2019
Earlier work this paper cites.
Volumetric Instance-Aware Semantic Mapping and 3D Object Discovery
M. Grinvald, F. Furrer, T. Novkovic, J. J. Chung, C. Cadena, R. Siegwart, and J. Nieto · 2019
Earlier work this paper cites.
LVIS: A dataset for large vocabulary instance segmentation
A. Gupta, P. Dollar, and R. Girshick · 2019
Earlier work this paper cites.
Visual representations for semantic target driven navigation
A. Mousavian, A. Toshev, M. Fišer, J. Košecká, A. Wahid, and J. Davidson · 2019
Earlier work this paper cites.
Multi-resolution pomdp planning for multi-object search in 3d
K. Zheng, Y. Sung, G. D. Konidaris, and S. Tellex · 2020
Earlier work this paper cites.
Object goal navigation using goal-oriented semantic exploration
D. S. Chaplot, D. Gandhi, A. Gupta, and R. Salakhutdinov · 2020
Earlier work this paper cites.
Kimera: From slam to spatial perception with 3d dynamic scene graphs
A. Rosinol, A. Violette, M. Abate, N. Hughes, Y. Chang, J. Shi, A. Gupta, and L. Carlone · 2021
Earlier work this paper cites.
Habitat-matterport 3d dataset (HM3d): 1000 large-scale 3d environments for embodied AI
S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. M. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang, M. Savva, Y. Zhao, and D. Batra · 2021
Earlier work this paper cites.
Habitat 2.0: Training home assistants to rearrange their habitat
A. Szot, A. Clegg, E. Undersander, E. Wijmans, Y. Zhao, J. Turner, N. Maestre, M. Mukadam, D. Chaplot, O. Maksymets, A. Gokaslan, V. Vondrus, S. Dharur, F. Meier, W. Galuba, A. Chang, Z. Kira, V. Koltun, J. Malik, M. Savva, and D. Batra · 2021
Earlier work this paper cites.
Learning object-conditioned exploration using distributed soft actor critic
A. Wahid, A. Stone, K. Chen, B. Ichter, and A. Toshev · 2021
Earlier work this paper cites.
Thda: Treasure hunt data augmentation for semantic navigation
O. Maksymets, V. Cartillier, A. Gokaslan, E. Wijmans, W. Galuba, S. Lee, and D. Batra · 2021
Earlier work this paper cites.
Auxiliary tasks and exploration enable objectgoal navigation
J. Ye, D. Batra, A. Das, and E. Wijmans · 2021
Cited alongside, same era.
Habitat challenge 2022
K. Yadav, S. K. Ramakrishnan, J. Turner, A. Gokaslan, O. Maksymets, R. Jain, R. Ramrakhya, A. X. Chang, A. Clegg, M. Savva, E. Undersander, D. S. Chaplot, and D. Batra · 2022
Cited alongside, same era.
Simple open-vocabulary object detection
M. Minderer, A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shen, X. Wang, X. Zhai, T. Kipf, and N. Houlsby · 2022
Cited alongside, same era.
Poni: Potential functions for objectgoal navigation with interaction-free learning
S. K. Ramakrishnan, D. S. Chaplot, Z. Al-Halah, J. Malik, and K. Grauman · 2022
Cited alongside, same era.
Zson: Zero-shot object-goal navigation using multimodal goal embeddings
A. Majumdar, G. Aggarwal, B. Devnani, J. Hoffman, and D. Batra · 2022
Cited alongside, same era.
A system for generalized 3d multi-object search
K. Zheng, A. Paul, and S. Tellex · 2023
Later among the works it cites.
How To Not Train Your Dragon: Training-free Embodied Object Goal Navigation with Semantic Frontiers
J. Chen, G. Li, S. Kumar, B. Ghanem, and F. Yu · 2023
Later among the works it cites.
Toward general-purpose robots via foundation models: A survey and meta-analysis, 2023
Y. Hu, Q. Xie, V. Jain, J. Francis, J. Patrikar, N. Keetha, S. Kim, Y. Xie, T. Zhang, S. Zhao, Y. Q. Chong, C. Wang, K. Sycara, M. Johnson-Roberson, D. Batra, X. Wang, S. Scherer, Z. Kira, F. Xia, and Y. Bisk · 2023
Later among the works it cites.
Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill, 2023
W. Cai, S. Huang, G. Cheng, Y. Long, P. Gao, C. Sun, and H. Dong · 2023
Later among the works it cites.
Orca 2: Teaching small language models how to reason, 2023
A. Mitra, L. D. Corro, S. Mahajan, A. Codas, C. Simoes, S. Agarwal, X. Chen, A. Razdaibiedina, E. Jones, K. Aggarwal, H. Palangi, G. Zheng, C. Rosset, H. Khanpour, and A. Awadallah · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Huang and K. C.-C. Chang · 2022
Cited alongside, same era.
Reasoning with scene graphs for robot planning under partial observability
S. Amiri, K. Chandan, and S. Zhang · 2022
Cited alongside, same era.
Hierarchical representations and explicit memory: Learning effective navigation policies on 3d scene graphs using graph neural networks
Z. Ravichandran, L. Peng, N. Hughes, J. D. Griffith, and L. Carlone · 2022
Cited alongside, same era.
Navgpt: Explicit reasoning in vision-and-language navigation with large language models, 2023
G. Zhou, Y. Hong, and Q. Wu · 2023
Cited alongside, same era.
L3mvn: Leveraging large language models for visual target navigation
B. Yu, H. Kasaei, and M. Cao · 2023
Cited alongside, same era.
Navigation with large language models: Semantic guesswork as a heuristic for planning
D. Shah, M. R. Equi, B. Osiński, F. Xia, B. Ichter, and S. Levine · 2023
Cited alongside, same era.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick · 2023
Cited alongside, same era.
Robots that ask for help: Uncertainty alignment for large language model planners
A. Z. Ren, A. Dixit, A. Bodrova, S. Singh, S. Tu, N. Brown, P. Xu, L. Takayama, F. Xia, J. Varley, Z. Xu, D. Sadigh, A. Zeng, and A. Majumdar · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Later among the works it cites.
Vlfm: Vision-language frontier maps for zero-shot semantic navigation
N. Yokoyama, S. Ha, D. Batra, J. Wang, and B. Bucher · 2024
Closest in time.
Mapgpt: Map-guided prompting for unified vision-and-language navigation, 2024
J. Chen, B. Lin, R. Xu, Z. Chai, X. Liang, and K.-Y. K. Wong · 2024
Closest in time.
Step-back prompting enables reasoning via abstraction in large language models
H. S. Zheng, S. Mishra, X. Chen, H.-T. Cheng, E. H. Chi, Q. V. Le, and D. Zhou · 2024
Closest in time.
Can an embodied agent find your “cat-shaped mug”? llm-based zero-shot object navigation
V. S. Dorbala, J. F. Mullen, and D. Manocha · 2024
Closest in time.
Language-grounded dynamic scene graphs for interactive object search with mobile manipulation
D. Honerkamp, M. Büchner, F. Despinoy, T. Welschehold, and A. Valada · 2024
Closest in time.
Task and motion planning in hierarchical 3d scene graphs, 2024
A. Ray, C. Bradley, L. Carlone, and N. Roy · 2024
Closest in time.
Indoor and outdoor 3d scene graph generation via language-enabled spatial ontologies
J. Strader, N. Hughes, W. Chen, A. Speranzon, and L. Carlone · 2024
Closest in time.
Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation
A. Werby, C. Huang, M. Büchner, A. Valada, and W. Burgard · 2024
Closest in time.
Explore until confident: Efficient exploration for embodied question answering, 2024
A. Z. Ren, J. Clark, A. Dixit, M. Itkina, A. Majumdar, and D. Sadigh · 2024
Closest in time.