Fetching the paper…
Reading the bibliography…
In Embodied Question Answering (EQA), agents must explore and develop a semantic understanding of an unseen environment to answer a situated question with confidence.
3d scene graph: A structure for unified semantics, 3d space, and camera, 2019
I. Armeni, Z.-Y. He, J. Gwak, A. R. Zamir, M. Fischer, J. Malik, and S. Savarese · 1910
Earlier work this paper cites.
Embodied question answering
A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra · 2018
Earlier work this paper cites.
Iqa: Visual question answering in interactive environments
D. Gordon, A. Kembhavi, M. Rastegari, J. Redmon, D. Fox, and A. Farhadi · 2018
Earlier work this paper cites.
3-d scene graph: A sparse and semantic representation of physical environments for intelligent agents
U.-H. Kim, J.-M. Park, T.-j. Song, and J.-H. Kim · 2019
Earlier work this paper cites.
Embodied question answering in photorealistic environments with point cloud perception
E. Wijmans, S. Datta, O. Maksymets, A. Das, G. Gkioxari, S. Lee, I. Essa, D. Parikh, and D. Batra · 2019
Earlier work this paper cites.
Learning 3d semantic scene graphs from 3d indoor reconstructions
J. Wald, H. Dhamo, N. Navab, and F. Tombari · 2020
Earlier work this paper cites.
Kimera: from slam to spatial perception with 3d dynamic scene graphs, 2021
A. Rosinol, A. Violette, M. Abate, N. Hughes, Y. Chang, J. Shi, A. Gupta, and L. Carlone · 2021
Earlier work this paper cites.
Scenegraphfusion: Incremental 3d scene graph prediction from rgb-d sequences
S.-C. Wu, J. Wald, K. Tateno, N. Navab, and F. Tombari · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Habitat 2.0: Training home assistants to rearrange their habitat
A. Szot, A. Clegg, E. Undersander, E. Wijmans, Y. Zhao, J. Turner, N. Maestre, M. Mukadam, D. Chaplot, O. Maksymets, A. Gokaslan, V. Vondrus, S. Dharur, F. Meier, W. Galuba, A. Chang, Z. Kira, V. Koltun, J. Malik, M. Savva, and D. Batra · 2021
Earlier work this paper cites.
Habitat-matterport 3d dataset (HM3d): 1000 large-scale 3d environments for embodied AI
S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. M. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang, M. Savva, Y. Zhao, and D. Batra · 2021
Earlier work this paper cites.
Hydra: A real-time spatial perception system for 3d scene graph construction and optimization, 2022
N. Hughes, Y. Chang, and L. Carlone · 2022
Earlier work this paper cites.
Habitat-matterport 3d semantics dataset
K. Yadav, R. Ramrakhya, S. K. Ramakrishnan, T. Gervet, J. Turner, A. Gokaslan, N. Maestre, A. X. Chang, D. Batra, M. Savva, et al · 2022
Earlier work this paper cites.
Taskography: Evaluating robot task planning over large 3d scene graphs
C. Agia, K. M. Jatavallabhula, M. Khodeir, O. Miksik, V. Vineet, M. Mukadam, L. Paull, and F. Shkurti · 2022
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances, 2022
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, M. Yan, and A. Zeng · 2022
Earlier work this paper cites.
Clip-fields: Weakly supervised semantic fields for robotic memory
N. M. M. Shafiullah, C. Paxton, L. Pinto, S. Chintala, and A. Szlam · 2022
Earlier work this paper cites.
Neural feature fusion fields: 3d distillation of self-supervised 2d image representations
V. Tschernezki, I. Laina, D. Larlus, and A. Vedaldi · 2022
Earlier work this paper cites.
Core challenges in embodied vision-language planning
J. Francis, N. Kitamura, F. Labelle, X. Lu, I. Navarro, and J. Oh · 2022
Earlier work this paper cites.
Detecting twenty-thousand classes using image-level supervision, 2022
X. Zhou, R. Girdhar, A. Joulin, P. Krähenbühl, and I. Misra · 2022
Earlier work this paper cites.
Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning, 2023
Q. Gu, A. Kuwajerwala, S. Morin, K. M. Jatavallabhula, B. Sen, A. Agarwal, C. Rivera, W. Paul, K. Ellis, R. Chellappa, C. Gan, C. M. de Melo, J. B. Tenenbaum, A. Torralba, F. Shkurti, and L. Paull · 2023
Earlier work this paper cites.
K. Rana, J. Haviland, S. Garg, J. Abou-Chakra, I. Reid, and N. Suenderhauf · 2023
Cited alongside, same era.
Context-aware entity grounding with open-vocabulary 3d scene graphs, 2023
H. Chang, K. Boyalakuntla, S. Lu, S. Cai, E. Jing, S. Keskar, S. Geng, A. Abbas, L. Zhou, K. Bekris, and A. Boularias · 2023
Cited alongside, same era.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Cited alongside, same era.
Open-vocabulary queryable scene representations for real world planning
B. Chen, F. Xia, B. Ichter, K. Rao, K. Gopalakrishnan, M. S. Ryoo, A. Stone, and D. Kappler · 2023
Cited alongside, same era.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2023
Remembr: Building and reasoning over long-horizon spatio-temporal memory for robot navigation, 2024
A. Anwar, J. Welsh, J. Biswas, S. Pouya, and Y. Chang · 2024
Closest in time.
Embodied-rag: General non-parametric embodied memory for retrieval and generation, 2024
Q. Xie, S. Y. Min, T. Zhang, K. Xu, A. Bajaj, R. Salakhutdinov, M. Johnson-Roberson, and Y. Bisk · 2024
Closest in time.
Mobility vla: Multimodal instruction navigation with long-context vlms and topological graphs
Z. Xu, H.-T. L. Chiang, Z. Fu, M. G. Jacob, T. Zhang, T.-W. E. Lee, W. Yu, C. Schenck, D. Rendleman, D. Shah, et al · 2024
Closest in time.
Openeqa: Embodied question answering in the era of foundation models
A. Majumdar, A. Ajay, X. Zhang, P. Putta, S. Yenamandra, M. Henaff, S. Silwal, P. Mcvay, O. Maksymets, S. Arnaud, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
C. H. Song, J. Wu, C. Washington, B. M. Sadler, W.-L. Chao, and Y. Su · 2023
Cited alongside, same era.
Pla: Language-driven open-vocabulary 3d scene understanding
R. Ding, J. Yang, C. Xue, W. Zhang, S. Bai, and X. Qi · 2023
Cited alongside, same era.
Openscene: 3d scene understanding with open vocabularies
S. Peng, K. Genova, C. Jiang, A. Tagliasacchi, M. Pollefeys, T. Funkhouser, et al · 2023
Cited alongside, same era.
Visual language maps for robot navigation
C. Huang, O. Mees, A. Zeng, and W. Burgard · 2023
Cited alongside, same era.
Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip
J. Zhang, R. Dong, and K. Ma · 2023
Cited alongside, same era.
Conceptfusion: Open-set multimodal 3d mapping
K. M. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, A. Maalouf, S. Li, G. Iyer, S. Saryazdi, N. Keetha, et al · 2023
Cited alongside, same era.
Lerf: Language embedded radiance fields
J. Kerr, C. M. Kim, K. Goldberg, A. Kanazawa, and M. Tancik · 2023
Cited alongside, same era.
R. Xu, Z. Huang, T. Wang, Y. Chen, J. Pang, and D. Lin · 2024
Closest in time.
Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation
A. Werby, C. Huang, M. Büchner, A. Valada, and W. Burgard · 2024
Closest in time.
Openeqa: Embodied question answering in the era of foundation models
A. Majumdar, A. Ajay, X. Zhang, P. Putta, S. Yenamandra, M. Henaff, S. Silwal, P. Mcvay, O. Maksymets, S. Arnaud, K. Yadav, Q. Li, B. Newman, M. Sharma, V. Berges, S. Zhang, P. Agrawal, Y. Bisk, D. Batra, M. Kalakrishnan, F. Meier, C. Paxton, S. Sax, and A. Rajeswaran · 2024
Closest in time.
Collaborative dynamic 3d scene graphs for automated driving
E. Greve, M. Büchner, N. Vödisch, W. Burgard, and A. Valada · 2024
Closest in time.
Llm-enhanced scene graph learning for household rearrangement, 2024
W. Li, Z. Yu, Q. She, Z. Yu, Y. Lan, C. Zhu, R. Hu, and K. Xu · 2024
Closest in time.
Clio: Real-time task-driven open-set 3d scene graphs
D. Maggio, Y. Chang, N. Hughes, M. Trang, D. Griffith, C. Dougherty, E. Cristofalo, L. Schmid, and L. Carlone · 2024
Closest in time.
Language-grounded dynamic scene graphs for interactive object search with mobile manipulation
D. Honerkamp, M. Büchner, F. Despinoy, T. Welschehold, and A. Valada · 2024
Closest in time.
Saynav: Grounding large language models for dynamic planning to navigation in new environments, 2024
A. Rajvanshi, K. Sikka, X. Lin, B. Lee, H.-P. Chiu, and A. Velasquez · 2024
Closest in time.
Dynamem: Online dynamic spatio-semantic memory for open world mobile manipulation
P. Liu, Z. Guo, M. Warke, S. Chintala, C. Paxton, N. M. M. Shafiullah, and L. Pinto · 2024
Closest in time.
Navid: Video-based vlm plans the next step for vision-and-language navigation, 2024
J. Zhang, K. Wang, R. Xu, G. Zhou, Y. Hong, X. Fang, Q. Wu, Z. Zhang, and H. Wang · 2024
Closest in time.
Vlfm: Vision-language frontier maps for zero-shot semantic navigation
N. Yokoyama, S. Ha, D. Batra, J. Wang, and B. Bucher · 2024
Closest in time.
Pivot: Iterative visual prompting elicits actionable knowledge for vlms
S. Nasiriany, F. Xia, W. Yu, T. Xiao, J. Liang, I. Dasgupta, A. Xie, D. Driess, A. Wahid, Z. Xu, et al · 2024
Closest in time.
Open-vocabulary mobile manipulation in unseen dynamic environments with 3d semantic maps, 2024
D. Qiu, W. Ma, Z. Pan, H. Xiong, and J. Liang · 2024
Closest in time.
Grounding embodied question-answering with state summaries from existing robot modules
S. Bustamante Gomez, M. W. Knauer, T. Jeremias, S. Schneyer, B. Weber, and F. Stulp · 2024
Closest in time.
Open scene graphs for open world object-goal navigation, 2024
J. Loo, Z. Wu, and D. Hsu · 2024
Closest in time.
Prismatic vlms: Investigating the design space of visually-conditioned language models
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh · 2024
Closest in time.