Fetching the paper…
Reading the bibliography…
For robots to perform a wide variety of tasks, they require a 3D representation of the world that is semantically rich, yet compact and efficient for task-driven perception and planning.
R. A. Newcombe, S. Izadi, O. Hilliges, D. Molyneaux, D. Kim, A. J. Davison, P. Kohi, J. Shotton, S. Hodges, and A. Fitzgibbon, “Kinectfusion: Real-time dense surface mapping and tracking,” in IEEE international symposium on mixed and augmented reality . IEEE, 2011, pp. 127–136
2011
Earlier work this paper cites.
M. Fisher, M. Savva, and P. Hanrahan, “Characterizing structural relationships in scenes using graph kernels,” ACM Trans. Graph. , vol. 30, no. 4, p. 34, 2011
2011
Earlier work this paper cites.
T. Whelan, S. Leutenegger, R. Salas-Moreno, B. Glocker, and A. Davison, “Elasticfusion: Dense slam without a pose graph,” in Robotics Science and Systems , 2015
2015
Earlier work this paper cites.
J. L. Schonberger and J.-M. Frahm, “Structure-from-motion revisited,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4104–4113
2016
Earlier work this paper cites.
J. L. Schönberger, E. Zheng, J.-M. Frahm, and M. Pollefeys, “Pixelwise view selection for unstructured multi-view stereo,” in European Conference on Computer Vision . Springer, 2016, pp. 501–518
2016
Earlier work this paper cites.
J. McCormac, A. Handa, A. Davison, and S. Leutenegger, “Semanticfusion: Dense 3d semantic mapping with convolutional neural networks,” in IEEE International Conference on Robotics and automation (ICRA) . IEEE, 2017, pp. 4628–4635
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Runz, M. Buffier, and L. Agapito, “Maskfusion: Real-time recognition, tracking and reconstruction of multiple moving objects,” in IEEE International Symposium on Mixed and Augmented Reality (ISMAR) . IEEE, 2018, pp. 10–20
2018
Earlier work this paper cites.
J. McCormac, R. Clark, M. Bloesch, A. Davison, and S. Leutenegger, “Fusion++: Volumetric object-level slam,” in international conference on 3D vision (3DV) . IEEE, 2018, pp. 32–41
2018
Earlier work this paper cites.
G. Narita, T. Seno, T. Ishikawa, and Y. Kaji, “Panopticfusion: Online volumetric semantic mapping at the level of stuff and things,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2019, pp. 4205–4212
2019
Earlier work this paper cites.
P. Gay, J. Stuart, and A. Del Bue, “Visual graphs from motion (vgfm): Scene understanding with object geometry reasoning,” in Asian Conference on Computer Vision . Springer, 2019
2019
Earlier work this paper cites.
I. Armeni, Z.-Y. He, J. Gwak, A. R. Zamir, M. Fischer, J. Malik, and S. Savarese, “3d scene graph: A structure for unified semantics, 3d space, and camera,” in Proceedings of International Conference on Computer Vision , October 2019
2019
Earlier work this paper cites.
U.-H. Kim, J.-M. Park, T.-J. Song, and J.-H. Kim, “3-d scene graph: A sparse and semantic representation of physical environments for intelligent agents,” IEEE transactions on cybernetics , vol. 50, no. 12, pp. 4921–4933, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Labbé and F. Michaud, “Rtab-map as an open-source lidar and visual simultaneous localization and mapping library for large-scale and long-term online operation,” Journal of Field Robotics , vol. 36, no. 2, pp. 416–446, 2019
2019
Earlier work this paper cites.
J. Wald, H. Dhamo, N. Navab, and F. Tombari, “Learning 3d semantic scene graphs from 3d indoor reconstructions,” in Proceedings of Computer Vision and Pattern Recognition , 2020
2020
Earlier work this paper cites.
Y. Du, S. Li, and I. Mordatch, “Compositional visual generation with energy based models,” in Neural Information Processing Systems , 2020
2020
Earlier work this paper cites.
E. Sucar, S. Liu, J. Ortiz, and A. J. Davison, “iMAP: Implicit mapping and positioning in real-time,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 6229–6238
2021
Earlier work this paper cites.
K. Li, D. DeTone, Y. F. S. Chen, M. Vo, I. Reid, H. Rezatofighi, C. Sweeney, J. Straub, and R. Newcombe, “ODAM: Object detection, association, and mapping using posed rgb video,” in Proceedings of International Conference on Computer Vision , 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning . PMLR, 2021
2021
Earlier work this paper cites.
M. Shridhar, L. Manuelli, and D. Fox, “CLIPort: What and where pathways for robotic manipulation,” in Conference on Robot Learning , vol. 164. PMLR, 2021, pp. 894–906
2021
Earlier work this paper cites.
A. Rosinol, A. Violette, M. Abate, N. Hughes, Y. Chang, J. Shi, A. Gupta, and L. Carlone, “Kimera: From slam to spatial perception with 3d dynamic scene graphs,” The International Journal of Robotics Research , vol. 40, no. 12-14, pp. 1510–1546, 2021
2021
Earlier work this paper cites.
S.-C. Wu, J. Wald, K. Tateno, N. Navab, and F. Tombari, “Scenegraphfusion: Incremental 3d scene graph prediction from rgb-d sequences,” in Proceedings of Computer Vision and Pattern Recognition , 2021
2021
Earlier work this paper cites.
Z. Zhu, S. Peng, V. Larsson, W. Xu, H. Bao, Z. Cui, M. R. Oswald, and M. Pollefeys, “Nice-slam: Neural implicit scalable encoding for slam,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 12 786–12 796
2022
Earlier work this paper cites.
J. Qian, V. Chatrath, J. Yang, J. Servos, A. P. Schoellig, and S. L. Waslander, “POCD: probabilistic object-level change detection and volumetric mapping in semi-static scenes,” in Robotics Science and Systems , K. Hauser, D. A. Shell, and S. Huang, Eds., 2022
2022
Cited alongside, same era.
M. Zins, G. Simon, and M.-O. Berger, “OA-SLAM: Leveraging objects for camera relocalization in visual slam,” in IEEE International Symposium on Mixed and Augmented Reality (ISMAR) . IEEE, 2022, pp. 720–728
2022
Cited alongside, same era.
V. Tschernezki, I. Laina, D. Larlus, and A. Vedaldi, “Neural feature fusion fields: 3d distillation of self-supervised 2d image representations,” in International Conference on 3D Vision (3DV) . IEEE, 2022
2022
Cited alongside, same era.
S. Kobayashi, E. Matsumoto, and V. Sitzmann, “Decomposing nerf for editing via feature field distillation,” Neural Information Processing Systems , vol. 35, pp. 23 311–23 330, 2022
2022
Cited alongside, same era.
N. M. M. Shafiullah, C. Paxton, L. Pinto, S. Chintala, and A. Szlam, “Clip-fields: Weakly supervised semantic fields for robotic memory,” in Robotics: Science and Systems , K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, Eds., 2023
2023
Closest in time.
2023
Closest in time.
J. Kerr, C. M. Kim, K. Goldberg, A. Kanazawa, and M. Tancik, “LERF: Language embedded radiance fields,” in International Conference on Computer Vision (ICCV) , 2023
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 10 684–10 695
2022
Cited alongside, same era.
Y. Hong, Y. Du, C. Lin, J. Tenenbaum, and C. Gan, “3d concept grounding on neural fields,” Neural Information Processing Systems , 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
C. Agia, K. M. Jatavallabhula, M. Khodeir, O. Miksik, V. Vineet, M. Mukadam, L. Paull, and F. Shkurti, “Taskography: Evaluating robot task planning over large 3d scene graphs,” in International Conference on Robot Learning . PMLR, 2022
2022
Cited alongside, same era.
T. Lüddecke and A. Ecker, “Image segmentation using text and image prompts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 7086–7096
2022
Cited alongside, same era.
B. Li, K. Q. Weinberger, S. Belongie, V. Koltun, and R. Ranftl, “Language-driven semantic segmentation,” in International Conference on Learning Representations , 2022
2022
Cited alongside, same era.
G. Ghiasi, X. Gu, Y. Cui, and T.-Y. Lin, “Scaling open-vocabulary image segmentation with image-level labels,” in European Conference on Computer Vision . Springer, 2022, pp. 540–557
2022
Cited alongside, same era.
W. Shen, G. Yang, A. Yu, J. Wong, L. P. Kaelbling, and P. Isola, “Distilled feature fields enable few-shot manipulation,” in International Conference on Robot Learning , 2023. [Online]. Available: https://openreview.net/forum?id=Rb0nGIt_kh5
2023
Closest in time.
F. Engelmann, F. Manhardt, M. Niemeyer, K. Tateno, M. Pollefeys, and F. Tombari, “Open-set 3d scene segmentation with rendered novel views,” 2023
2023
Closest in time.
K. Mazur, E. Sucar, and A. J. Davison, “Feature-realistic neural fusion for real-time, open set scene understanding,” in IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023
2023
Closest in time.
OpenAI, “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023
2023
Closest in time.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al. , “Segment anything,” Proceedings of International Conference on Computer Vision , 2023
2023
Closest in time.
2023
Closest in time.
Y. Hong, C. Lin, Y. Du, Z. Chen, J. B. Tenenbaum, and C. Gan, “3d concept learning and reasoning from multi-view images,” in Proceedings of Computer Vision and Pattern Recognition , 2023
2023
Closest in time.
Y. Hong, H. Zhen, P. Chen, S. Zheng, Y. Du, Z. Chen, and C. Gan, “3d-llm: Injecting the 3d world into large language models,” Neural Information Processing Systems , 2023
2023
Closest in time.
S. Sharma, A. Rashid, C. M. Kim, J. Kerr, L. Y. Chen, A. Kanazawa, and K. Goldberg, “Language embedded radiance fields for zero-shot task-oriented grasping,” in International Conference on Robot Learning , 2023. [Online]. Available: https://openreview.net/forum?id=k-Fg8JDQmc
2023
Closest in time.
D. Shah, B. Osiński, S. Levine et al. , “Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,” in International Conference on Robot Learning . PMLR, 2023
2023
Closest in time.
K. Rana, J. Abou-Chakra, S. Garg, J. Haviland, I. Reid, and N. Suenderhauf, “Sayplan: Grounding large language models using 3d scene graphs for scalable task planning,” in International Conference on Robot Learning , 2023. [Online]. Available: https://openreview.net/forum?id=wMpOMO0Ss7a
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
S. Levine and D. Shah, “Learning robotic navigation from experience: principles, methods and recent results,” Philosophical Transactions of the Royal Society B , vol. 378, no. 1869, p. 20210447, 2023
2023
Closest in time.
2023
Closest in time.
S. Lu, H. Chang, E. P. Jing, A. Boularias, and K. Bekris, “OVIR-3d: Open-vocabulary 3d instance retrieval without training on 3d data,” in International Conference on Robot Learning , 2023. [Online]. Available: https://openreview.net/forum?id=gVBvtRqU1_
2023
Closest in time.
H. Chang, K. Boyalakuntla, S. Lu, S. Cai, E. P. Jing, S. Keskar, S. Geng, A. Abbas, L. Zhou, K. Bekris, and A. Boularious, “Context-aware entity grounding with open-vocabulary 3d scene graphs,” in International Conference on Robot Learning , 2023. [Online]. Available: https://openreview.net/forum?id=cjEI5qXoT0
2023
Closest in time.
Z. Sun, S. Shen, S. Cao, H. Liu, C. Li, Y. Shen, C. Gan, L.-Y. Gui, Y.-X. Wang, Y. Yang, K. Keutzer, and T. Darrell, “Aligning large multimodal models with factually augmented rlhf,” 2023
2023
Closest in time.