Fetching the paper…
Reading the bibliography…
Open-vocabulary 3D scene understanding (OV-3D) aims to localize and classify novel objects beyond the closed set of object classes.
The Replica dataset: A digital replica of indoor spaces
Straub, J.; Whelan, T.; Ma, L.; Chen, Y.; Wijmans, E.; Green, S.; Engel, J. J.; Mur-Artal, R.; Ren, C.; Verma, S.; et al. 2019 · 1906
Earlier work this paper cites.
A density-based algorithm for discovering clusters in large spatial databases with noise
Ester, M.; Kriegel, H.-P.; Sander, J.; and Xu, X. 1996 · 1996
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
Everingham, M.; Eslami, S. A.; Van Gool, L.; Williams, C. K.; Winn, J.; and Zisserman, A. 2015 · 2015
Earlier work this paper cites.
3d semantic parsing of large-scale indoor spaces
Armeni, I.; Sener, O.; Zamir, A. R.; Jiang, H.; Brilakis, I.; Fischer, M.; and Savarese, S. 2016 · 2016
Earlier work this paper cites.
End to end learning for self-driving cars
Bojarski, M.; Del Testa, D.; Dworakowski, D.; Firner, B.; Flepp, B.; Goyal, P.; Jackel, L. D.; Monfort, M.; Muller, U.; Zhang, J.; et al. 2016 · 2016
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; and Schiele, B. 2016 · 2016
Earlier work this paper cites.
Matterport3D: Learning from RGB-D Data in Indoor Environments
Chang, A.; Dai, A.; Funkhouser, T.; Halber, M.; Niessner, M.; Savva, M.; Song, S.; Zeng, A.; and Zhang, Y. 2017 · 2017
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Dai, A.; Chang, A. X.; Savva, M.; Halber, M.; Funkhouser, T.; and Nießner, M. 2017 · 2017
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
Speer, R.; Chin, J.; and Havasi, C. 2017 · 2017
Earlier work this paper cites.
3d semantic segmentation with submanifold sparse convolutional networks
Graham, B.; Engelcke, M.; and Van Der Maaten, L. 2018 · 2018
Earlier work this paper cites.
Learning synergies between pushing and grasping with self-supervised deep reinforcement learning
Zeng, A.; Song, S.; Welker, S.; Lee, J.; Rodriguez, A.; and Funkhouser, T. 2018 · 2018
Earlier work this paper cites.
4d spatio-temporal convnets: Minkowski convolutional neural networks
Choy, C.; Gwak, J.; and Savarese, S. 2019 · 2019
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
Gupta, A.; Dollar, P.; and Girshick, R. 2019 · 2019
Earlier work this paper cites.
Semantic understanding of scenes through the ade20k dataset
Zhou, B.; Zhao, H.; Puig, X.; Xiao, T.; Fidler, S.; Barriuso, A.; and Torralba, A. 2019 · 2019
Earlier work this paper cites.
Scanrefer: 3d object localization in rgb-d scans using natural language
Chen, D. Z.; Chang, A. X.; and Nießner, M. 2020 · 2020
Earlier work this paper cites.
ARKitScenes - A Diverse Real-World Dataset for 3D Indoor Scene Understanding Using Mobile RGB-D Data
Baruch, G.; Chen, Z.; Dehghan, A.; Dimry, T.; Feigin, Y.; Fu, P.; Gebauer, T.; Joffe, B.; Kurz, D.; Schwartz, A.; and Shulman, E. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Scaling open-vocabulary image segmentation with image-level labels
Ghiasi, G.; Gu, X.; Cui, Y.; and Lin, T.-Y. 2022 · 2022
Cited alongside, same era.
Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
Gu, X.; Lin, T.-Y.; Kuo, W.; and Cui, Y. 2022 · 2022
Cited alongside, same era.
Language-Grounded Indoor 3D Semantic Segmentation in the Wild
Rozenberszki, D.; Litany, O.; Dai, A.; and Dai, A. 2022 · 2022
Cited alongside, same era.
Regionclip: Region-based language-image pretraining
Zhong, Y.; Yang, J.; Zhang, P.; Li, C.; Codella, N.; Li, L. H.; Zhou, L.; Dai, X.; Yuan, L.; Li, Y.; et al. 2022 · 2022
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Scannet++: A high-fidelity dataset of 3d indoor scenes
Yeshwanth, C.; Liu, Y.-C.; Nießner, M.; and Dai, A. 2023 · 2023
Later among the works it cites.
The devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understanding
Bianchi, L.; Carrara, F.; Messina, N.; Gennaro, C.; and Falchi, F. 2024 · 2024
Closest in time.
Scenefun3d: Fine-grained functionality and affordance understanding in 3d scenes
Delitzas, A.; Takmaz, A.; Tombari, F.; Sumner, R.; Pollefeys, M.; and Engelmann, F. 2024 · 2024
Closest in time.
Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation
Huang, Z.; Wu, X.; Chen, X.; Zhao, H.; Zhu, L.; and Lasenby, J. 2024 · 2024
Closest in time.
Segment and recognize anything at any granularity
Li, F.; Zhang, H.; Sun, P.; Zou, X.; Liu, S.; Li, C.; Yang, J.; Zhang, L.; and Gao, J. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Cited alongside, same era.
PLA: Language-Driven Open-Vocabulary 3D Scene Understanding
Ding, R.; Yang, J.; Xue, C.; Zhang, W.; Bai, S.; and Qi, X. 2023 · 2023
Cited alongside, same era.
Segment Anything
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; Dollar, P.; and Girshick, R. 2023 · 2023
Cited alongside, same era.
Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data
Lu, S.; Chang, H.; Jing, E. P.; Boularias, A.; and Bekris, K. 2023 · 2023
Cited alongside, same era.
ISBNet: A 3D Point Cloud Instance Segmentation Network With Instance-Aware Sampling and Box-Aware Dynamic Convolution
Ngo, T. D.; Hua, B.-S.; and Nguyen, K. 2023 · 2023
Cited alongside, same era.
Openscene: 3d scene understanding with open vocabularies
Peng, S.; Genova, K.; Jiang, C.; Tagliasacchi, A.; Pollefeys, M.; Funkhouser, T.; et al. 2023 · 2023
Cited alongside, same era.
High Quality Entity Segmentation
Qi, L.; Kuen, J.; Shen, T.; Gu, J.; Li, W.; Guo, W.; Jia, J.; Lin, Z.; and Yang, M.-H. 2023 · 2023
Cited alongside, same era.
Liu, S.; Zeng, Z.; Ren, T.; Li, F.; Zhang, H.; Yang, J.; Jiang, Q.; Li, C.; Yang, J.; Su, H.; et al. 2024 · 2024
Closest in time.
Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annotations
Lyu, R.; Lin, J.; Wang, T.; Mao, X.; Chen, Y.; Xu, R.; Huang, H.; Zhu, C.; Lin, D.; and Pang, J. 2024 · 2024
Closest in time.
Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance
Nguyen, P.; Ngo, T. D.; Kalogerakis, E.; Gan, C.; Tran, A.; Pham, C.; and Nguyen, K. 2024 · 2024
Closest in time.
Grounded sam: Assembling open-world models for diverse visual tasks
Ren, T.; Liu, S.; Zeng, A.; Lin, J.; Li, K.; Cao, H.; Chen, J.; Huang, X.; Chen, Y.; Yan, F.; et al. 2024 · 2024
Closest in time.
A Unified Framework for 3D Scene Understanding
Xu, W.; Shi, C.; Tu, S.; Zhou, X.; Liang, D.; and Bai, X. 2024 · 2024
Closest in time.
Maskclustering: View consensus based mask graph clustering for open-vocabulary 3d instance segmentation
Yan, M.; Zhang, J.; Zhu, Y.; and Wang, H. 2024 · 2024
Closest in time.
Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding
Yang, J.; Ding, R.; Deng, W.; Wang, Z.; and Qi, X. 2024 · 2024
Closest in time.
How to Evaluate the Generalization of Detection? A Benchmark for Comprehensive Open-Vocabulary Detection
Yao, Y.; Liu, P.; Zhao, T.; Zhang, Q.; Liao, J.; Fang, C.; Lee, K.; and Wang, Q. 2024 · 2024
Closest in time.
Sai3d: Segment any instance in 3d scenes
Yin, Y.; Liu, Y.; Xiao, Y.; Cohen-Or, D.; Huang, J.; and Chen, B. 2024 · 2024
Closest in time.
Reason3d: Searching and reasoning 3d segmentation via large language model
Huang, K.-C.; Li, X.; Qi, L.; Yan, S.; and Yang, M.-H. 2025 · 2025
Closest in time.
Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object Detection
Zhao, Y.; Lin, J.; and Lau, R. W. 2025 · 2025
Closest in time.