Fetching the paper…
Reading the bibliography…
Prior studies on 3D scene understanding have primarily developed specialized models for specific tasks or required task-specific fine-tuning.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in CVPR , 2017, pp. 5828–5839
2017
Earlier work this paper cites.
B. Graham, M. Engelcke, and L. Van Der Maaten, “3d semantic segmentation with submanifold sparse convolutional networks,” in CVPR , 2018, pp. 9224–9232
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
L. Jiang, H. Zhao, S. Shi, S. Liu, C.-W. Fu, and J. Jia, “Pointgroup: Dual-set point grouping for 3d instance segmentation,” in CVPR , 2020, pp. 4867–4876
2020
Earlier work this paper cites.
D. Z. Chen, A. X. Chang, and M. Nießner, “Scanrefer: 3d object localization in rgb-d scans using natural language,” in ECCV . Springer, 2020, pp. 202–221
2020
Earlier work this paper cites.
P. Achlioptas, A. Abdelreheem, F. Xia, M. Elhoseiny, and L. Guibas, “Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes,” in ECCV . Springer, 2020, pp. 422–440
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV . Springer, 2020, pp. 213–229
2020
Earlier work this paper cites.
Z. Liu, Z. Zhang, Y. Cao, H. Hu, and X. Tong, “Group-free 3d object detection via transformers,” in ICCV , 2021, pp. 2949–2958
2021
Earlier work this paper cites.
L. Zhao, D. Cai, L. Sheng, and D. Xu, “3dvg-transformer: Relation modeling for visual grounding on point clouds,” in ICCV , 2021, pp. 2928–2937
2021
Earlier work this paper cites.
D. Z. Chen, Q. Wu, M. Nießner, and A. X. Chang, “D3net: A speaker-listener architecture for semi-supervised dense captioning and visual grounding in rgb-d scans,” 2021
2021
Earlier work this paper cites.
Z. Chen, A. Gholami, M. Nießner, and A. X. Chang, “Scan2cap: Context-aware dense captioning in rgb-d scans,” in CVPR , 2021, pp. 3193–3203
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Zareian, K. D. Rosa, D. H. Hu, and S.-F. Chang, “Open-vocabulary object detection using captions,” in CVPR , 2021, pp. 14 393–14 402
2021
Earlier work this paper cites.
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion, “Mdetr-modulated detection for end-to-end multi-modal understanding,” in ICCV , 2021, pp. 1780–1790
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
T. Vu, K. Kim, T. M. Luu, T. Nguyen, and C. D. Yoo, “Softgroup for 3d instance segmentation on point clouds,” in CVPR , 2022, pp. 2708–2717
2022
Earlier work this paper cites.
S. Huang, Y. Chen, J. Jia, and L. Wang, “Multi-view transformer for 3d visual grounding,” in CVPR , 2022, pp. 15 524–15 533
2022
Earlier work this paper cites.
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev, “Language conditioned spatial relation reasoning for 3d object grounding,” NeurIPS , vol. 35, pp. 20 522–20 535, 2022
2022
Earlier work this paper cites.
D. Azuma, T. Miyanishi, S. Kurita, and M. Kawanabe, “Scanqa: 3d question answering for spatial scene understanding,” in CVPR , 2022, pp. 19 129–19 139
2022
Earlier work this paper cites.
H. Wang, C. Zhang, J. Yu, and W. Cai, “Spatiality-guided transformer for 3D dense captioning on point clouds,” in IJCAI , 2022
2022
Earlier work this paper cites.
D. Cai, L. Zhao, J. Zhang, L. Sheng, and D. Xu, “3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds,” in CVPR , 2022, pp. 16 464–16 473
2022
Earlier work this paper cites.
A. Jain, N. Gkanatsios, I. Mediratta, and K. Fragkiadaki, “Bottom up top down detection transformers for language grounding in images and point clouds,” in ECCV . Springer, 2022, pp. 417–433
2022
Earlier work this paper cites.
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang et al. , “Grounded language-image pre-training,” in CVPR , 2022, pp. 10 965–10 975
2022
Earlier work this paper cites.
L. Yao, J. Han, Y. Wen, X. Liang, D. Xu, W. Zhang, Z. Li, C. Xu, and H. Xu, “Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection,” NeurIPS , vol. 35, pp. 9125–9138, 2022
2022
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” NeurIPS , vol. 35, pp. 23 716–23 736, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Z. Yuan, X. Yan, Y. Liao, Y. Guo, G. Li, S. Cui, and Z. Li, “X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning,” in CVPR , 2022, pp. 8563–8573
2022
Earlier work this paper cites.
Y. Jiao, S. Chen, Z. Jie, J. Chen, L. Ma, and Y.-G. Jiang, “More: Multi-order relation mining for dense captioning in 3d scenes,” in ECCV . Springer, 2022, pp. 528–545
2022
Cited alongside, same era.
X. Yu, L. Tang, Y. Rao, T. Huang, J. Zhou, and J. Lu, “Point-bert: Pre-training 3d point cloud transformers with masked point modeling,” in CVPR , 2022, pp. 19 313–19 322
2022
Cited alongside, same era.
2022
Cited alongside, same era.
B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” in CVPR , 2022, pp. 1290–1299
2022
Cited alongside, same era.
R. Chen, Y. Liu, L. Kong, X. Zhu, Y. Ma, Y. Li, Y. Hou, Y. Qiao, and W. Wang, “Clip2scene: Towards label-efficient 3d scene understanding by clip,” in CVPR , 2023, pp. 7020–7030
2023
Later among the works it cites.
2023
Later among the works it cites.
——, “Pla: Language-driven open-vocabulary 3d scene understanding,” in CVPR , 2023, pp. 7010–7019
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Peng, K. Genova, C. M. Jiang, A. Tagliasacchi, M. Pollefeys, and T. Funkhouser, “Openscene: 3d scene understanding with open vocabularies,” in CVPR , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
J. Schult, F. Engelmann, A. Hermans, O. Litany, S. Tang, and B. Leibe, “Mask3d: Mask transformer for 3d semantic instance segmentation,” in ICRA . IEEE, 2023, pp. 8216–8223
2023
Cited alongside, same era.
2023
Cited alongside, same era.
M. Parelli, A. Delitzas, N. Hars, G. Vlassis, S. Anagnostidis, G. Bachmann, and T. Hofmann, “Clip-guided vision-language pre-training for question answering in 3d scenes,” in CVPR , 2023, pp. 5606–5611
2023
Cited alongside, same era.
Z. Jin, M. Hayat, Y. Yang, Y. Guo, and Y. Lei, “Context-aware alignment and mutual masking for 3d-language pre-training,” in CVPR , 2023, pp. 10 984–10 994
2023
Cited alongside, same era.
Z. Chen, R. Hu, X. Chen, M. Nießner, and A. X. Chang, “Unit3d: A unified transformer for 3d dense captioning and visual grounding,” in ICCV , 2023, pp. 18 109–18 119
2023
Cited alongside, same era.
S. Chen, H. Zhu, X. Chen, Y. Lei, G. Yu, and T. Chen, “End-to-end 3d dense captioning with vote2cap-detr,” in CVPR , 2023, pp. 11 124–11 133
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
M. Parelli, A. Delitzas, N. Hars, G. Vlassis, S. Anagnostidis, G. Bachmann, and T. Hofmann, “Clip-guided vision-language pre-training for question answering in 3d scenes,” in CVPR , 2023, pp. 5606–5611
2023
Later among the works it cites.
M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. VanderBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi, “Objaverse: A universe of annotated 3d objects,” in CVPR , 2023, pp. 13 142–13 153
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” NeurIPS , vol. 36, 2024
2024
Closest in time.
W. Wang, Z. Chen, X. Chen, J. Wu, X. Zhu, G. Zeng, P. Luo, T. Lu, J. Zhou, Y. Qiao et al. , “Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,” NeurIPS , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Sadigh, L. Guibas, and F. Xia, “Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,” in CVPR , 2024, pp. 14 455–14 465
2024
Closest in time.
M. El Banani, A. Raj, K.-K. Maninis, A. Kar, Y. Li, M. Rubinstein, D. Sun, L. Guibas, J. Johnson, and V. Jampani, “Probing the 3d awareness of visual foundation models,” in CVPR , 2024, pp. 21 795–21 806
2024
Closest in time.
T. Wang, X. Mao, C. Zhu, R. Xu, R. Lyu, P. Li, X. Chen, W. Zhang, K. Chen, T. Xue, X. Liu, C. Lu, D. Lin, and J. Pang, “Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai,” in CVPR , 2024
2024
Closest in time.
J. Yang, R. Ding, W. Deng, Z. Wang, and X. Qi, “Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding,” in CVPR , 2024, pp. 19 823–19 832
2024
Closest in time.
Y. Man, L.-Y. Gui, and Y.-X. Wang, “Situational awareness matters in 3d vision language reasoning,” in CVPR , 2024
2024
Closest in time.
M. Liu, R. Shi, K. Kuang, Y. Zhu, X. Li, S. Han, H. Cai, F. Porikli, and H. Su, “Openshape: Scaling up 3d shape representation towards open-world understanding,” NeurIPS , vol. 36, 2024
2024
Closest in time.
J. Zhou, J. Wang, B. Ma, Y.-S. Liu, T. Huang, and X. Wang, “Uni3d: Exploring unified 3d representation at scale,” in ICLR , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.