Fetching the paper…
Reading the bibliography…
Developing vision-language models (VLMs) capable of understanding 3D scenes has been a longstanding research goal.
B. Graham, “Sparse 3d convolutional neural networks,” in
2015
Earlier work this paper cites.
D. Maturana and S. Scherer, “Voxnet: A 3d convolutional neural network for real-time object recognition,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in
2017
Earlier work this paper cites.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in
2017
Earlier work this paper cites.
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” in
2017
Earlier work this paper cites.
G. Riegler, A. Osman Ulusoy, and A. Geiger, “Octnet: Learning deep 3d representations at high resolutions,” in
2017
Earlier work this paper cites.
M. Tatarchenko, A. Dosovitskiy, and T. Brox, “Octree generating networks: Efficient convolutional architectures for high-resolution 3d outputs,” in
2017
Earlier work this paper cites.
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang, “Matterport3d: Learning from rgb-d data in indoor environments,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,”
2017
Earlier work this paper cites.
A. V. Phan, M. Le Nguyen, Y. L. H. Nguyen, and L. T. Bui, “Dgcnn: A convolutional neural network over large-scale labeled graphs,”
2018
Earlier work this paper cites.
Y. Li, R. Bu, M. Sun, W. Wu, X. Di, and B. Chen, “Pointcnn: Convolution on x-transformed points,” in
2018
Earlier work this paper cites.
B. Graham, M. Engelcke, and L. Van Der Maaten, “3d semantic segmentation with submanifold sparse convolutional networks,” in
2018
Earlier work this paper cites.
A. Dai and M. Nießner, “3dmv: Joint 3d-multi-view prediction for 3d semantic scene segmentation,” in
2018
Earlier work this paper cites.
J. Wald, A. Avetisyan, N. Navab, F. Tombari, and M. Nießner, “Rio: 3d object instance re-localization in changing indoor environments,” in
2019
Earlier work this paper cites.
W. Wu, Z. Qi, and L. Fuxin, “Pointconv: Deep convolutional networks on 3d point clouds,” in
2019
Earlier work this paper cites.
C. R. Qi, O. Litany, K. He, and L. J. Guibas, “Deep hough voting for 3d object detection in point clouds,” in
2019
Earlier work this paper cites.
C. Choy, J. Gwak, and S. Savarese, “4d spatio-temporal convnets: Minkowski convolutional neural networks,” in
2019
Earlier work this paper cites.
M. Tancik, P. Srinivasan, B. Mildenhall, S. Fridovich-Keil, N. Raghavan, U. Singhal, R. Ramamoorthi, J. Barron, and R. Ng, “Fourier features let networks learn high frequency functions in low dimensional domains,” in
2020
Earlier work this paper cites.
J. Zheng, J. Zhang, J. Li, R. Tang, S. Gao, and Z. Zhou, “Structured3d: A large photo-realistic dataset for structured 3d modeling,” in
2020
Earlier work this paper cites.
D. Z. Chen, A. X. Chang, and M. Nießner, “Scanrefer: 3d object localization in rgb-d scans using natural language,” in
2020
Earlier work this paper cites.
P. Achlioptas, A. Abdelreheem, F. Xia, M. Elhoseiny, and L. Guibas, “Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes,” in
2020
Earlier work this paper cites.
G. Baruch, Z. Chen, A. Dehghan, T. Dimry, Y. Feigin, P. Fu, T. Gebauer, B. Joffe, D. Kurz, A. Schwartz
2021
Earlier work this paper cites.
H. Zhao, L. Jiang, J. Jia, P. H. Torr, and V. Koltun, “Point transformer,” in
2021
Earlier work this paper cites.
S. Huang, Y. Xie, S.-C. Zhu, and Y. Zhu, “Spatio-temporal self-supervised representation learning for 3d point clouds,” in
2021
Earlier work this paper cites.
L. Zhao, D. Cai, L. Sheng, and D. Xu, “3dvg-transformer: Relation modeling for visual grounding on point clouds,” in
2021
Earlier work this paper cites.
S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang
2021
Earlier work this paper cites.
H. Fu, B. Cai, L. Gao, L.-X. Zhang, J. Wang, C. Li, Q. Zeng, C. Sun, R. Jia, B. Zhao
2021
Earlier work this paper cites.
Z. Chen, A. Gholami, M. Nießner, and A. X. Chang, “Scan2cap: Context-aware dense captioning in rgb-d scans,” in
2021
Earlier work this paper cites.
Y. Mao, Y. Zhang, H. Jiang, A. Chang, and M. Savva, “Multiscan: Scalable rgbd scanning for 3d environments with articulated objects,” in
2022
Earlier work this paper cites.
X. Yu, L. Tang, Y. Rao, T. Huang, J. Zhou, and J. Lu, “Point-bert: Pre-training 3d point cloud transformers with masked point modeling,” in
2022
Earlier work this paper cites.
S. Huang, Y. Chen, J. Jia, and L. Wang, “Multi-view transformer for 3d visual grounding,” in
2022
Earlier work this paper cites.
A. Abdelreheem, U. Upadhyay, I. Skorokhodov, R. Al Yahya, J. Chen, and M. Elhoseiny, “3dreftransformer: Fine-grained object identification in real-world scenes using natural language,” in
2022
Earlier work this paper cites.
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev, “Language conditioned spatial relation reasoning for 3d object grounding,” in
2022
Earlier work this paper cites.
M. Deitke, E. VanderBilt, A. Herrasti, L. Weihs, K. Ehsani, J. Salvador, W. Han, E. Kolve, A. Kembhavi, and R. Mottaghi, “Procthor: Large-scale embodied ai using procedural generation,” in
2022
Earlier work this paper cites.
D. Azuma, T. Miyanishi, S. Kurita, and M. Kawanabe, “Scanqa: 3d question answering for spatial scene understanding,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in
2022
Earlier work this paper cites.
D. Cai, L. Zhao, J. Zhang, L. Sheng, and D. Xu, “3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds,” in
2022
Earlier work this paper cites.
OpenAI, “Gpt-4 technical report,”
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
Z. Zhu, X. Ma, Y. Chen, Z. Deng, S. Huang, and Q. Li, “3d-vista: Pre-trained transformer for 3d vision and text alignment,” in
2023
Cited alongside, same era.
T. Luo, C. Rockwell, H. Lee, and J. Johnson, “Scalable 3d captioning with pretrained models,” in
2023
Cited alongside, same era.
J. Schult, F. Engelmann, A. Hermans, O. Litany, S. Tang, and B. Leibe, “Mask3d: Mask transformer for 3d semantic instance segmentation,” in
2023
Cited alongside, same era.
M. Liu, R. Shi, K. Kuang, Y. Zhu, X. Li, S. Han, H. Cai, F. Porikli, and H. Su, “Openshape: Scaling up 3d shape representation towards open-world understanding,” in
2023
Cited alongside, same era.
W. Yuan, R. Y. Pang, K. Cho, S. Sukhbaatar, J. Xu, and J. Weston, “Self-rewarding language models,” in
2024
Later among the works it cites.
R. Y. Pang, W. Yuan, K. Cho, H. He, S. Sukhbaatar, and J. Weston, “Iterative reasoning preference optimization,” in
2024
Later among the works it cites.
Z. Liu, M. Lu, S. Zhang, B. Liu, H. Guo, Y. Yang, J. Blanchet, and Z. Wang, “Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer,” in
2024
Later among the works it cites.
S. Li, R. Lin, and S. Pei, “Multi-modal preference alignment remedies regression of visual instruction tuning on language model,” in
2024
Later among the works it cites.
R. Pi, T. Han, W. Xiong, J. Zhang, R. Liu, R. Pan, and T. Zhang, “Strengthening multimodal large language model with bootstrapped preference optimization,” in
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Hong, H. Zhen, P. Chen, S. Zheng, Y. Du, Z. Chen, and C. Gan, “3d-llm: Injecting the 3d world into large language models,” in
2023
Cited alongside, same era.
S. Peng, K. Genova, C. Jiang, A. Tagliasacchi, M. Pollefeys, T. Funkhouser
2023
Cited alongside, same era.
N. M. M. Shafiullah, C. Paxton, L. Pinto, S. Chintala, and A. Szlam, “Clip-fields: Weakly supervised semantic fields for robotic memory,” in
2023
Cited alongside, same era.
R. Ding, J. Yang, C. Xue, W. Zhang, S. Bai, and X. Qi, “Pla: Language-driven open-vocabulary 3d scene understanding,” in
2023
Cited alongside, same era.
K. M. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, A. Maalouf, S. Li, G. Iyer, S. Saryazdi, N. Keetha
2023
Cited alongside, same era.
C. Huang, O. Mees, A. Zeng, and W. Burgard, “Visual language maps for robot navigation,” in
2023
Cited alongside, same era.
C. Yeshwanth, Y.-C. Liu, M. Nießner, and A. Dai, “Scannet++: A high-fidelity dataset of 3d indoor scenes,” in
2023
Cited alongside, same era.
Later among the works it cites.
F. Wang, W. Zhou, J. Y. Huang, N. Xu, S. Zhang, H. Poon, and M. Chen, “mdpo: Conditional preference optimization for multimodal large language models,” in
2024
Later among the works it cites.
Y. Xie, G. Li, X. Xu, and M.-Y. Kan, “V-dpo: Mitigating hallucination in large vision language models via vision-guided direct preference optimization,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Chen, H. Zhu, M. Li, X. Chen, P. Guo, Y. Lei, G. Yu, T. Li, and T. Chen, “Vote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning,”
2024
Later among the works it cites.
T. Zhang, S. He, T. Dai, Z. Wang, B. Chen, and S.-T. Xia, “Vision-language pre-training with object contrastive learning for 3d scene understanding,” in
2024
Later among the works it cites.
C. Zhu, T. Wang, W. Zhang, K. Chen, and X. Liu, “Scanreason: Empowering 3d visual grounding with reasoning capabilities,” in
2024
Later among the works it cites.
J. Yang, S. Yang, A. W. Gupta, R. Han, L. Fei-Fei, and S. Xie, “Thinking in space: How multimodal large language models see, remember, and recall spaces,” in
2025
Closest in time.
C. H. Song, V. Blukis, J. Tremblay, S. Tyree, Y. Su, and S. Birchfield, “Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics,” in
2025
Closest in time.
2025
Closest in time.
J. Huang, B. Jia, Y. Wang, Z. Zhu, X. Linghu, Q. Li, S.-C. Zhu, and S. Huang, “Unveiling the mist over 3d vision-language understanding: Object-centric evaluation with chain-of-analysis,” in
2025
Closest in time.
H. Zhu, H. Yang, X. Wu, D. Huang, S. Zhang, X. He, H. Zhao, C. Shen, Y. Qiao, T. He
2025
Closest in time.
Y. Wang, B. Jia, Z. Zhu, and S. Huang, “Masked point-entity contrast for open-vocabulary 3d scene understanding,” in
2025
Closest in time.
R. Fu, J. Liu, X. Chen, Y. Nie, and W. Xiong, “Scene-llm: Extending language model for 3d visual understanding and reasoning,” in
2025
Closest in time.
J. Deng, T. He, L. Jiang, T. Wang, F. Dayoub, and I. Reid, “3d-llava: Towards generalist 3d lmms with omni superpoint transformer,” in
2025
Closest in time.
J. Luo, Y. Liu, W. Chen, Z. Li, Y. Wang, G. Li, and L. Lin, “Dspnet: Dual-vision scene perception for robust 3d question answering,” in
2025
Closest in time.
C. Zhu, T. Wang, W. Zhang, J. Pang, and X. Liu, “Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness,” in
2025
Closest in time.
D. Zheng, S. Huang, and L. Wang, “Video-3d llm: Learning position-aware video representation for 3d scene understanding,” in
2025
Closest in time.
H. Zhi, P. Chen, J. Li, S. Ma, X. Sun, T. Xiang, Y. Lei, M. Tan, and C. Gan, “Lscenellm: Enhancing large 3d scene understanding using adaptive visual preferences,” in
2025
Closest in time.
H. Yu, W. Li, S. Wang, J. Chen, and J. Zhu, “Inst3d-lmm: Instance-aware 3d scene understanding with multi-modal instruction tuning,” in
2025
Closest in time.
A. Thai, S. Peng, K. Genova, L. Guibas, and T. Funkhouser, “Splattalk: 3d vqa with gaussian splatting,” in
2025
Closest in time.
D. Zheng, S. Huang, Y. Li, and L. Wang, “Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors,” in
2025
Closest in time.
X. Huang, J. Wu, Q. Xie, and K. Han, “Mllms need 3d-aware representation supervision for scene understanding,” in
2025
Closest in time.
H.-W. Huang, F.-C. Chen, W. Chai, C.-C. Su, L. Xia, S. Jung, C.-Y. Yang, J.-N. Hwang, M. Sun, and C.-H. Kuo, “Zero-shot 3d question answering via voxel-based dynamic token compression,” in
2025
Closest in time.
2025
Closest in time.
J. Yang, X. Chen, N. Madaan, M. Iyengar, S. Qian, D. F. Fouhey, and J. Chai, “3d-grand: A million-scale dataset for 3d-llms with better grounding and less hallucination,” in
2025
Closest in time.
J. Zhang, Y. Chen, Y. Zhou, Y. Xu, Z. Huang, J. Mei, J. Chen, Y.-J. Yuan, X. Cai, G. Huang
2025
Closest in time.
2025
Closest in time.
W. Zhang, R. Peng, C. Gao, J. Fang, X. Zeng, K. Li, Z. Wang, J. Cui, X. Wang, X. Chen
2025
Closest in time.
W. Kang, H. Huang, Y. Shang, M. Shah, and Y. Yan, “Robin3d: Improving 3d large language model via robust instruction tuning,” in
2025
Closest in time.
T. Chu, Y. Zhai, J. Yang, S. Tong, S. Xie, D. Schuurmans, Q. V. Le, S. Levine, and Y. Ma, “Sft memorizes, rl generalizes: A comparative study of foundation model post-training,” in
2025
Closest in time.
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi
2025
Closest in time.
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao
2025
Closest in time.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang
2025
Closest in time.
Z. Qi, Z. Zhang, Y. Fang, J. Wang, and H. Zhao, “Gpt4scene: Understand 3d scenes from videos with vision-language models,” in
2026
Closest in time.
Y. Zhao, J. Lin, S. Ye, Q. Pang, and R. W. Lau, “Openscan: A benchmark for generalized open-vocabulary 3d scene understanding,” in
2026
Closest in time.
X. Linghu, J. Huang, Z. Zhu, B. Jia, and S. Huang, “Scenecot: Eliciting grounded chain-of-thought reasoning in 3d scenes,” in
2026
Closest in time.