Fetching the paper…
Reading the bibliography…
Recent advancements in 3D Large Language Models (LLMs) have demonstrated promising capabilities for 3D scene understanding.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner · 2017
Earlier work this paper cites.
Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
P. Achlioptas, A. Abdelreheem, F. Xia, M. Elhoseiny, and L. Guibas · 2020
Earlier work this paper cites.
Scanrefer: 3d object localization in rgb-d scans using natural language
D. Z. Chen, A. X. Chang, and M. Nießner · 2020
Earlier work this paper cites.
Pointgroup: Dual-set point grouping for 3d instance segmentation
L. Jiang, H. Zhao, S. Shi, S. Liu, C.-W. Fu, and J. Jia · 2020
Earlier work this paper cites.
D3net: a speaker-listener architecture for semi-supervised dense captioning and visual grounding in rgb-d scans
D. Z. Chen, Q. Wu, M. Nießner, and A. X. Chang · 2021
Earlier work this paper cites.
Scan2cap: Context-aware dense captioning in rgb-d scans
Z. Chen, A. Gholami, M. Nießner, and A. X. Chang · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Text-guided graph neural networks for referring 3d instance segmentation
P.-H. Huang, H.-H. Lee, H.-T. Chen, and T.-L. Liu · 2021
Earlier work this paper cites.
Sat: 2d semantics assisted training for 3d visual grounding
Z. Yang, S. Zhang, L. Wang, and J. Luo · 2021
Earlier work this paper cites.
Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring
Z. Yuan, X. Yan, Y. Liao, R. Zhang, S. Wang, Z. Li, and S. Cui · 2021
Earlier work this paper cites.
3dvg-transformer: Relation modeling for visual grounding on point clouds
L. Zhao, D. Cai, L. Sheng, and D. Xu · 2021
Earlier work this paper cites.
Scanqa: 3d question answering for spatial scene understanding
D. Azuma, T. Miyanishi, S. Kurita, and M. Kawanabe · 2022
Earlier work this paper cites.
3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds
D. Cai, L. Zhao, J. Zhang, L. Sheng, and D. Xu · 2022
Earlier work this paper cites.
Ham: Hierarchical attention model with high performance for 3d visual grounding
J. Chen, W. Luo, X. Wei, L. Ma, and W. Zhang · 2022
Earlier work this paper cites.
Language conditioned spatial relation reasoning for 3d object grounding
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Earlier work this paper cites.
Multi-view transformer for 3d visual grounding
S. Huang, Y. Chen, J. Jia, and L. Wang · 2022
Earlier work this paper cites.
Bottom up top down detection transformers for language grounding in images and point clouds
A. Jain, N. Gkanatsios, I. Mediratta, and K. Fragkiadaki · 2022
Earlier work this paper cites.
More: Multi-order relation mining for dense captioning in 3d scenes
Y. Jiao, S. Chen, Z. Jie, J. Chen, L. Ma, and Y.-G. Jiang · 2022
Earlier work this paper cites.
3d-sps: Single-stage 3d visual grounding via referred point progressive selection
J. Luo, J. Fu, X. Kong, C. Gao, H. Ren, H. Shen, H. Xia, and S. Liu · 2022
Earlier work this paper cites.
Sqa3d: Situated question answering in 3d scenes
X. Ma, S. Yong, Z. Zheng, Q. Li, Y. Liang, S.-C. Zhu, and S. Huang · 2022
Earlier work this paper cites.
Softgroup for 3d instance segmentation on point clouds
T. Vu, K. Kim, T. M. Luu, T. Nguyen, and C. D. Yoo · 2022
Cited alongside, same era.
X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning
Z. Yuan, X. Yan, Y. Liao, Y. Guo, G. Li, S. Cui, and Z. Li · 2022
Cited alongside, same era.
Ll3da: Visual interactive instruction tuning for omni-3d understanding, reasoning, and planning
S. Chen, X. Chen, C. Zhang, M. Li, G. Yu, H. Fei, H. Zhu, J. Fan, and T. Chen · 2023
Cited alongside, same era.
End-to-end 3d dense captioning with vote2cap-detr
S. Chen, H. Zhu, X. Chen, Y. Lei, G. Yu, and T. Chen · 2023
Cited alongside, same era.
Focalformer3d: focusing on hard instance for 3d object detection
Y. Chen, Z. Yu, Y. Chen, S. Lan, A. Anandkumar, J. Jia, and J. M. Alvarez · 2023
Cited alongside, same era.
V-detr: Detr with vertex relative position encoding for 3d object detection
Y. Shen, Z. Geng, Y. Yuan, Y. Lin, Z. Liu, C. Wang, H. Hu, N. Zheng, and B. Guo · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Three ways to improve verbo-visual fusion for dense 3d visual grounding
O. Unal, C. Sakaridis, S. Saha, F. Yu, and L. Van Gool · 2023
Closest in time.
3drp-net: 3d relative position-aware network for 3d visual grounding
Z. Wang, H. Huang, Y. Zhao, L. Li, X. Cheng, Y. Zhu, A. Yin, and Z. Zhao · 2023
Closest in time.
Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. K. Cheng, S. W. Oh, B. Price, A. Schwing, and J.-Y. Lee · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, et al · 2023
Cited alongside, same era.
Z. Guo, R. Zhang, X. Zhu, Y. Tang, X. Ma, J. Han, K. Chen, P. Gao, X. Li, H. Li, et al · 2023
Cited alongside, same era.
Imagebind-llm: Multi-modality instruction tuning
J. Han, R. Zhang, W. Shao, P. Gao, P. Xu, H. Xiao, K. Zhang, C. Liu, S. Wen, Z. Guo, et al · 2023
Cited alongside, same era.
3d-llm: Injecting the 3d world into large language models
Y. Hong, H. Zhen, P. Chen, S. Zheng, Y. Du, Z. Chen, and C. Gan · 2023
Cited alongside, same era.
An embodied generalist agent in 3d world
J. Huang, S. Yong, X. Ma, X. Linghu, P. Li, Y. Wang, Q. Li, S.-C. Zhu, B. Jia, and S. Huang · 2023
Cited alongside, same era.
Context-aware alignment and mutual masking for 3d-language pre-training
Z. Jin, M. Hayat, Y. Yang, Y. Guo, and Y. Lei · 2023
Cited alongside, same era.
Z. Wang, H. Huang, Y. Zhao, L. Li, X. Cheng, Y. Zhu, A. Yin, and Z. Zhao · 2023
Closest in time.
Chat-3d: Data-efficiently tuning large language model for universal dialogue of 3d scenes
Z. Wang, H. Huang, Y. Zhao, Z. Zhang, and Z. Zhao · 2023
Closest in time.
Eda: Explicit text-decoupling and dense alignment for 3d visual grounding
Y. Wu, X. Cheng, R. Zhang, Z. Cheng, and J. Zhang · 2023
Closest in time.
Pointllm: Empowering large language models to understand point clouds
R. Xu, X. Wang, T. Wang, Y. Chen, J. Pang, and D. Lin · 2023
Closest in time.
Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent
J. Yang, X. Chen, S. Qian, N. Madaan, M. Iyengar, D. F. Fouhey, and J. Chai · 2023
Closest in time.
mplug-owl: Modularization empowers large language models with multimodality
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi, et al · 2023
Closest in time.
Ferret: Refer and ground anything anywhere at any granularity
H. You, H. Zhang, Z. Gan, X. Du, B. Zhang, Z. Wang, L. Cao, S.-F. Chang, and Y. Yang · 2023
Closest in time.
Sam3d: Zero-shot 3d object detection via segment anything model
D. Zhang, D. Liang, H. Yang, Z. Zou, X. Ye, Z. Liu, and X. Bai · 2023
Closest in time.
Multi3drefer: Grounding text description to multiple 3d objects
Y. Zhang, Z. Gong, and A. X. Chang · 2023
Closest in time.
Bubogpt: Enabling visual grounding in multi-modal llms
Y. Zhao, Z. Lin, D. Zhou, Z. Huang, J. Feng, and B. Kang · 2023
Closest in time.
Uni3d: Exploring unified 3d representation at scale
J. Zhou, J. Wang, B. Ma, Y.-S. Liu, T. Huang, and X. Wang · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Closest in time.
3d-vista: Pre-trained transformer for 3d vision and text alignment
Z. Zhu, X. Ma, Y. Chen, Z. Deng, S. Huang, and Q. Li · 2023
Closest in time.
Vote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning
S. Chen, H. Zhu, M. Li, X. Chen, P. Guo, Y. Lei, Y. Gang, T. Li, and T. Chen · 2024
Closest in time.
Scene-llm: Extending language model for 3d visual understanding and reasoning
R. Fu, J. Liu, X. Chen, Y. Nie, and W. Xiong · 2024
Closest in time.
Dora: 3d visual grounding with order-aware referring
T.-Y. Wu, S.-Y. Huang, and Y.-C. F. Wang · 2024
Closest in time.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Z. Yin, J. Wang, J. Cao, Z. Shi, D. Liu, M. Li, X. Huang, Z. Wang, L. Sheng, L. Bai, et al · 2024
Closest in time.
Ferret-v2: An improved baseline for referring and grounding with large language models
H. Zhang, H. You, P. Dufter, B. Zhang, C. Chen, H.-Y. Chen, T.-J. Fu, W. Y. Wang, S.-F. Chang, Z. Gan, et al · 2024
Closest in time.