Fetching the paper…
Reading the bibliography…
With the emergence of LLMs and their integration with other data modalities, multi-modal 3D perception attracts more attention due to its connectivity to the physical world and makes rapid progress.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Sun rgb-d: A rgb-d scene understanding benchmark suite
S. Song, S. P. Lichtenberg, and J. Xiao · 2015
Earlier work this paper cites.
Matterport3D: Learning from RGB-D data in indoor environments
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang · 2017
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner · 2017
Earlier work this paper cites.
Deep hough voting for 3d object detection in point clouds
C. R. Qi, O. Litany, K. He, and L. J. Guibas · 2019
Earlier work this paper cites.
Rio: 3d object instance re-localization in changing indoor environments
J. Wald, A. Avetisyan, N. Navab, F. Tombari, and M. Nießner · 2019
Earlier work this paper cites.
Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
P. Achlioptas, A. Abdelreheem, F. Xia, M. Elhoseiny, and L. Guibas · 2020
Earlier work this paper cites.
Scanrefer: 3d object localization in rgb-d scans using natural language
D. Z. Chen, A. X. Chang, and M. Nießner · 2020
Earlier work this paper cites.
ARKitscenes - a diverse real-world dataset for 3d indoor scene understanding using mobile RGB-d data
G. Baruch, Z. Chen, A. Dehghan, T. Dimry, Y. Feigin, P. Fu, T. Gebauer, B. Joffe, D. Kurz, A. Schwartz, and E. Shulman · 2021
Earlier work this paper cites.
D3net: A speaker-listener architecture for semi-supervised dense captioning and visual grounding in rgb-d scans, 2021
D. Z. Chen, Q. Wu, M. Nießner, and A. X. Chang · 2021
Earlier work this paper cites.
Scan2cap: Context-aware dense captioning in rgb-d scans
Z. Chen, A. Gholami, M. Nießner, and A. X. Chang · 2021
Earlier work this paper cites.
Habitat-matterport 3d dataset (HM3d): 1000 large-scale 3d environments for embodied AI
S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. M. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang, M. Savva, Y. Zhao, and D. Batra · 2021
Earlier work this paper cites.
3dvg-transformer: Relation modeling for visual grounding on point clouds
L. Zhao, D. Cai, L. Sheng, and D. Xu · 2021
Earlier work this paper cites.
Scanqa: 3d question answering for spatial scene understanding
D. Azuma, T. Miyanishi, S. Kurita, and M. Kawanabe · 2022
Earlier work this paper cites.
3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds
D. Cai, L. Zhao, J. Zhang, L. Sheng, and D. Xu · 2022
Earlier work this paper cites.
Language conditioned spatial relation reasoning for 3d object grounding
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev · 2022
Earlier work this paper cites.
Multi-view transformer for 3d visual grounding, 2022
S. Huang, Y. Chen, J. Jia, and L. Wang · 2022
Earlier work this paper cites.
Bottom up top down detection transformers for language grounding in images and point clouds
A. Jain, N. Gkanatsios, I. Mediratta, and K. Fragkiadaki · 2022
Earlier work this paper cites.
More: Multi-order relation mining for dense captioning in 3d scenes
Y. Jiao, S. Chen, Z. Jie, J. Chen, L. Ma, and Y.-G. Jiang · 2022
Cited alongside, same era.
Sqa3d: Situated question answering in 3d scenes
X. Ma, S. Yong, Z. Zheng, Q. Li, Y. Liang, S.-C. Zhu, and S. Huang · 2022
Cited alongside, same era.
X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning
Z. Yuan, X. Yan, Y. Liao, Y. Guo, G. Li, S. Cui, and Z. Li · 2022
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Cited alongside, same era.
Ll3da: Visual interactive instruction tuning for omni-3d understanding, reasoning, and planning
S. Chen, X. Chen, C. Zhang, M. Li, G. Yu, H. Fei, H. Zhu, J. Fan, and T. Chen · 2023
Three ways to improve verbo-visual fusion for dense 3d visual grounding
O. Unal, C. Sakaridis, S. Saha, F. Yu, and L. Van Gool · 2023
Later among the works it cites.
Cogvlm: Visual expert for pretrained language models
W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, X. Song, et al · 2023
Later among the works it cites.
Chat-3d: Data-efficiently tuning large language model for universal dialogue of 3d scenes
Z. Wang, H. Huang, Y. Zhao, Z. Zhang, and Z. Zhao · 2023
Later among the works it cites.
Pointllm: Empowering large language models to understand point clouds
R. Xu, X. Wang, T. Wang, Y. Chen, J. Pang, and D. Lin · 2023
Later among the works it cites.
Multi3drefer: Grounding text description to multiple 3d objects
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
End-to-end 3d dense captioning with vote2cap-detr
S. Chen, H. Zhu, X. Chen, Y. Lei, G. Yu, and T. Chen · 2023
Cited alongside, same era.
Opencompass: A universal evaluation platform for foundation models
O. Contributors · 2023
Cited alongside, same era.
Z. Guo, R. Zhang, X. Zhu, Y. Tang, X. Ma, J. Han, K. Chen, P. Gao, X. Li, H. Li, et al · 2023
Cited alongside, same era.
3d-llm: Injecting the 3d world into large language models
Y. Hong, H. Zhen, P. Chen, S. Zheng, Y. Du, Z. Chen, and C. Gan · 2023
Cited alongside, same era.
Chat-3d v2: Bridging 3d scene and large language models with object identifiers
H. Huang, Z. Wang, R. Huang, L. Liu, X. Cheng, Y. Zhao, T. Jin, and Z. Zhao · 2023
Cited alongside, same era.
An embodied generalist agent in 3d world
J. Huang, S. Yong, X. Ma, X. Linghu, P. Li, Y. Wang, Q. Li, S.-C. Zhu, B. Jia, and S. Huang · 2023
Cited alongside, same era.
Context-aware alignment and mutual masking for 3d-language pre-training
Z. Jin, M. Hayat, Y. Yang, Y. Guo, and Y. Lei · 2023
Cited alongside, same era.
Y. Zhang, Z. Gong, and A. X. Chang · 2023
Later among the works it cites.
Object2scene: Putting objects in context for open-vocabulary 3d detection
C. Zhu, W. Zhang, T. Wang, X. Liu, and K. Chen · 2023
Later among the works it cites.
3d-vista: Pre-trained transformer for 3d vision and text alignment
Z. Zhu, X. Ma, Y. Chen, Z. Deng, S. Huang, and Q. Li · 2023
Later among the works it cites.
Grounded 3d-llm with referent tokens
Y. Chen, S. Yang, H. Huang, T. Wang, R. Lyu, R. Xu, D. Lin, and J. Pang · 2024
Closest in time.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Z. Chen, W. Wang, H. Tian, S. Ye, Z. Gao, E. Cui, W. Tong, K. Hu, J. Luo, Z. Ma, et al · 2024
Closest in time.
X. Dong, P. Zhang, Y. Zang, Y. Cao, B. Wang, L. Ouyang, X. Wei, S. Zhang, H. Duan, M. Cao, W. Zhang, Y. Li, H. Yan, Y. Gao, X. Zhang, W. Li, J. Li, K. Chen, C. He, X. Zhang, Y. Qiao, D. Lin, and J. Wang · 2024
Closest in time.
Sceneverse: Scaling 3d vision-language learning for grounded scene understanding
B. Jia, Y. Chen, H. Yu, Y. Wang, X. Niu, T. Liu, Q. Li, and S. Huang · 2024
Closest in time.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2024
Closest in time.
Cross3dvg: Cross-dataset 3d visual grounding on different rgb-d scans, 2024
T. Miyanishi, D. Azuma, S. Kurita, and M. Kawanabe · 2024
Closest in time.
Shapellm: Universal 3d object understanding for embodied interaction
Z. Qi, R. Dong, S. Zhang, H. Geng, C. Han, Z. Ge, L. Yi, and K. Ma · 2024
Closest in time.
Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai
T. Wang, X. Mao, C. Zhu, R. Xu, R. Lyu, P. Li, X. Chen, W. Zhang, K. Chen, T. Xue, X. Liu, C. Lu, D. Lin, and J. Pang · 2024
Closest in time.
Scanreason: Empowering 3d visual grounding with reasoning capabilities, 2024
C. Zhu, T. Wang, W. Zhang, K. Chen, and X. Liu · 2024
Closest in time.
Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness, 2024
C. Zhu, T. Wang, W. Zhang, J. Pang, and X. Liu · 2024
Closest in time.