Fetching the paper…
Reading the bibliography…
Enabling Large Language Models (LLMs) to comprehend the 3D physical world remains a significant challenge.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N.; and Gurevych, I. 2019 · 1908
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Wu, Z.; Song, S.; Khosla, A.; Yu, F.; Zhang, L.; Tang, X.; and Xiao, J. 2015 · 1920
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Generating wikipedia by summarizing long sequences
Liu, P. J.; Saleh, M.; Pot, E.; Goodrich, B.; Sepassi, R.; Kaiser, L.; and Shazeer, N. 2018 · 2018
Earlier work this paper cites.
Simcse: Simple contrastive learning of sentence embeddings
Gao, T.; Yao, X.; and Chen, D. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. 2022 · 2022
Earlier work this paper cites.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y.; Kordi, Y.; Mishra, S.; Liu, A.; Smith, N. A.; Khashabi, D.; and Hajishirzi, H. 2022 · 2022
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; Choromanski, K.; Ding, T.; Driess, D.; Dubey, A.; Finn, C.; et al. 2023 · 2023
Earlier work this paper cites.
Objaverse: A universe of annotated 3d objects
Deitke, M.; Schwenk, D.; Salvador, J.; Weihs, L.; Michel, O.; VanderBilt, E.; Schmidt, L.; Ehsani, K.; Kembhavi, A.; and Farhadi, A. 2023 · 2023
Earlier work this paper cites.
Palm-e: An embodied multimodal language model
Driess, D.; Xia, F.; Sajjadi, M. S.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. 2023 · 2023
Cited alongside, same era.
Guo, Z.; Zhang, R.; Zhu, X.; Tang, Y.; Ma, X.; Han, J.; Chen, K.; Gao, P.; Li, X.; Li, H.; et al. 2023 · 2023
Cited alongside, same era.
3d-llm: Injecting the 3d world into large language models
Hong, Y.; Zhen, H.; Chen, P.; Zheng, S.; Du, Y.; Chen, Z.; and Gan, C. 2023 · 2023
Cited alongside, same era.
Clip2point: Transfer clip to point cloud classification with image-depth pre-training
Huang, T.; Dong, B.; Yang, Y.; Huang, X.; Lau, R. W.; Ouyang, W.; and Zuo, W. 2023 · 2023
Cited alongside, same era.
Gpt-4 technical report. arxiv 2303.08774
OpenAI, R. 2023 · 2023
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
Abdin, M.; Jacobs, S. A.; Awan, A. A.; Aneja, J.; Awadallah, A.; Awadalla, H.; Bach, N.; Bahree, A.; Bakhtiari, A.; Behl, H.; et al. 2024 · 2024
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P. N.; and Hoi, S. 2024 · 2024
Closest in time.
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024 · 2024
Closest in time.
Drive like a human: Rethinking autonomous driving with large language models
Fu, D.; Li, X.; Wen, L.; Dou, M.; Cai, P.; Shi, B.; and Qiao, Y. 2024 · 2024
Closest in time.
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
Gao, Y.; Wang, Z.; Zheng, W.-S.; Xie, C.; and Zhou, Y. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Contrast with reconstruct: Contrastive 3d representation learning guided by generative pretraining
Qi, Z.; Dong, R.; Fan, G.; Ge, Z.; Zhang, X.; Ma, K.; and Yi, L. 2023 · 2023
Cited alongside, same era.
Eva-clip: Improved training techniques for clip at scale
Sun, Q.; Fang, Y.; Wu, L.; Wang, X.; and Cao, Y. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Cited alongside, same era.
Pointllm: Empowering large language models to understand point clouds
Xu, R.; Wang, X.; Wang, T.; Chen, Y.; Pang, J.; and Lin, D. 2023 · 2023
Cited alongside, same era.
Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding
Xue, L.; Gao, M.; Xing, C.; Martín-Martín, R.; Wu, J.; Xiong, C.; Xu, R.; Niebles, J. C.; and Savarese, S. 2023 · 2023
Cited alongside, same era.
Uni3d: Exploring unified 3d representation at scale
Zhou, J.; Wang, J.; Ma, B.; Liu, Y.-S.; Huang, T.; and Wang, X. 2023 · 2023
Cited alongside, same era.
LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding Reasoning and Planning
Chen, S.; Chen, X.; Zhang, C.; Li, M.; Yu, G.; Fei, H.; Zhu, H.; Fan, J.; and Chen, T. 2024a
Cited in the paper.
Closest in time.
GPT-4o mini: advancing cost-efficient intelligence
Jacob, M.; Kevin, L.; Shengjia, Z.; Eric, W.; Hongyu, R.; Haitang, H.; Nick, S.; and Felipe, P. S. 2024 · 2024
Closest in time.
Improved baselines with visual instruction tuning
Liu, H.; Li, C.; Li, Y.; and Lee, Y. J. 2024 · 2024
Closest in time.
View selection for 3d captioning via diffusion ranking
Luo, T.; Johnson, J.; and Lee, H. 2024 · 2024
Closest in time.
Scalable 3d captioning with pretrained models
Luo, T.; Rockwell, C.; Lee, H.; and Johnson, J. 2024 · 2024
Closest in time.
MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors
Tang, Y.; Han, X.; Li, X.; Yu, Q.; Hao, Y.; Hu, L.; and Chen, M. 2024 · 2024
Closest in time.
Ulip-2: Towards scalable multimodal pre-training for 3d understanding
Xue, L.; Yu, N.; Zhang, S.; Panagopoulou, A.; Li, J.; Martín-Martín, R.; Wu, J.; Xiong, C.; Xu, R.; Niebles, J. C.; et al. 2024 · 2024
Closest in time.