Fetching the paper…
Reading the bibliography…
This paper presents ShapeLLM, the first 3D Multimodal Large Language Model (LLM) designed for embodied interaction, exploring a universal 3D object understanding with 3D point clouds and languages.
Kuhn, H.W.: The hungarian method for the assignment problem. Naval research logistics quarterly 2
1955
Earlier work this paper cites.
Boser, B.E., Guyon, I., Vapnik, V.: A training algorithm for optimal margin classifiers. In: ACM Conf. Comput. Learn. Theory (COLT). pp. 144–152. ACM (1992)
1992
Earlier work this paper cites.
Bradski, G., Grossberg, S.: Recognition of 3-d objects from multiple 2-d views by a self-organizing neural architecture. In: From Statistics to Neural Networks: Theory and Pattern Recognition Applications, pp. 349–375. Springer (1994)
1994
Earlier work this paper cites.
Kanade, T., Okutomi, M.: A stereo matching algorithm with an adaptive window: Theory and experiment. IEEE Trans. Pattern Anal. Mach. Intell. 16
1994
Earlier work this paper cites.
Vapnik, V.: Statistical learning theory. Wiley (1998)
1998
Earlier work this paper cites.
Papineni, K., Roukos, S., Ward, T., Zhu, W.: Bleu: a method for automatic evaluation of machine translation (2002)
2002
Earlier work this paper cites.
Lin, C.Y.: Rouge: A package for automatic evaluation of summaries. In: Proc. Workshop on Text Summariation Branches Out, Post-Conference Workshop of ACL 2004 (2004)
2004
Earlier work this paper cites.
Banerjee, S., Lavie, A.: METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In: Proceedings of the Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization@ACL 2005, Ann Arbor, Michigan, USA, June 29, 2005 (2005)
2005
Earlier work this paper cites.
Bronstein, A.M., Bronstein, M.M., Guibas, L.J., Ovsjanikov, M.: Shape google: Geometric words and expressions for invariant shape retrieval. ACM Trans. Graph. 30
2011
Earlier work this paper cites.
Grabner, H., Gall, J., Gool, L.V.: What makes a chair a chair? In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2011)
2011
Earlier work this paper cites.
Gupta, S., Girshick, R.B., Arbeláez, P.A., Malik, J.: Learning rich features from RGB-D images for object detection and segmentation. In: Eur. Conf. Comput. Vis. (ECCV) (2014)
2014
Earlier work this paper cites.
Kim, V.G., Chaudhuri, S., Guibas, L.J., Funkhouser, T.A.: Shape2pose: human-centric shape analysis. ACM Trans. Graph. 33
2014
Earlier work this paper cites.
Zhao, X., Wang, H., Komura, T.: Indexing 3d scenes using the interaction bisector surface. ACM Trans. Graph. 33
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Hu, R., Zhu, C., van Kaick, O., Liu, L., Shamir, A., Zhang, H.: Interaction context (ICON): towards a geometric functionality descriptor. ACM Trans. Graph. 34
2015
Earlier work this paper cites.
Maturana, D., Scherer, S.A.: Voxnet: A 3d convolutional neural network for real-time object recognition. In: IEEE/RSJ Int. Conf. Intell. Robot. and Syst. (IROS). pp. 922–928. IEEE (2015)
2015
Earlier work this paper cites.
Su, H., Maji, S., Kalogerakis, E., Learned-Miller, E.G.: Multi-view convolutional neural networks for 3d shape recognition. In: Int. Conf. Comput. Vis. (ICCV) (2015)
2015
Earlier work this paper cites.
Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., Xiao, J.: 3d shapenets: A deep representation for volumetric shapes. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR). pp. 1912–1920 (2015)
2015
Earlier work this paper cites.
Armeni, I., Sener, O., Zamir, A.R., Jiang, H., Brilakis, I., Fischer, M., Savarese, S.: 3d semantic parsing of large-scale indoor spaces. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2016)
2016
Earlier work this paper cites.
Hendrycks, D., Gimpel, K.: Gaussian error linear units (gelus). CoRR abs/1606.08415
2016
Earlier work this paper cites.
Hu, R., van Kaick, O., Wu, B., Huang, H., Shamir, A., Zhang, H.: Learning how objects function via co-analysis of interactions. ACM Trans. Graph. 35
2016
Earlier work this paper cites.
Yi, L., Kim, V.G., Ceylan, D., Shen, I.C., Yan, M., Su, H., Lu, C., Huang, Q., Sheffer, A., Guibas, L.: A scalable active framework for region annotation in 3d shape collections. ACM Trans. Graph. 35
2016
Earlier work this paper cites.
Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T., Nießner, M.: Scannet: Richly-annotated 3d reconstructions of indoor scenes. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2017)
2017
Earlier work this paper cites.
Fan, H., Su, H., Guibas, L.J.: A point set generation network for 3d object reconstruction from a single image. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2017)
2017
Earlier work this paper cites.
Hu, R., Li, W., van Kaick, O., Shamir, A., Zhang, H., Huang, H.: Learning to predict part mobility from a single static snapshot. ACM Trans. Graph. 36
2017
Earlier work this paper cites.
Loshchilov, I., Hutter, F.: SGDR: stochastic gradient descent with warm restarts. In: Int. Conf. Learn. Represent. (ICLR) (2017)
2017
Earlier work this paper cites.
MacLeod, H., Bennett, C.L., Morris, M.R., Cutrell, E.: Understanding blind people’s experiences with computer-generated captions of social media images. In: Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems. p. 5988–5999. CHI ’17, Association for Computing Machinery, New York, NY, USA (2017)
2017
Earlier work this paper cites.
Pirk, S., Krs, V., Hu, K., Rajasekaran, S.D., Kang, H., Yoshiyasu, Y., Benes, B., Guibas, L.J.: Understanding and exploiting object interaction landscapes. ACM Trans. Graph. 36
2017
Earlier work this paper cites.
Qi, C.R., Su, H., Mo, K., Guibas, L.J.: Pointnet: Deep learning on point sets for 3d classification and segmentation. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR). pp. 77–85 (2017)
2017
Earlier work this paper cites.
Qi, C.R., Yi, L., Su, H., Guibas, L.J.: Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In: Adv. Neural Inform. Process. Syst. (NIPS). pp. 5099–5108 (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: Adv. Neural Inform. Process. Syst. (NIPS). pp. 5998–6008 (2017)
2017
Earlier work this paper cites.
Yi, L., Su, H., Guo, X., Guibas, L.J.: Syncspeccnn: Synchronized spectral CNN for 3d shape segmentation. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2017)
2017
Earlier work this paper cites.
Achlioptas, P., Diamanti, O., Mitliagkas, I., Guibas, L.J.: Learning representations and generative models for 3d point clouds. In: Int. Conf. Mach. Learn. (ICML) (2018)
2018
Earlier work this paper cites.
Lu, C., Su, H., Li, Y., Lu, Y., Yi, L., Tang, C., Guibas, L.J.: Beyond holistic object recognition: Enriching image understanding with part states. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2018)
2018
Earlier work this paper cites.
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I.: Improving language understanding by generative pre-training (2018)
2018
Earlier work this paper cites.
Rohrbach, A., Hendricks, L.A., Burns, K., Darrell, T., Saenko, K.: Object hallucination in image captioning. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 (2018)
2018
Earlier work this paper cites.
Rohrbach, A., Hendricks, L.A., Burns, K., Darrell, T., Saenko, K.: Object hallucination in image captioning. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (2018)
2018
Earlier work this paper cites.
Yi, L., Huang, H., Liu, D., Kalogerakis, E., Su, H., Guibas, L.J.: Deep part induction from articulated object pairs. ACM Trans. Graph. 37
2018
Earlier work this paper cites.
Das, A., Kottur, S., Gupta, K., Singh, A., Yadav, D., Lee, S., Moura, J.M.F., Parikh, D., Batra, D.: Visual dialog. IEEE Trans. Pattern Anal. Mach. Intell. (TPAMI) 41
2019
Earlier work this paper cites.
Davison, J., Feldman, J., Rush, A.M.: Commonsense knowledge mining from pretrained models. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019 (2019)
2019
Earlier work this paper cites.
Devlin, J., Chang, M., Lee, K., Toutanova, K.: BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers). pp. 4171–4186. Association for Computational Linguistics (2019)
2019
Earlier work this paper cites.
Goyal, P., Mahajan, D., Gupta, A., Misra, I.: Scaling and benchmarking self-supervised visual representation learning. In: Int. Conf. Comput. Vis. (ICCV). pp. 6390–6399. IEEE (2019)
2019
Earlier work this paper cites.
Gupta, A., Dollar, P., Girshick, R.: Lvis: A dataset for large vocabulary instance segmentation. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2019)
2019
Earlier work this paper cites.
Liu, Y., Fan, B., Xiang, S., Pan, C.: Relation-shape convolutional neural network for point cloud analysis. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2019)
2019
Earlier work this paper cites.
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: Int. Conf. Learn. Represent. (ICLR) (2019)
2019
Earlier work this paper cites.
Mo, K., Zhu, S., Chang, A.X., Yi, L., Tripathi, S., Guibas, L.J., Su, H.: Partnet: A large-scale benchmark for fine-grained and hierarchical part-level 3d object understanding. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2019)
2019
Earlier work this paper cites.
Petroni, F., Rocktäschel, T., Riedel, S., Lewis, P.S.H., Bakhtin, A., Wu, Y., Miller, A.H.: Language models as knowledge bases? In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019 (2019)
2019
Earlier work this paper cites.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I.: Language models are unsupervised multitask learners. OpenAI blog 1
2019
Earlier work this paper cites.
Reimers, N., Gurevych, I.: Sentence-bert: Sentence embeddings using siamese bert-networks. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019 (2019)
2019
Earlier work this paper cites.
Uy, M.A., Pham, Q.H., Hua, B.S., Nguyen, T., Yeung, S.K.: Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR). pp. 1588–1597 (2019)
2019
Earlier work this paper cites.
Wang, H., Sridhar, S., Huang, J., Valentin, J., Song, S., Guibas, L.J.: Normalized object coordinate space for category-level 6d object pose and size estimation. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2019)
2019
Earlier work this paper cites.
Wang, Y., Sun, Y., Liu, Z., Sarma, S.E., Bronstein, M.M., Solomon, J.M.: Dynamic graph CNN for learning on point clouds. ACM Trans. Graph. 38
2019
Earlier work this paper cites.
Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D.M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., Amodei, D.: Language models are few-shot learners. In: Adv. Neural Inform. Process. Syst. (NeurIPS) (2020)
2020
Earlier work this paper cites.
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers. In: Eur. Conf. Comput. Vis. (ECCV) (2020)
2020
Earlier work this paper cites.
Chen, D.Z., Chang, A.X., Nießner, M.: Scanrefer: 3d object localization in RGB-D scans using natural language. In: Eur. Conf. Comput. Vis. (ECCV) (2020)
2020
Earlier work this paper cites.
Huang, W., Mordatch, I., Pathak, D.: One policy to control them all: Shared modular policies for agent-agnostic control. In: Int. Conf. Mach. Learn. (ICML) (2020)
2020
Earlier work this paper cites.
Jiang, Z., Xu, F.F., Araki, J., Neubig, G.: How can we know what language models know. Trans. Assoc. Comput. Linguistics 8
2020
Earlier work this paper cites.
Li, X., Wang, H., Yi, L., Guibas, L.J., Abbott, A.L., Song, S.: Category-level articulated object pose estimation. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2020)
2020
Earlier work this paper cites.
Xie, S., Gu, J., Guo, D., Qi, C.R., Guibas, L.J., Litany, O.: Pointcontrast: Unsupervised pre-training for 3d point cloud understanding. In: Eur. Conf. Comput. Vis. (ECCV). Lecture Notes in Computer Science, vol. 12348, pp. 574–591. Springer (2020)
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Chen, D.Z., Gholami, A., Nießner, M., Chang, A.X.: Scan2cap: Context-aware dense captioning in RGB-D scans. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2021)
2021
Earlier work this paper cites.
Deng, J., Shi, S., Li, P., Zhou, W., Zhang, Y., Li, H.: Voxel r-cnn: Towards high performance voxel-based 3d object detection. In: AAAI Conf. Artif. Intell. (AAAI) (2021)
2021
Earlier work this paper cites.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An image is worth 16x16 words: Transformers for image recognition at scale. In: Int. Conf. Learn. Represent. (ICLR) (2021)
2021
Earlier work this paper cites.
Engel, N., Belagiannis, V., Dietmayer, K.: Point transformer. IEEE Access 9
2021
Earlier work this paper cites.
Fu, H., Jia, R., Gao, L., Gong, M., Zhao, B., Maybank, S., Tao, D.: 3d-future: 3d furniture shape with texture. International Journal of Computer Vision 129
2021
Earlier work this paper cites.
Gao, T., Yao, X., Chen, D.: Simcse: Simple contrastive learning of sentence embeddings. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021 (2021)
2021
Earlier work this paper cites.
Hamdi, A., Giancola, S., Ghanem, B.: MVTN: multi-view transformation network for 3d shape recognition. In: Int. Conf. Comput. Vis. (ICCV). pp. 1–11. IEEE (2021)
2021
Earlier work this paper cites.
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., Neubig, G.: Towards a unified view of parameter-efficient transfer learning. In: Int. Conf. Learn. Represent. (ICLR) (2021)
2021
Earlier work this paper cites.
Hou, J., Xie, S., Graham, B., Dai, A., Nießner, M.: Pri3d: Can 3d priors help 2d representation learning? In: Int. Conf. Comput. Vis. (ICCV). pp. 5673–5682. IEEE (2021)
2021
Earlier work this paper cites.
Ilharco, G., Wortsman, M., Wightman, R., Gordon, C., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., Schmidt, L.: Openclip (Jul 2021)
2021
Cited alongside, same era.
Li, X.L., Liang, P.: Prefix-tuning: Optimizing continuous prompts for generation. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (2021)
2021
Cited alongside, same era.
Liu, Z., Zhang, Z., Cao, Y., Hu, H., Tong, X.: Group-free 3d object detection via transformers. In: Int. Conf. Comput. Vis. (ICCV) (2021)
2021
Cited alongside, same era.
Mao, J., Xue, Y., Niu, M., Bai, H., Feng, J., Liang, X., Xu, H., Xu, C.: Voxel transformer for 3d object detection. In: Int. Conf. Comput. Vis. (ICCV) (2021)
2021
Cited alongside, same era.
OpenAI: Gpt-4v(ision) system card (2023), https://openai.com/research/gpt-4v-system-card
2023
Later among the works it cites.
Peng, B., Li, C., He, P., Galley, M., Gao, J.: Instruction tuning with GPT-4. CoRR abs/2304.03277
2023
Later among the works it cites.
Peng, S., Genova, K., Jiang, C.M., Tagliasacchi, A., Pollefeys, M., Funkhouser, T.A.: Openscene: 3d scene understanding with open vocabularies. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Qi, H., Kumar, A., Calandra, R., Ma, Y., Malik, J.: In-hand object rotation via rapid motor adaptation. In: Annu. Conf. Robot. Learn. (CoRL) (2023)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: Int. Conf. Mach. Learn. (ICML). Proceedings of Machine Learning Research, vol. 139, pp. 8748–8763. PMLR (2021)
2021
Cited alongside, same era.
Weng, Y., Wang, H., Zhou, Q., Qin, Y., Duan, Y., Fan, Q., Chen, B., Su, H., Guibas, L.J.: CAPTRA: category-level pose tracking for rigid and articulated objects from point clouds. In: Int. Conf. Comput. Vis. (ICCV) (2021)
2021
Cited alongside, same era.
Alayrac, J., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., Simonyan, K.: Flamingo: a visual language model for few-shot learning. In: Adv. Neural Inform. Process. Syst. (NeurIPS) (2022)
2022
Cited alongside, same era.
Collins, J., Goel, S., Deng, K., Luthra, A., Xu, L., Gundogdu, E., Zhang, X., Vicente, T.F.Y., Dideriksen, T., Arora, H., et al.: Abo: Dataset and benchmarks for real-world 3d object understanding. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2022)
2022
Cited alongside, same era.
Dao, T., Fu, D., Ermon, S., Rudra, A., Ré, C.: Flashattention: Fast and memory-efficient exact attention with io-awareness. In: Adv. Neural Inform. Process. Syst. (NeurIPS) (2022)
2022
Cited alongside, same era.
Dong, R., Tan, Z., Wu, M., Zhang, L., Ma, K.: Finding the task-optimal low-bit sub-distribution in deep neural networks. In: Int. Conf. Mach. Learn. (ICML) (2022)
2022
Cited alongside, same era.
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., Girshick, R.B.: Masked autoencoders are scalable vision learners. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2022)
2022
Cited alongside, same era.
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: Lora: Low-rank adaptation of large language models. In: Int. Conf. Learn. Represent. (ICLR) (2022)
2022
Cited alongside, same era.
2023
Later among the works it cites.
Qi, Z., Dong, R., Fan, G., Ge, Z., Zhang, X., Ma, K., Yi, L.: Contrast with reconstruct: Contrastive 3d representation learning guided by generative pretraining. In: Int. Conf. Mach. Learn. (ICML) (2023)
2023
Later among the works it cites.
Qi, Z., Yu, M., Dong, R., Ma, K.: VPP: efficient conditional 3d generation via voxel-point progressive representation. In: Adv. Neural Inform. Process. Syst. (NeurIPS) (2023)
2023
Later among the works it cites.
Shen, W., Yang, G., Yu, A., Wong, J., Kaelbling, L.P., Isola, P.: Distilled feature fields enable few-shot language-guided manipulation. In: Annu. Conf. Robot. Learn. (CoRL) (2023)
2023
Later among the works it cites.
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., Zhuang, Y.: Hugginggpt: Solving AI tasks with chatgpt and its friends in huggingface. In: Adv. Neural Inform. Process. Syst. (NeurIPS) (2023)
2023
Later among the works it cites.
Shi, H., Xu, H., Clarke, S., Li, Y., Wu, J.: Robocook: Long-horizon elasto-plastic object manipulation with diverse tools. In: Annu. Conf. Robot. Learn. (CoRL) (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Surís, D., Menon, S., Vondrick, C.: Vipergpt: Visual inference via python execution for reasoning. In: Int. Conf. Comput. Vis. (ICCV) (2023)
2023
Later among the works it cites.
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., Hashimoto, T.B.: Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Wan, W., Geng, H., Liu, Y., Shan, Z., Yang, Y., Yi, L., Wang, H.: Unidexgrasp++: Improving dexterous grasping policy learning via geometry-aware curriculum and iterative generalist-specialist learning. In: Int. Conf. Comput. Vis. (ICCV) (2023)
2023
Later among the works it cites.
Wang, Z., Yu, X., Rao, Y., Zhou, J., Lu, J.: Take-a-photo: 3d-to-2d generative pre-training of point cloud models. In: Int. Conf. Comput. Vis. (ICCV) (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Xu, Y., Wan, W., Zhang, J., Liu, H., Shan, Z., Shen, H., Wang, R., Geng, H., Weng, Y., Chen, J., Liu, T., Yi, L., Wang, H.: Unidexgrasp: Universal robotic dexterous grasping via learning diverse proposal generation and goal-conditioned policy. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2023)
2023
Later among the works it cites.
Xu, Z., Shen, Y., Huang, L.: Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL) (Volume 1: Long Papers) (2023)
2023
Later among the works it cites.
Xue, L., Gao, M., Xing, C., Martín-Martín, R., Wu, J., Xiong, C., Xu, R., Niebles, J.C., Savarese, S.: ULIP: learning unified representation of language, image and point cloud for 3d understanding. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2023)
2023
Later among the works it cites.
Yang, R., Song, L., Li, Y., Zhao, S., Ge, Y., Li, X., Shan, Y.: Gpt4tools: Teaching large language model to use tools via self-instruction. In: Adv. Neural Inform. Process. Syst. (NeurIPS) (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Zeid, K.A., Schult, J., Hermans, A., Leibe, B.: Point2vec for self-supervised representation learning on point clouds. In: DAGM German Conference on Pattern Recognition. pp. 131–146. Springer (2023)
2023
Later among the works it cites.
Zhang, J., Dong, R., Ma, K.: CLIP-FO3D: learning free open-world 3d scene representations from 2d dense CLIP. In: Int. Conf. Comput. Vis. Worksh. (ICCV Workshop) (2023)
2023
Later among the works it cites.
Zhang, L., Chen, X., Dong, R., Ma, K.: Region-aware knowledge distillation for efficient image-to-image translation. In: Brit. Mach. Vis. Conf. (BMVC) (2023)
2023
Later among the works it cites.
Zhang, L., Dong, R., Tai, H., Ma, K.: Pointdistiller: Structured knowledge distillation towards efficient and compact 3d detection. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2023)
2023
Later among the works it cites.
Zhang, R., Wang, L., Qiao, Y., Gao, P., Li, H.: Learning 3d representations from 2d pre-trained models via image-to-point masked autoencoders. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Zheng, J., Zheng, Q., Fang, L., Liu, Y., Yi, L.: CAMS: canonicalized manipulation spaces for category-level functional hand-object manipulation synthesis. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2023)
2023
Later among the works it cites.
Zhu, X., Zhang, R., He, B., Guo, Z., Zeng, Z., Qin, Z., Zhang, S., Gao, P.: Pointclip v2: Prompting clip and gpt for powerful 3d open-world learning. In: Int. Conf. Comput. Vis. (ICCV) (2023)
2023
Later among the works it cites.
Zhu, Z., Ma, X., Chen, Y., Deng, Z., Huang, S., Li, Q.: 3d-vista: Pre-trained transformer for 3d vision and text alignment. In: Int. Conf. Comput. Vis. (ICCV) (2023)
2023
Later among the works it cites.
Bai, Y., Geng, X., Mangalam, K., Bar, A., Yuille, A.L., Darrell, T., Malik, J., Efros, A.A.: Sequential modeling enables scalable learning for large vision models. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2024)
2024
Closest in time.
Chang, M., Gervet, T., Khanna, M., Yenamandra, S., Shah, D., Min, S.Y., Shah, K., Paxton, C., Gupta, S., Batra, D., Mottaghi, R., Malik, J., Chaplot, D.S.: GOAT: GO to any thing. In: Robotics: Science and Systems (RSS) (2024)
2024
Closest in time.
Chen, B., Xu, Z., Kirmani, S., Ichter, B., Driess, D., Florence, P., Sadigh, D., Guibas, L., Xia, F.: Spatialvlm: Endowing vision-language models with spatial reasoning capabilities. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2024)
2024
Closest in time.
Chen, S., Garcia, R., Laptev, I., Schmid, C.: Sugar: Pre-training 3d visual representations for robotics. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 18049–18060 (2024)
2024
Closest in time.
Ding, R., Yang, J., Xue, C., Zhang, W., Bai, S., Qi, X.: Lowis3d: Language-driven open-world instance-level 3d scene understanding. IEEE Trans. Pattern Anal. Mach. Intell. (TPAMI) pp. 1–16 (2024)
2024
Closest in time.
Dong, R., Han, C., Peng, Y., Qi, Z., Ge, Z., Yang, J., Zhao, L., Sun, J., Zhou, H., Wei, H., Kong, X., Zhang, X., Ma, K., Yi, L.: DreamLLM: Synergistic multimodal comprehension and creation. In: Int. Conf. Learn. Represent. (ICLR) (2024)
2024
Closest in time.
Gao, Y., Wang, Z., Zheng, W.S., Xie, C., Zhou, Y.: Sculpting holistic 3d representation in contrastive language-image-3d pre-training. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2024)
2024
Closest in time.
Ge, Y., Ge, Y., Zeng, Z., Wang, X., Shan, Y.: Planting a SEED of vision in large language model. In: Int. Conf. Learn. Represent. (ICLR) (2024)
2024
Closest in time.
Geng, H., Wei, S., Deng, C., Shen, B., Wang, H., Guibas, L.: Sage: Bridging semantic and actionable parts for generalizable articulated-object manipulation under language instructions. In: Robotics: Science and Systems (RSS) (2024)
2024
Closest in time.
Gunjal, A., Yin, J., Bas, E.: Detecting and preventing hallucinations in large vision language models. In: AAAI Conf. Artif. Intell. (AAAI) (2024)
2024
Closest in time.
Huang, J., Yong, S., Ma, X., Linghu, X., Li, P., Wang, Y., Li, Q., Zhu, S., Jia, B., Huang, S.: An embodied generalist agent in 3d world. In: Int. Conf. Mach. Learn. (ICML) (2024)
2024
Closest in time.
Jaiswal, A., Gan, Z., Du, X., Zhang, B., Wang, Z., Yang, Y.: Compressing llms: The truth is rarely pure and never simple. In: Int. Conf. Learn. Represent. (ICLR) (2024)
2024
Closest in time.
Liang, Y., Wu, C., Song, T., Wu, W., Xia, Y., Liu, Y., Ou, Y., Lu, S., Ji, L., Mao, S., et al.: Taskmatrix. ai: Completing tasks by connecting foundation models with millions of apis. Intelligent Computing 3
2024
Closest in time.
Liu, H., Li, C., Li, Y., Lee, Y.J.: Improved baselines with visual instruction tuning. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2024)
2024
Closest in time.
Liu, X., Yi, L.: GeneOH diffusion: Towards generalizable hand-object interaction denoising via denoising diffusion. In: Int. Conf. Learn. Represent. (ICLR) (2024)
2024
Closest in time.
Liu, Y., Lin, C., Zeng, Z., Long, X., Liu, L., Komura, T., Wang, W.: Syncdreamer: Generating multiview-consistent images from a single-view image. In: Int. Conf. Learn. Represent. (ICLR) (2024)
2024
Closest in time.
OpenAI: Introducing gpt-4o and more tools to chatgpt free users (2024), https://openai.com/index/gpt-4o-and-more-tools-to-chatgpt-free/
2024
Closest in time.
Pan, X., Dong, L., Huang, S., Peng, Z., Chen, W., Wei, F.: Kosmos-g: Generating images in context with multimodal large language models. In: Int. Conf. Learn. Represent. (ICLR) (2024)
2024
Closest in time.
2024
Closest in time.
Sun, Q., Cui, Y., Zhang, X., Zhang, F., Yu, Q., Luo, Z., Wang, Y., Rao, Y., Liu, J., Huang, T., Wang, X.: Generative multimodal models are in-context learners. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2024)
2024
Closest in time.
Sun, Q., Yu, Q., Cui, Y., Zhang, F., Zhang, X., Wang, Y., Gao, H., Liu, J., Huang, T., Wang, X.: Emu: Generative pretraining in multimodality. In: Int. Conf. Learn. Represent. (ICLR) (2024)
2024
Closest in time.
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., Anandkumar, A.: Voyager: An open-ended embodied agent with large language models. T. Mach. Learn. Res. (TMLR) (2024)
2024
Closest in time.
Wu, S., Fei, H., Qu, L., Ji, W., Chua, T.: Next-gpt: Any-to-any multimodal LLM. In: Int. Conf. Mach. Learn. (ICML) (2024)
2024
Closest in time.
Wu, T., Yang, G., Li, Z., Zhang, K., Liu, Z., Guibas, L.J., Lin, D., Wetzstein, G.: Gpt-4v(ision) is a human-aligned evaluator for text-to-3d generation. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2024)
2024
Closest in time.
Xue, L., Yu, N., Zhang, S., Li, J., Martín-Martín, R., Wu, J., Xiong, C., Xu, R., Niebles, J.C., Savarese, S.: ULIP-2: towards scalable multimodal pre-training for 3d understanding. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2024)
2024
Closest in time.
Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., Wang, L.: Mm-vet: Evaluating large multimodal models for integrated capabilities. In: Int. Conf. Mach. Learn. (ICML) (2024)
2024
Closest in time.
Zhang, R., Han, J., Zhou, A., Hu, X., Yan, S., Lu, P., Li, H., Gao, P., Qiao, Y.: Llama-adapter: Efficient fine-tuning of language models with zero-init attention. In: Int. Conf. Learn. Represent. (ICLR) (2024)
2024
Closest in time.
Zhang, Z., Cao, S., Wang, Y.: TAMM: triadapter multi-modal learning for 3d shape understanding. In: IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR) (2024)
2024
Closest in time.
Zhao, L., Yu, E., Ge, Z., Yang, J., Wei, H., Zhou, H., Sun, J., Peng, Y., Dong, R., Han, C., Zhang, X.: Chatspot: Bootstrapping multimodal llms via precise referring instruction tuning. In: Int. Joint Conf. Artif. Intell. (IJCAI) (2024)
2024
Closest in time.
Zheng, L., Chiang, W., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E.P., Zhang, H., Gonzalez, J.E., Stoica, I.: Judging llm-as-a-judge with mt-bench and chatbot arena. In: Adv. Neural Inform. Process. Syst. (NeurIPS) (2024)
2024
Closest in time.
Zhou, J., Wang, J., Ma, B., Liu, Y., Huang, T., Wang, X.: Uni3d: Exploring unified 3d representation at scale. In: Int. Conf. Learn. Represent. (ICLR) (2024)
2024
Closest in time.
Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., Yao, H.: Analyzing and mitigating object hallucination in large vision-language models. In: Int. Conf. Learn. Represent. (ICLR) (2024)
2024
Closest in time.
Zhu, D., Chen, J., Shen, X., Li, X., Elhoseiny, M.: Minigpt-4: Enhancing vision-language understanding with advanced large language models. In: Int. Conf. Learn. Represent. (ICLR) (2024)
2024
Closest in time.