Fetching the paper…
Reading the bibliography…
We introduce Duoduo CLIP, a model for 3D representation learning that learns shape encodings from multi-view images instead of point clouds.
Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data
Mikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Thanh Nguyen, and Sai-Kit Yeung · 1908
Earlier work this paper cites.
Recognition of 3-D objects from multiple 2-D views by a self-organizing neural architecture
Gary Bradski and Stephen Grossberg · 1994
Earlier work this paper cites.
3D model search engine based on lightfield descriptors
Yu-Te Shen, Ding-Yun Chen, Xiao-Pei Tian, and Ming Ouhyoung · 2003
Earlier work this paper cites.
Quantifying attention flow in transformers
Samira Abnar and Willem Zuidema · 2005
Earlier work this paper cites.
PointContrast: Unsupervised pre-training for 3D point cloud understanding
Saining Xie, Jiatao Gu, Demi Guo, Charles R Qi, Leonidas Guibas, and Or Litany · 2007
Earlier work this paper cites.
3D-FUTURE: 3D furniture shape with texture
Huan Fu, Rongfei Jia, Lin Gao, Mingming Gong, Binqiang Zhao, Steve Maybank, and Dacheng Tao · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
Contrastive learning of medical visual representations from paired images and text
Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D Manning, and Curtis P Langlotz · 2010
Earlier work this paper cites.
MVTN: Multi-view transformation network for 3D shape recognition
Abdullah Hamdi, Silvio Giancola, and Bernard Ghanem · 2011
Earlier work this paper cites.
WSABIE: Scaling up to large vocabulary image annotation
Jason Weston, Samy Bengio, and Nicolas Usunier · 2011
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Earlier work this paper cites.
ShapeNet: An information-rich 3D model repository
Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al · 2015
Earlier work this paper cites.
Multi-view convolutional neural networks for 3D shape recognition
Hang Su, Subhransu Maji, Evangelos Kalogerakis, and Erik Learned-Miller · 2015
Earlier work this paper cites.
SceneNN: A scene meshes dataset with annotations
Binh-Son Hua, Quang-Hieu Pham, Duc Thanh Nguyen, Minh-Khoi Tran, Lap-Fai Yu, and Sai-Kit Yeung · 2016
Earlier work this paper cites.
ScanNet: Richly-annotated 3D reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
PointNet: Deep learning on point sets for 3D classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
VSE++: Improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Text2shape: Generating shapes from natural language by learning joint embeddings
Kevin Chen, Christopher B Choy, Manolis Savva, Angel X Chang, Thomas Funkhouser, and Silvio Savarese · 2019
Earlier work this paper cites.
Learning implicit fields for generative shape modeling
Zhiqin Chen and Hao Zhang · 2019
Earlier work this paper cites.
Y2Seq2Seq: Cross-modal representation learning for 3D shape and text by joint reconstruction and prediction of view and word sequences
Zhizhong Han, Mingyang Shang, Xiyang Wang, Yu-Shen Liu, and Matthias Zwicker · 2019
Earlier work this paper cites.
View-GCN: View-based graph convolutional network for 3D shape analysis
Xin Wei, Ruixuan Yu, and Jian Sun · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Cited alongside, same era.
OpenCLIP, July 2021
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Contrast with reconstruct: Contrastive 3D representation learning guided by generative pretraining
Zekun Qi, Runpei Dong, Guofan Fan, Zheng Ge, Xiangyu Zhang, Kaisheng Ma, and Li Yi · 2023
Later among the works it cites.
MVDream: Multi-view diffusion for 3D generation
Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang · 2023
Later among the works it cites.
MV-CLIP: Multi-view CLIP for zero-shot 3D shape recognition
Dan Song, Xinwei Fu, Weizhi Nie, Wenhui Li, and Anan Liu · 2023
Later among the works it cites.
Parts2words: Learning joint embedding of point clouds and texts by bidirectional matching between parts and words
Chuan Tang, Xi Yang, Bojian Wu, Zhizhong Han, and Yi Chang · 2023
Later among the works it cites.
MXM-CLR: A unified framework for contrastive learning of multifold cross-modal representations
Ye Wang, Bowei Jiang, Changqing Zou, and Rui Ma · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ABO: Dataset and benchmarks for real-world 3D object understanding
Jasmine Collins, Shubham Goel, Kenan Deng, Achleshwar Luthra, Leon Xu, Erhan Gundogdu, Xi Zhang, Tomas F Yago Vicente, Thomas Dideriksen, Himanshu Arora, et al · 2022
Cited alongside, same era.
CyCLIP: Cyclic contrastive language-image pretraining
Shashank Goel, Hritik Bansal, Sumit Bhatia, Ryan Rossi, Vishwa Vinay, and Aditya Grover · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Cited alongside, same era.
SLIP: Self-supervision meets language-image pre-training
Norman Mu, Alexander Kirillov, David Wagner, and Saining Xie · 2022
Cited alongside, same era.
PointNeXt: Revisiting PointNet++ with improved training and scaling strategies
Guocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai, Hasan Hammoud, Mohamed Elhoseiny, and Bernard Ghanem · 2022
Cited alongside, same era.
CLIPort: What and where pathways for robotic manipulation
Mohit Shridhar, Lucas Manuelli, and Dieter Fox · 2022
Cited alongside, same era.
Language grounding with 3D objects
Jesse Thomason, Mohit Shridhar, Yonatan Bisk, Chris Paxton, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Later among the works it cites.
ULIP: Learning a unified representation of language, images, and point clouds for 3D understanding
Le Xue, Mingfei Gao, Chen Xing, Roberto Martín-Martín, Jiajun Wu, Caiming Xiong, Ran Xu, Juan Carlos Niebles, and Silvio Savarese · 2023
Later among the works it cites.
MVImgNet: A large-scale dataset of multi-view images
Xianggang Yu, Mutian Xu, Yidan Zhang, Haolin Liu, Chongjie Ye, Yushuang Wu, Zizheng Yan, Chenming Zhu, Zhangyang Xiong, Tianyou Liang, et al · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Later among the works it cites.
PointCLIP v2: Prompting CLIP and GPT for powerful 3D open-world learning
Xiangyang Zhu, Renrui Zhang, Bowei He, Ziyu Guo, Ziyao Zeng, Zipeng Qin, Shanghang Zhang, and Peng Gao · 2023
Later among the works it cites.
Probing the 3D awareness of visual foundation models
Mohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis, Abhishek Kar, Yuanzhen Li, Michael Rubinstein, Deqing Sun, Leonidas Guibas, Justin Johnson, and Varun Jampani · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al · 2024
Closest in time.
Sculpting holistic 3D representation in contrastive language-image-3D pre-training
Yipeng Gao, Zeyu Wang, Wei-Shi Zheng, Cihang Xie, and Yuyin Zhou · 2024
Closest in time.
RoboHop: Segment-based topological map representation for open-world visual navigation
Sourav Garg, Krishan Rana, Mehdi Hosseinzadeh, Lachlan Mares, Niko Sünderhauf, Feras Dayoub, and Ian Reid · 2024
Closest in time.
VIT-lens: Towards omni-modal representations
Weixian Lei, Yixiao Ge, Jianfeng Zhang, Dylan Sun, Kun Yi, Ying Shan, and Mike Zheng Shou · 2024
Closest in time.
PEVA-net: Prompt-enhanced view aggregation network for zero/few-shot multi-view 3D shape recognition
Dongyun Lin, Yi Cheng, Shangbo Mao, Aiyuan Guo, and Yiqun Li · 2024
Closest in time.
Scalable 3D captioning with pretrained models
Tiange Luo, Chris Rockwell, Honglak Lee, and Justin Johnson · 2024
Closest in time.
ShapeLLM: Universal 3D object understanding for embodied interaction
Zekun Qi, Runpei Dong, Shaochen Zhang, Haoran Geng, Chunrui Han, Zheng Ge, Li Yi, and Kaisheng Ma · 2024
Closest in time.
TriCoLo: Trimodal contrastive loss for text to shape retrieval
Yue Ruan, Han-Hung Lee, Yiming Zhang, Ke Zhang, and Angel X Chang · 2024
Closest in time.
Find what you want: learning demand-conditioned object attribute space for demand-driven navigation
Hongcheng Wang, Andy Guan Hong Chen, Xiaoqi Li, Mingdong Wu, and Hao Dong · 2024
Closest in time.
ULIP-2: Towards scalable multimodal pre-training for 3D understanding
Le Xue, Ning Yu, Shu Zhang, Junnan Li, Roberto Martín-Martín, Jiajun Wu, Caiming Xiong, Ran Xu, Juan Carlos Niebles, and Silvio Savarese · 2024
Closest in time.
TAMM: Triadapter multi-modal learning for 3D shape understanding
Zhihao Zhang, Shengcao Cao, and Yu-Xiong Wang · 2024
Closest in time.
Uni3D: Exploring unified 3D representation at scale
Junsheng Zhou, Jinsheng Wang, Baorui Ma, Yu-Shen Liu, Tiejun Huang, and Xinlong Wang · 2024
Closest in time.