Fetching the paper…
Reading the bibliography…
The pretrain-finetune paradigm has achieved great success in NLP and 2D image fields because of the high-quality representation ability and transferability of their pretrained models.
A short note on the kinetics-700 human action dataset
Carreira, J.; Noland, E.; Hillier, C.; and Zisserman, A. 2019 · 1907
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Wu, Z.; Song, S.; Khosla, A.; Yu, F.; Zhang, L.; Tang, X.; and Xiao, J. 2015 · 1920
Earlier work this paper cites.
Elementary differential geometry
Pressley, A. N. 2010 · 2010
Earlier work this paper cites.
ShapeNet: An Information-Rich 3D Model Repository
Chang, A. X.; Funkhouser, T.; Guibas, L.; Hanrahan, P.; Huang, Q.; Li, Z.; Savarese, S.; Savva, M.; Song, S.; Su, H.; Xiao, J.; Yi, L.; and Yu, F. 2015 · 2015
Earlier work this paper cites.
Predicting deep zero-shot convolutional neural networks using textual descriptions
Lei Ba, J.; Swersky, K.; Fidler, S.; et al. 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Vinyals, O.; Toshev, A.; Bengio, S.; and Erhan, D. 2015 · 2015
Earlier work this paper cites.
3d semantic parsing of large-scale indoor spaces
Armeni, I.; Sener, O.; Zamir, A. R.; Jiang, H.; Brilakis, I.; Fischer, M.; and Savarese, S. 2016 · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Dai, A.; Chang, A. X.; Savva, M.; Halber, M.; Funkhouser, T.; and Nießner, M. 2017 · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Qi, C. R.; Su, H.; Mo, K.; and Guibas, L. J. 2017 · 2017
Earlier work this paper cites.
Octnet: Learning deep 3d representations at high resolutions
Riegler, G.; Osman Ulusoy, A.; and Geiger, A. 2017 · 2017
Earlier work this paper cites.
Semantickitti: A dataset for semantic scene understanding of lidar sequences
Behley, J.; Garbade, M.; Milioto, A.; Quenzel, J.; Behnke, S.; Stachniss, C.; and Gall, J. 2019 · 2019
Earlier work this paper cites.
4d spatio-temporal convnets: Minkowski convolutional neural networks
Choy, C.; Gwak, J.; and Savarese, S. 2019 · 2019
Earlier work this paper cites.
Do better imagenet models transfer better?
Kornblith, S.; Shlens, J.; and Le, Q. V. 2019 · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
Recht, B.; Roelofs, R.; Schmidt, L.; and Shankar, V. 2019 · 2019
Earlier work this paper cites.
Kpconv: Flexible and deformable convolution for point clouds
Thomas, H.; Qi, C. R.; Deschaud, J.-E.; Marcotegui, B.; Goulette, F.; and Guibas, L. J. 2019 · 2019
Earlier work this paper cites.
Dynamic graph cnn for learning on point clouds
Wang, Y.; Sun, Y.; Liu, Z.; Sarma, S. E.; Bronstein, M. M.; and Solomon, J. M. 2019 · 2019
Earlier work this paper cites.
Pointcontrast: Unsupervised pre-training for 3d point cloud understanding
Xie, S.; Gu, J.; Guo, D.; Qi, C. R.; Guibas, L.; and Litany, O. 2020 · 2020
Earlier work this paper cites.
Emerging Properties in Self-Supervised Vision Transformers
Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; and Joulin, A. 2021 · 2021
Earlier work this paper cites.
Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers
Chefer, H.; Gur, S.; and Wolf, L. 2021 · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
Ppt: Pre-trained prompt tuning for few-shot learning
Gu, Y.; Han, X.; Liu, Z.; and Huang, M. 2021 · 2021
Cited alongside, same era.
Pct: Point cloud transformer
Guo, M.-H.; Cai, J.-X.; Liu, Z.-N.; Mu, T.-J.; Martin, R. R.; and Hu, S.-M. 2021 · 2021
Cited alongside, same era.
An end-to-end transformer model for 3d object detection
Misra, I.; Girdhar, R.; and Joulin, A. 2021 · 2021
Cited alongside, same era.
Jia, M.; Tang, L.; Chen, B.-C.; Cardie, C.; Belongie, S.; Hariharan, B.; and Lim, S.-N. 2022 · 2022
Closest in time.
Frozen CLIP Models are Efficient Video Learners
Lin, Z.; Geng, S.; Zhang, R.; Gao, P.; de Melo, G.; Wang, X.; Dai, J.; Qiao, Y.; and Li, H. 2022 · 2022
Closest in time.
Masked Discrimination for Self-Supervised Learning on Point Clouds
Liu, H.; Cai, M.; and Lee, Y. J. 2022 · 2022
Closest in time.
Unsupervised Point Cloud Pre-Training Via Contrasting and Clustering
Mei, G.; Huang, X.; Liu, J.; Zhang, J.; and Wu, Q. 2022 · 2022
Closest in time.
Masked autoencoders for point cloud self-supervised learning
Pang, Y.; Wang, W.; Tay, F. E.; Liu, W.; Tian, Y.; and Yuan, L. 2022 · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mokady, R.; Hertz, A.; and Bermano, A. H. 2021 · 2021
Cited alongside, same era.
Intriguing properties of vision transformers
Naseer, M. M.; Ranasinghe, K.; Khan, S. H.; Hayat, M.; Shahbaz Khan, F.; and Yang, M.-H. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
How Much Can CLIP Benefit Vision-and-Language Tasks?
Shen, S.; Li, L. H.; Tan, H.; Bansal, M.; Rohrbach, A.; Chang, K.-W.; Yao, Z.; and Keutzer, K. 2021 · 2021
Cited alongside, same era.
Unsupervised point cloud pre-training via occlusion completion
Wang, H.; Liu, Q.; Yue, X.; Lasenby, J.; and Kusner, M. J. 2021 · 2021
Cited alongside, same era.
Rpvnet: A deep and efficient range-point-voxel fusion network for lidar point cloud segmentation
Xu, J.; Zhang, R.; Dou, J.; Zhu, Y.; Sun, J.; and Pu, S. 2021 · 2021
Cited alongside, same era.
Point transformer
Zhao, H.; Jiang, L.; Jia, J.; Torr, P. H.; and Koltun, V. 2021 · 2021
Cited alongside, same era.
How do vision transformers work?
Park, N.; and Kim, S. 2022 · 2022
Closest in time.
Pix4Point: Image Pretrained Transformers for 3D Point Cloud Understanding
Qian, G.; Zhang, X.; Hamdi, A.; and Ghanem, B. 2022 · 2022
Closest in time.
Softgroup for 3d instance segmentation on point clouds
Vu, T.; Kim, K.; Luu, T. M.; Nguyen, T.; and Yoo, C. D. 2022 · 2022
Closest in time.
Image2point: 3d point-cloud understanding with 2d image pretrained models
Xu, C.; Yang, S.; Galanti, T.; Wu, B.; Yue, X.; Zhai, B.; Zhan, W.; Vajda, P.; Keutzer, K.; and Tomizuka, M. 2022 · 2022
Closest in time.
2dpass: 2d priors assisted semantic segmentation on lidar point clouds
Yan, X.; Gao, J.; Zheng, C.; Zheng, C.; Zhang, R.; Cui, S.; and Li, Z. 2022 · 2022
Closest in time.
Point-bert: Pre-training 3d point cloud transformers with masked point modeling
Yu, X.; Tang, L.; Rao, Y.; Huang, T.; Zhou, J.; and Lu, J. 2022 · 2022
Closest in time.
Pointclip: Point cloud understanding by clip
Zhang, R.; Guo, Z.; Zhang, W.; Li, K.; Miao, X.; Cui, B.; Qiao, Y.; Gao, P.; and Li, H. 2022 · 2022
Closest in time.
Cylinder3d: An effective 3d framework for driving-scene lidar semantic segmentation
Zhou, H.; Zhu, X.; Song, X.; Ma, Y.; Wang, Z.; Li, H.; and Lin, D. 2020 · 2022
Closest in time.
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; Dollár, P.; and Girshick, R. 2023 · 2023
Closest in time.
Rethinking range view representation for lidar segmentation
Kong, L.; Liu, Y.; Chen, R.; Ma, Y.; Zhu, X.; Li, Y.; Hou, Y.; Qiao, Y.; and Liu, Z. 2023 · 2023
Closest in time.
Take-A-Photo: 3D-to-2D Generative Pre-training of Point Cloud Models
Wang, Z.; Yu, X.; Rao, Y.; Zhou, J.; and Lu, J. 2023 · 2023
Closest in time.
LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark
Yin, Z.; Wang, J.; Cao, J.; Shi, Z.; Liu, D.; Li, M.; Sheng, L.; Bai, L.; Huang, X.; Wang, Z.; et al. 2023 · 2023
Closest in time.