Fetching the paper…
Reading the bibliography…
The use of pretrained backbones with fine-tuning has been successful for 2D vision and natural language processing tasks, showing advantages over task-specific networks.
ShapeNet: An information-rich 3D model repository
Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu · 2015
Earlier work this paper cites.
3D semantic parsing of large-scale indoor spaces
Iro Armeni, Ozan Sener, Amir Roshan Zamir, Helen Jiang, Ioannis K. Brilakis, Martin Fischer, and Silvio Savarese · 2016
Earlier work this paper cites.
ScanNet: Richly-annotated 3D reconstructions of indoor scenes
Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
O-CNN: Octree-based convolutional neural networks for 3D shape analysis
Peng-Shuai Wang, Yang Liu, Yu-Xiao Guo, Chun-Yu Sun, and Xin Tong · 2017
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
4D spatio-temporal ConvNets: Minkowski convolutional neural networks
Christopher Choy, JunYoung Gwak, and Silvio Savarese · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
KPConv: Flexible and deformable convolution for point clouds
Hugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui, François Goulette, and Leonidas J. Guibas · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Generative sparse detection networks for 3D single-shot object detection
JunYoung Gwak, Christopher Choy, and Silvio Savarese · 2020
Earlier work this paper cites.
Generative sparse detection networks for 3D single-shot object detection
JunYoung Gwak, Christopher B Choy, and Silvio Savarese · 2020
Earlier work this paper cites.
Unsupervised 3D learning for shape analysis via multiresolution instance discrimination
Peng-Shuai Wang, Yu-Qi Yang, Qian-Fang Zou, Zhirong Wu, Yang Liu, and Xin Tong · 2020
Earlier work this paper cites.
PointContrast: Unsupervised pre-training for 3D point cloud understanding
Saining Xie, Jiatao Gu, Demi Guo, Charles R Qi, Leonidas J Guibas, and Or Litany · 2020
Earlier work this paper cites.
Structured3D: A large photo-realistic dataset for structured 3D modeling
Jia Zheng, Junfei Zhang, Jing Li, Rui Tang, Shenghua Gao, and Zihan Zhou · 2020
Earlier work this paper cites.
ARKitScenes: A diverse real-world dataset For 3D indoor scene understanding using mobile RGB-D data
Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, and Elad Shulman · 2021
Earlier work this paper cites.
Twins: Revisiting the design of spatial attention in vision transformers
Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Earlier work this paper cites.
PCT: Point cloud transformer
Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R Martin, and Shi-Min Hu · 2021
Earlier work this paper cites.
Exploring data-efficient 3D scene understanding with contrastive scene contexts
Ji Hou, Benjamin Graham, Matthias Nießner, and Saining Xie · 2021
Earlier work this paper cites.
Transformers in vision: A survey
Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Earlier work this paper cites.
Mix3D: Out-of-context data augmentation for 3D scenes
Alexey Nekrasov, Jonas Schult, Or Litany, Bastian Leibe, and Francis Engelmann · 2021
Earlier work this paper cites.
Yongming Rao, Benlin Liu, Yi Wei, Jiwen Lu, Cho-Jui Hsieh, and Jie Zhou · 2021
Cited alongside, same era.
FCAF3D: Fully convolutional anchor-free 3D object detection
Danila Rukhovich, Anna Vorontsova, and Anton Konushin · 2021
Cited alongside, same era.
Rethinking and improving relative position encoding for vision transformer
Kan Wu, Houwen Peng, Minghao Chen, Jianlong Fu, and Hongyang Chao · 2021
Cited alongside, same era.
Focal self-attention for local-global interactions in vision transformers
Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai, Bin Xiao, Lu Yuan, and Jianfeng Gao · 2021
Cited alongside, same era.
Swin transformer V2: Scaling up capacity and resolution
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al · 2022
Later among the works it cites.
Masked autoencoders for point cloud self-supervised learning
Yatian Pang, Wenxiao Wang, Francis E. H. Tay, W. Liu, Yonghong Tian, and Liuliang Yuan · 2022
Later among the works it cites.
Fast point transformer
Chunghyun Park, Yoonwoo Jeong, Minsu Cho, and Jaesik Park · 2022
Later among the works it cites.
PointNeXt: Revisiting PointNet++ with improved training and scaling strategies
Guocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai, Hasan Abed Al Kader Hammoud, Mohamed Elhoseiny, and Bernard Ghanem · 2022
Later among the works it cites.
Surface representation for point clouds
Haoxi Ran, Jun Liu, and Chengjie Wang · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuhui Yuan, Rao Fu, Lang Huang, Weihong Lin, Chao Zhang, Xilin Chen, and Jingdong Wang · 2021
Cited alongside, same era.
Self-supervised pretraining of 3D features on any point-cloud
Zaiwei Zhang, Rohit Girdhar, Armand Joulin, and Ishan Misra · 2021
Cited alongside, same era.
Point transformer
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun · 2021
Cited alongside, same era.
Chuhang Zou, Jheng-Wei Su, Chi-Han Peng, Alex Colburn, Qi Shan, Peter Wonka, Hung-Kuo Chu, and Derek Hoiem · 2021
Cited alongside, same era.
BEiT: BERT pre-training of image transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei · 2022
Cited alongside, same era.
MixFormer: Mixing features across windows and dimensions
Qiang Chen, Qiman Wu, Jian Wang, Qinghao Hu, Tao Hu, Errui Ding, Jian Cheng, and Jingdong Wang · 2022
Cited alongside, same era.
PointMixer: MLP-Mixer for point cloud understanding
Jaesung Choe, Chunghyun Park, Francois Rameau, Jaesik Park, and In So Kweon · 2022
Cited alongside, same era.
CSWin transformer: A general vision transformer backbone with cross-shaped windows
Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo · 2022
Cited alongside, same era.
Later among the works it cites.
MaxViT: Multi-axis vision transformer
Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li · 2022
Later among the works it cites.
SoftGroup for 3D instance segmentation on point clouds
Thang Vu, Kookhoi Kim, Tung M Luu, Thanh Nguyen, and Chang D Yoo · 2022
Later among the works it cites.
CAGroup3D: Class-aware grouping for 3D object detection on point clouds
Haiyang Wang, Lihe Ding, Shaocong Dong, Shaoshuai Shi, Aoxue Li, Jianan Li, Zhenguo Li, and Liwei Wang · 2022
Later among the works it cites.
Window normalization: Enhancing point cloud understanding by unifying inconsistent point densities
Qi Wang, Sheng Shi, Jiahui Li, Wuming Jiang, and Xiangde Zhang · 2022
Later among the works it cites.
Crossformer: A versatile vision transformer hinging on cross-scale attention
Wenxiao Wang, Lu Yao, Long Chen, Binbin Lin, Deng Cai, Xiaofei He, and Wei Liu · 2022
Later among the works it cites.
P2P: Tuning pre-trained image models for point cloud analysis with point-to-pixel prompting
Ziyi Wang, Xumin Yu, Yongming Rao, Jie Zhou, and Jiwen Lu · 2022
Later among the works it cites.
Pale transformer: A general vision transformer backbone with pale-shaped attention
Sitong Wu, Tianyi Wu, Haoru Tan, and Guodong Guo · 2022
Later among the works it cites.
Point Transformer V2: Grouped vector attention and partition-based pooling
Xiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu, and Hengshuang Zhao · 2022
Later among the works it cites.
Point-BERT: Pre-training 3D point cloud transformers with masked point modeling
Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang, Jie Zhou, and Jiwen Lu · 2022
Later among the works it cites.
VSA: Learning varied-size window attention in vision transformers
Qiming Zhang, Yufei Xu, Jing Zhang, and Dacheng Tao · 2022
Later among the works it cites.
Point-M2AE: Multi-scale masked autoencoders for hierarchical point cloud pre-training
Renrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang, Bin Zhao, Dong Wang, Yu Qiao, and Hongsheng Li · 2022
Later among the works it cites.
LargeKernel3D: Scaling up kernels in 3D CNNs , 2023
Yukang Chen, Jianhui Liu, Xiaojuan Qi, Xiangyu Zhang, Jian Sun, and Jiaya Jia · 2023
Closest in time.
Runpei Dong, Zekun Qi, Linfeng Zhang, Junbo Zhang, Jianjian Sun, Zheng Ge, Li Yi, and Kaisheng Ma · 2023
Closest in time.
Meta architecture for point cloud analysis
Haojia Lin, Xiawu Zheng, Lijiang Li, Fei Chao, Shanshan Wang, Yan Wang, Yonghong Tian, and Rongrong Ji · 2023
Closest in time.
OctFormer: Octree-based transformers for 3D point clouds
Peng-Shuai Wang · 2023
Closest in time.
PointConvFormer: Revenge of the point-based convolution
Wenxuan Wu, Qi Shan, and Li Fuxin · 2023
Closest in time.
Masked scene contrast: A scalable framework for unsupervised 3D representation learning
Xiaoyang Wu, Xin Wen, Xihui Liu, and Hengshuang Zhao · 2023
Closest in time.