Fetching the paper…
Reading the bibliography…
Inspired by the great success achieved by CNN in image recognition, view-based methods applied CNNs to model the projected views for 3D object understanding and achieved excellent performance.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
VoxNet: A 3d convolutional neural network for real-time object recognition
Daniel Maturana and Sebastian A. Scherer · 2015
Earlier work this paper cites.
Multi-view convolutional neural networks for 3d shape recognition
Hang Su, Subhransu Maji, Evangelos Kalogerakis, and Erik G. Learned-Miller · 2015
Earlier work this paper cites.
3D ShapeNets: A deep representation for volumetric shapes
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
GIFT: A real-time and scalable 3d shape search engine
Song Bai, Xiang Bai, Zhichao Zhou, Zhaoxiang Zhang, and Longin Jan Latecki · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Pairwise decomposition of image sequences for active multi-view recognition
Edward Johns, Stefan Leutenegger, and Andrew J. Davison · 2016
Earlier work this paper cites.
Volumetric and multi-view CNNs for object classification on 3d data
Charles Ruizhongtai Qi, Hao Su, Matthias Nießner, Angela Dai, Mengyuan Yan, and Leonidas J. Guibas · 2016
Earlier work this paper cites.
PointNet: Deep learning on point sets for 3d classification and segmentation
Charles Ruizhongtai Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas · 2017
Earlier work this paper cites.
3d-a-nets: 3d deep dense descriptor for volumetric shapes with adversarial networks
Mengwei Ren, Liang Niu, and Yi Fang · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Dominant set clustering and pooling for multi-view 3d object recognition
Chu Wang, Marcello Pelillo, and Kaleem Siddiqi · 2017
Cited alongside, same era.
3DmFV: Three-dimensional point cloud classification in real-time using convolutional neural networks
Yizhak Ben-Shabat, Michael Lindenbaum, and Anath Fischer · 2018
Cited alongside, same era.
GVCNN: group-view convolutional neural networks for 3d shape recognition
Yifan Feng, Zizhao Zhang, Xibin Zhao, Rongrong Ji, and Yue Gao · 2018
Cited alongside, same era.
RotationNet: Joint object categorization and pose estimation using multiviews from unsupervised viewpoints
Asako Kanezaki, Yasuyuki Matsushita, and Yoshifumi Nishida · 2018
Cited alongside, same era.
Multi-view harmonized bilinear network for 3d object recognition
Tan Yu, Jingjing Meng, and Junsong Yuan · 2018
Cited alongside, same era.
Twins: Revisiting the design of spatial attention in vision transformers
Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Container: Context aggregation network
Peng Gao, Jiasen Lu, Hongsheng Li, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2021
Closest in time.
Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang · 2021
Closest in time.
Rethinking spatial dimensions of vision transformers
Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, and Seong Joon Oh · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
DeepCCFV: Camera constraint-free multi-view convolutional neural network for 3d object retrieval
Zhengyue Huang, Zhehui Zhao, Hengguang Zhou, Xibin Zhao, and Yue Gao · 2019
Cited alongside, same era.
LP-3DCNN: unveiling local phase in 3d convolutional neural networks
Sudhakar Kumawat and Shanmuganathan Raman · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
VV-Net: Voxel VAE net with group convolutions for point cloud segmentation
Hsien-Yu Meng, Lin Gao, Yu-Kun Lai, and Dinesh Manocha · 2019
Cited alongside, same era.
Learning relationships for multi-view 3d object recognition
Ze Yang and Liwei Wang · 2019
Cited alongside, same era.
View-GCN: View-based graph convolutional network for 3d shape analysis
Xin Wei, Ruixuan Yu, and Jian Sun · 2020
Cited alongside, same era.
Closest in time.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Closest in time.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Closest in time.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2021
Closest in time.
CvT: Introducing convolutions to vision transformers
Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang · 2021
Closest in time.
Multi-view 3d shape recognition via correspondence-aware deep learning
Yong Xu, Chaoda Zheng, Ruotao Xu, Yuhui Quan, and Haibin Ling · 2021
Closest in time.
3d object representation learning: A set-to-set matching perspective
Tan Yu, Jingjing Meng, Ming Yang, and Junsong Yuan · 2021
Closest in time.
Tokens-to-token ViT: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Closest in time.