Fetching the paper…
Reading the bibliography…
The sparse transformer can reduce the computational complexity of the self-attention layers to $O(n)$, whilst still being a universal approximator of continuous sequence-to-sequence functions.
The monte carlo method
Metropolis, N. and Ulam, S · 1949
Earlier work this paper cites.
API design for machine learning software: experiences from the scikit-learn project
Buitinck, L., Louppe, G., Blondel, M., Pedregosa, F., Mueller, A., Grisel, O., Niculae, V., Prettenhofer, P., Gramfort, A., Grobler, J., Layton, R., VanderPlas, J., Joly, A., Holt, B., and Varoquaux, G · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
Chang, A. X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., Su, H., et al · 2015
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Deep sets
Zaheer, M., Kottur, S., Ravanbakhsh, S., Poczos, B., Salakhutdinov, R. R., and Smola, A. J · 2017
Earlier work this paper cites.
Pointcnn: Convolution on x-transformed points
Li, Y., Bu, R., Sun, M., Wu, W., Di, X., and Chen, B · 2018
Earlier work this paper cites.
Attentional shapecontextnet for point cloud recognition
Xie, S., Liu, S., Chen, Z., and Tu, Z · 2018
Earlier work this paper cites.
Spidercnn: Deep learning on point sets with parameterized convolutional filters
Xu, Y., Fan, T., Xu, M., Zeng, L., and Qiao, Y · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Guo, Q., Qiu, X., Liu, P., Shao, Y., Xue, X., and Zhang, Z · 2019
Earlier work this paper cites.
Set transformer: A framework for attention-based permutation-invariant neural networks
Lee, J., Lee, Y., Kim, J., Kosiorek, A., Choi, S., and Teh, Y. W · 2019
Earlier work this paper cites.
Stand-alone self-attention in vision models
Ramachandran, P., Parmar, N., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J · 2019
Cited alongside, same era.
Kpconv: Flexible and deformable convolution for point clouds
Thomas, H., Qi, C. R., Deschaud, J.-E., Marcotegui, B., Goulette, F., and Guibas, L. J · 2019
Cited alongside, same era.
Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data
Uy, M. A., Pham, Q.-H., Hua, B.-S., Nguyen, T., and Yeung, S.-K · 2019
Cited alongside, same era.
Are transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S. J., and Kumar, S · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Voxel transformer for 3d object detection
Mao, J., Xue, Y., Niu, M., Bai, H., Feng, J., Liang, X., Xu, H., and Xu, C · 2021
Later among the works it cites.
Cloud transformers: A universal approach to point cloud processing tasks
Mazur, K. and Lempitsky, V · 2021
Later among the works it cites.
An end-to-end transformer model for 3d object detection
Misra, I., Girdhar, R., and Joulin, A · 2021
Later among the works it cites.
Sparsebert: Rethinking the importance analysis in self-attention
Shi, H., Gao, J., Ren, X., Xu, H., Liang, X., Li, Z., and Kwok, J. T.-Y · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Later among the works it cites.
Unsupervised point cloud pre-training via occlusion completion
Wang, H., Liu, Q., Yue, X., Lasenby, J., and Kusner, M. J · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Se (3)-transformers: 3d roto-translation equivariant attention networks
Fuchs, F., Worrall, D., Fischer, V., and Welling, M · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al · 2020
Cited alongside, same era.
Self-supervised few-shot learning on point clouds
Sharma, C. and Kaul, M · 2020
Cited alongside, same era.
O (n) connections are expressive enough: Universal approximability of sparse transformers
Yun, C., Chang, Y.-W., Bhojanapalli, S., Rawat, A. S., Reddi, S., and Kumar, S · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Cited alongside, same era.
Beit: Bert pre-training of image transformers
Bao, H., Dong, L., and Wei, F · 2021
Cited alongside, same era.
Pct: Point cloud transformer
Guo, M.-H., Cai, J.-X., Liu, Z.-N., Mu, T.-J., Martin, R. R., and Hu, S.-M · 2021
Cited alongside, same era.
Later among the works it cites.
Pvt: Point-voxel transformer for point cloud learning
Zhang, C., Wan, H., Shen, X., and Wu, Z · 2021
Later among the works it cites.
Point transformer
Zhao, H., Jiang, L., Jia, J., Torr, P. H., and Koltun, V · 2021
Later among the works it cites.
Dual transformer for point cloud analysis
Han, X.-F., Jin, Y.-F., Cheng, H.-X., and Xiao, G.-Q · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Later among the works it cites.
Stratified transformer for 3d point cloud segmentation
Lai, X., Liu, J., Jiang, L., Wang, L., Zhao, H., Liu, S., Qi, X., and Jia, J · 2022
Later among the works it cites.
Masked autoencoders for point cloud self-supervised learning
Pang, Y., Wang, W., Tay, F. E., Liu, W., Tian, Y., and Yuan, L · 2022
Later among the works it cites.
Sinkformers: Transformers with doubly stochastic attention
Sander, M. E., Ablin, P., Blondel, M., and Peyré, G · 2022
Later among the works it cites.
Implicit autoencoder for point cloud self-supervised representation learning
Yan, S., Yang, Z., Li, H., Guan, L., Kang, H., Hua, G., and Huang, Q · 2022
Later among the works it cites.
Point-bert: Pre-training 3d point cloud transformers with masked point modeling
Yu, X., Tang, L., Rao, Y., Huang, T., Zhou, J., and Lu, J · 2022
Later among the works it cites.