Fetching the paper…
Reading the bibliography…
Vision Transformers (ViTs) have achieved impressive performance over various computer vision tasks.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in ICML , 2010
2010
Earlier work this paper cites.
2013
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in NeurIPSW , 2014
2014
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” IJCV , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML , 2015, pp. 448–456
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS , pp. 5998–6008, 2017
2017
Earlier work this paper cites.
P. Molchanov, S. Tyree, T. Karras, T. Aila, and J. Kautz, “Pruning convolutional neural networks for resource efficient inference,” in ICLR , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS , pp. 5998–6008, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
P. Michel, O. Levy, and G. Neubig, “Are sixteen heads really better than one?” NeurIPS , vol. 32, 2019
2019
Earlier work this paper cites.
R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel, “Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness,” in ICLR , 2019
2019
Earlier work this paper cites.
W. Brendel and M. Bethge, “Approximating cnns with bag-of-local-features models works surprisingly well on imagenet,” in ICLR , 2019
2019
Earlier work this paper cites.
G. Lample, A. Sablayrolles, M. Ranzato, L. Denoyer, and H. Jégou, “Large memory layers with product keys,” NeurIPS , 2019
2019
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV , 2020, pp. 213–229
2020
Earlier work this paper cites.
J.-B. Cordonnier, A. Loukas, and M. Jaggi, “On the relationship between self-attention and convolutional layers,” in ICLR , 2020
2020
Earlier work this paper cites.
M. Behnke and K. Heafield, “Losing heads in the lottery: Pruning transformer attention in neural machine translation,” in EMNLP , 2020, pp. 2664–2674
2020
Earlier work this paper cites.
B. Li, Z. Kong, T. Zhang, J. Li, Z. Li, H. Liu, and C. Ding, “Efficient transformer-based large scale language representations using hardware-friendly block structured pruning,” in EMNLP , 2020
2020
Earlier work this paper cites.
B. R. Bartoldson, A. S. Morcos, A. Barbu, and G. Erlebacher, “The generalization-stability tradeoff in neural network pruning,” NeurIPS , 2020
2020
Earlier work this paper cites.
T. Chen, J. Frankle, S. Chang, S. Liu, Y. Zhang, Z. Wang, and M. Carbin, “The lottery ticket hypothesis for pre-trained bert networks,” NeurIPS , 2020
2020
Earlier work this paper cites.
S. Prasanna, A. Rogers, and A. Rumshisky, “When bert plays the lottery, all tickets are winning,” in EMNLP , 2020
2020
Earlier work this paper cites.
U. Evci, T. Gale, J. Menick, P. S. Castro, and E. Elsen, “Rigging the lottery: Making all tickets winners,” in ICML . PMLR, 2020, pp. 2943–2952
2020
Earlier work this paper cites.
D. Stamoulis, R. Ding, D. Wang, D. Lymberopoulos, B. Priyantha, J. Liu, and D. Marculescu, “Single-path nas: Designing hardware-efficient convnets in less than 4 hours,” in ECML PKDD , 2020
2020
Earlier work this paper cites.
Z. Guo, X. Zhang, H. Mu, W. Heng, Z. Liu, Y. Wei, and J. Sun, “Single path one-shot neural architecture search with uniform sampling,” in ECCV , 2020, pp. 544–560
2020
Earlier work this paper cites.
H. Cai, C. Gan, T. Wang, Z. Zhang, and S. Han, “Once-for-all: Train one network and specialize it for efficient deployment,” in ICLR , 2020
2020
Earlier work this paper cites.
Y. Li, G. Hu, Y. Wang, T. Hospedales, N. M. Robertson, and Y. Yang, “Differentiable automatic data augmentation,” in ECCV , 2020, pp. 580–595
2020
Earlier work this paper cites.
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollár, “Designing network design spaces,” in CVPR , 2020, pp. 10 428–10 436
2020
Earlier work this paper cites.
S. Goyal, A. R. Choudhury, S. Raje, V. Chakaravarthy, Y. Sabharwal, and A. Verma, “Power-bert: Accelerating bert inference via progressive word-vector elimination,” in ICML , 2020, pp. 3690–3699
2020
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Cited alongside, same era.
Z. Pan, B. Zhuang, J. Liu, H. He, and J. Cai, “Scalable visual transformers with hierarchical pooling,” in ICCV , 2021
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV , 2021
2021
Cited alongside, same era.
M. Raghu, T. Unterthiner, S. Kornblith, C. Zhang, and A. Dosovitskiy, “Do vision transformers see like convolutional neural networks?” NeurIPS , 2021
2021
Closest in time.
2021
Closest in time.
M. Chen, H. Peng, J. Fu, and H. Ling, “Autoformer: Searching transformers for visual recognition,” in ICCV , 2021, pp. 12 270–12 280
2021
Closest in time.
2021
Closest in time.
J. Li, R. Cotterell, and M. Sachan, “Differentiable subset pruning of transformer heads,” TACL , 2021
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in ICML , 2021, pp. 10 347–10 357
2021
Cited alongside, same era.
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. Torr et al. , “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” in CVPR , 2021, pp. 6881–6890
2021
Cited alongside, same era.
B. Cheng, A. Schwing, and A. Kirillov, “Per-pixel classification is not all you need for semantic segmentation,” NeurIPS , vol. 34, 2021
2021
Cited alongside, same era.
H. Wang, Y. Zhu, H. Adam, A. Yuille, and L.-C. Chen, “Max-deeplab: End-to-end panoptic segmentation with mask transformers,” in CVPR , 2021, pp. 5463–5474
2021
Cited alongside, same era.
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable {detr}: Deformable transformers for end-to-end object detection,” in ICLR , 2021
2021
Cited alongside, same era.
P. Gao, M. Zheng, X. Wang, J. Dai, and H. Li, “Fast convergence of detr with spatially modulated co-attention,” in ICCV , 2021, pp. 3621–3630
2021
Cited alongside, same era.
M. Zhu, K. Han, and Y. Tang, “Visual transformer pruning,” in KDDW , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Closest in time.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in ICCV , 2021
2021
Closest in time.
C. Li, G. Wang, B. Wang, X. Liang, Z. Li, and X. Chang, “Dynamic slimmable network,” in CVPR , 2021, pp. 8607–8617
2021
Closest in time.
Y. Dong, J.-B. Cordonnier, and A. Loukas, “Attention is not all you need: Pure attention loses rank doubly exponentially with depth,” in ICML , 2021
2021
Closest in time.
H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jégou, “Going deeper with image transformers,” in ICCV , 2021
2021
Closest in time.
2021
Closest in time.
X. Zhang, X. Zhou, M. Lin, and J. Sun, “Shufflenet: An extremely efficient convolutional neural network for mobile devices,” in CVPR , 2018, pp. 6848–6856
2021
Closest in time.
Q. Zhang, Y. Xu, J. Zhang, and D. Tao, “Vsa: Learning varied-size window attention in vision transformers,” in ECCV , 2022
2022
Closest in time.
J. Guo, K. Han, H. Wu, C. Xu, Y. Tang, C. Xu, and Y. Wang, “Cmt: Convolutional neural networks meet vision transformers,” in CVPR , 2022
2022
Closest in time.
S. Yu, T. Chen, J. Shen, H. Yuan, J. Tan, S. Yang, J. Liu, and Z. Wang, “Unified visual transformer compression,” in ICLR , 2022
2022
Closest in time.
Z. Hou and S.-Y. Kung, “Multi-dimensional model compression of vision transformer,” in ICME , 2022
2022
Closest in time.
S. Mehta and M. Rastegari, “Mobilevit: light-weight, general-purpose, and mobile-friendly vision transformer,” in ICLR , 2022
2022
Closest in time.
P. Gao, T. Ma, H. Li, Z. Lin, J. Dai, and Y. Qiao, “Convmae: Masked convolution meets masked autoencoders,” NeurIPS , 2022
2022
Closest in time.
Z. Lu, H. Xie, C. Liu, and Y. Zhang, “Bridging the gap between vision transformers and convolutional neural networks on small datasets,” NeurIPS , vol. 35, pp. 14 663–14 677, 2022
2022
Closest in time.
Z. Pan, B. Zhuang, H. He, J. Liu, and J. Cai, “Less is more: Pay less attention in vision transformers,” in AAAI , 2022
2022
Closest in time.
Y. Liang, C. Ge, Z. Tong, Y. Song, J. Wang, and P. Xie, “Not all patches are what you need: Expediting vision transformers via token reorganizations,” in ICLR , 2022
2022
Closest in time.
H. Yang, H. Yin, P. Molchanov, H. Li, and J. Kautz, “Nvit: Vision transformer compression and parameter redistribution,” in CVPR , 2023
2023
Closest in time.
D. Bolya, C.-Y. Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman, “Token merging: Your vit but faster,” in ICLR , 2023
2023
Closest in time.
L. Yu and W. Xiang, “X-pruner: explainable pruning for vision transformers,” in CVPR , 2023, pp. 24 355–24 363
2023
Closest in time.
N. Zhang, F. Nex, G. Vosselman, and N. Kerle, “Lite-mono: A lightweight cnn and transformer architecture for self-supervised monocular depth estimation,” in CVPR , 2023, pp. 18 537–18 546
2023
Closest in time.
Y. Zhang, X. Guo, M. Poggi, Z. Zhu, G. Huang, and S. Mattoccia, “Completionformer: Depth completion with convolutions and vision transformers,” in CVPR , 2023, pp. 18 527–18 536
2023
Closest in time.
P. K. A. Vasu, J. Gabriel, J. Zhu, O. Tuzel, and A. Ranjan, “Fastvit: A fast hybrid vision transformer using structural reparameterization,” in ICCV , 2023
2023
Closest in time.
J. Liu, B. Zhuang, P. Chen, Y. Guo, C. Shen, J. Cai, and M. Tan, “Single-path bit sharing for automatic loss-aware model compression,” TPAMI , 2023
2023
Closest in time.
S. Wei, T. Ye, S. Zhang, Y. Tang, and J. Liang, “Joint token pruning and squeezing towards more aggressive compression of vision transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 2092–2101
2092
Closest in time.