Fetching the paper…
Reading the bibliography…
ViTs are often too computationally expensive to be fitted onto real-world resource-constrained devices, due to (1) their quadratically increased complexity with the number of input tokens and (2) their overparameterized self-attention heads and model depth.
Robust sparse regularization: Simultaneously optimizing neural network robustness and compactness
Rakin, A. S.; He, Z.; Yang, L.; Wang, Y.; Wang, L.; and Fan, D. 2019 · 1905
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016 · 1937
Earlier work this paper cites.
Hu, T.-K.; Chen, T.; Wang, H.; and Wang, Z. 2020 · 2002
Earlier work this paper cites.
Hydra: Pruning adversarially robust neural networks
Sehwag, V.; Wang, S.; Mittal, P.; and Jana, S. 2020 · 2002
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Fu, Y.; You, H.; Zhao, Y.; Wang, Y.; Li, C.; Gopalakrishnan, K.; Wang, Z.; and Lin, Y. 2020 · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Adversarial examples in the physical world
Kurakin, A.; Goodfellow, I.; Bengio, S.; et al. 2016 · 2016
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks
Teerapittayanon, S.; McDanel, B.; and Kung, H.-T. 2016 · 2016
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; and Adam, H. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A.; Makelov, A.; Schmidt, L.; Tsipras, D.; and Vladu, A. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Athalye, A.; Carlini, N.; and Wagner, D. 2018 · 2018
Cited alongside, same era.
Rakin, A. S.; Yi, J.; Gong, B.; and Fan, D. 2018 · 2018
Cited alongside, same era.
Skipnet: Learning dynamic routing in convolutional networks
Wang, X.; Yu, F.; Dou, Z.-Y.; Darrell, T.; and Gonzalez, J. E. 2018 · 2018
Cited alongside, same era.
Blockdrop: Dynamic inference paths in residual networks
Wu, Z.; Nagarajan, T.; Kumar, A.; Rennie, S.; Davis, L. S.; Grauman, K.; and Feris, R. 2018 · 2018
Cited alongside, same era.
LeViT: a Vision Transformer in ConvNet’s Clothing for Faster Inference
Graham, B.; El-Nouby, A.; Touvron, H.; Stock, P.; Joulin, A.; Jégou, H.; and Douze, M. 2021 · 2021
Closest in time.
Rethinking spatial dimensions of vision transformers
Heo, B.; Yun, S.; Han, D.; Chun, S.; Choe, J.; and Oh, S. J. 2021 · 2021
Closest in time.
Localvit: Bringing locality to vision transformers
Li, Y.; Zhang, K.; Cao, J.; Timofte, R.; and Van Gool, L. 2021 · 2021
Closest in time.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Closest in time.
Rethinking the Design Principles of Robust Vision Transformer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shallow-deep networks: Understanding and mitigating network overthinking
Kaya, Y.; Hong, S.; and Dumitras, T. 2019 · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M.; and Le, Q. 2019 · 2019
Cited alongside, same era.
Adversarial robustness vs. model compression, or both?
Ye, S.; Xu, K.; Liu, S.; Cheng, H.; Lambrechts, J.-H.; Zhang, H.; Zhou, A.; Ma, K.; Wang, Y.; and Lin, X. 2019 · 2019
Cited alongside, same era.
Robust overfitting may be mitigated by properly learned smoothening
Chen, T.; Zhang, Z.; Liu, S.; Chang, S.; and Wang, Z. 2020 · 2020
Cited alongside, same era.
Fractional skipping: Towards finer-grained dynamic cnn inference
Shen, J.; Wang, Y.; Xu, P.; Fu, Y.; Wang, Z.; and Lin, Y. 2020 · 2020
Cited alongside, same era.
Dual dynamic inference: Enabling more efficient, adaptive, and controllable deep inference
Wang, Y.; Shen, J.; Hu, T.-K.; Xu, P.; Nguyen, T.; Baraniuk, R.; Wang, Z.; and Lin, Y. 2020 · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; and Joulin, A. 2021 · 2021
Cited alongside, same era.
Mao, X.; Qi, G.; Chen, Y.; Li, X.; Ye, S.; He, Y.; and Xue, H. 2021 · 2021
Closest in time.
DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; and Hsieh, C.-J. 2021 · 2021
Closest in time.
Imagenet-21k pretraining for the masses
Ridnik, T.; Ben-Baruch, E.; Noy, A.; and Zelnik-Manor, L. 2021 · 2021
Closest in time.
On the adversarial robustness of visual transformers
Shao, R.; Shi, Z.; Yi, J.; Chen, P.-Y.; and Hsieh, C.-J. 2021 · 2021
Closest in time.
How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
Steiner, A.; Kolesnikov, A.; Zhai, X.; Wightman, R.; Uszkoreit, J.; and Beyer, L. 2021 · 2021
Closest in time.
Segmenter: Transformer for Semantic Segmentation
Strudel, R.; Garcia, R.; Laptev, I.; and Schmid, C. 2021 · 2021
Closest in time.
Training data-efficient image transformers & distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021 · 2021
Closest in time.
Co-scale conv-attentional image transformers
Xu, W.; Xu, Y.; Chang, T.; and Tu, Z. 2021 · 2021
Closest in time.
Deepvit: Towards deeper vision transformer
Zhou, D.; Kang, B.; Jin, X.; Yang, L.; Lian, X.; Jiang, Z.; Hou, Q.; and Feng, J. 2021 · 2021
Closest in time.