Fetching the paper…
Reading the bibliography…
With Vision Transformers (ViTs) making great advances in a variety of computer vision tasks, recent literature have proposed various variants of vanilla ViTs to achieve better efficiency and efficacy.
Dynet: Dynamic convolution for accelerating convolutional neural networks
Zhang, Y.; Zhang, J.; Wang, Q.; and Zhong, Z. 2020 · 2004
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions
Chollet, F. 2017 · 2017
Earlier work this paper cites.
Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
Hendrycks, D.; and Dietterich, T. 2018 · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2018 · 2018
Earlier work this paper cites.
Classification-driven dynamic image enhancement
Sharma, V.; Diba, A.; Neven, D.; Brown, M. S.; Van Gool, L.; and Stiefelhagen, R. 2018 · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, H.; Cisse, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2018 · 2018
Earlier work this paper cites.
Autoaugment: Learning augmentation policies from data
Cubuk, E. D.; Zoph, B.; Mane, D.; Vasudevan, V.; and Le, Q. V. 2019 · 2019
Earlier work this paper cites.
Augmix: A simple data processing method to improve robustness and uncertainty
Hendrycks, D.; Mu, N.; Cubuk, E. D.; Zoph, B.; Gilmer, J.; and Lakshminarayanan, B. 2019 · 2019
Earlier work this paper cites.
How much Position Information Do Convolutional Neural Networks Encode?
Islam, M. A.; Jia, S.; and Bruce, N. D. 2019 · 2019
Earlier work this paper cites.
Condconv: Conditionally parameterized convolutions for efficient inference
Yang, B.; Bender, G.; Le, Q. V.; and Ngiam, J. 2019 · 2019
Earlier work this paper cites.
A fourier perspective on model robustness in computer vision
Yin, D.; Gontijo Lopes, R.; Shlens, J.; Cubuk, E. D.; and Gilmer, J. 2019 · 2019
Earlier work this paper cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S.; Han, D.; Oh, S. J.; Chun, S.; Choe, J.; and Yoo, Y. 2019 · 2019
Earlier work this paper cites.
Adversarial AutoAugment
Zhang, X.; Wang, Q.; Zhang, J.; and Zhong, Z. 2019 · 2019
Earlier work this paper cites.
Dynamic convolution: Attention over convolution kernels
Chen, Y.; Dai, X.; Liu, M.; Chen, D.; Yuan, L.; and Liu, Z. 2020 · 2020
Earlier work this paper cites.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D.; Zoph, B.; Shlens, J.; and Le, Q. V. 2020 · 2020
Earlier work this paper cites.
Augmix: A simple data processing method to improve robustness and uncertainty
Hendrycks, D.; Mu, N.; Cubuk, E. D.; Zoph, B.; Gilmer, J.; and Lakshminarayanan, B. 2020 · 2020
Earlier work this paper cites.
A simple way to make neural networks robust against diverse image corruptions
Rusak, E.; Schott, L.; Zimmermann, R. S.; Bitterwolf, J.; Bringmann, O.; Bethge, M.; and Brendel, W. 2020 · 2020
Cited alongside, same era.
Adversarial examples improve image recognition
Xie, C.; Tan, M.; Gong, B.; Wang, J.; Yuille, A. L.; and Le, Q. V. 2020 · 2020
Cited alongside, same era.
Xcit: Cross-covariance image transformers
Ali, A.; Touvron, H.; Caron, M.; Bojanowski, P.; Douze, M.; Joulin, A.; Laptev, I.; Neverova, N.; Synnaeve, G.; Verbeek, J.; et al. 2021 · 2021
Cited alongside, same era.
Are Transformers more robust than CNNs?
Bai, Y.; Mei, J.; Yuille, A. L.; and Xie, C. 2021 · 2021
Cited alongside, same era.
Understanding robustness of transformers for image classification
Bhojanapalli, S.; Chakrabarti, A.; Glasner, D.; Li, D.; Unterthiner, T.; and Veit, A. 2021 · 2021
Cited alongside, same era.
Crossvit: Cross-attention multi-scale vision transformer for image classification
Intriguing properties of vision transformers
Naseer, M. M.; Ranasinghe, K.; Khan, S. H.; Hayat, M.; Shahbaz Khan, F.; and Yang, M.-H. 2021 · 2021
Later among the works it cites.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; and Hsieh, C.-J. 2021 · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021 · 2021
Later among the works it cites.
Cvt: Introducing convolutions to vision transformers
Wu, H.; Xiao, B.; Codella, N.; Liu, M.; Dai, X.; Yuan, L.; and Zhang, L. 2021 · 2021
Later among the works it cites.
A fourier-based framework for domain generalization
Xu, Q.; Zhang, R.; Zhang, Y.; Wang, Y.; and Tian, Q. 2021 · 2021
Later among the works it cites.
Dynamic Resolution Network
Zhu, M.; Han, K.; Wu, E.; Zhang, Q.; Nie, Y.; Lan, Z.; and Wang, Y. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, C.-F. R.; Fan, Q.; and Panda, R. 2021 · 2021
Cited alongside, same era.
Amplitude-phase recombination: Rethinking robustness of convolutional neural networks in frequency domain
Chen, G.; Peng, P.; Ma, L.; Li, J.; Du, L.; and Tian, Y. 2021 · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
Robust Contrastive Learning Using Negative Samples with Diminished Semantics
Ge, S.; Mishra, S.; Li, C.-L.; Wang, H.; and Jacobs, D. 2021 · 2021
Cited alongside, same era.
Dynamic neural networks: A survey
Han, Y.; Huang, G.; Song, S.; Yang, L.; Wang, H.; and Wang, Y. 2021 · 2021
Cited alongside, same era.
Escaping the big data paradigm with compact transformers
Hassani, A.; Walton, S.; Shah, N.; Abuduweili, A.; Li, J.; and Shi, H. 2021 · 2021
Cited alongside, same era.
Rethinking spatial dimensions of vision transformers
Heo, B.; Yun, S.; Han, D.; Chun, S.; Choe, J.; and Oh, S. J. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
RegionViT: Regional-to-Local Attention for Vision Transformers
Chen, C.-F.; Panda, R.; and Fan, Q. 2022 · 2022
Closest in time.
Cswin transformer: A general vision transformer backbone with cross-shaped windows
Dong, X.; Bao, J.; Chen, D.; Zhang, W.; Yu, N.; Yuan, L.; Chen, D.; and Guo, B. 2022 · 2022
Closest in time.
Masked autoencoders are scalable vision learners
He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; and Girshick, R. 2022 · 2022
Closest in time.
Pyramid Adversarial Training Improves ViT Performance
Herrmann, C.; Sargent, K.; Jiang, L.; Zabih, R.; Chang, H.; Liu, C.; Krishnan, D.; and Sun, D. 2022 · 2022
Closest in time.
3D Common Corruptions and Data Augmentation
Kar, O. F.; Yeo, T.; Atanov, A.; and Zamir, A. 2022 · 2022
Closest in time.
Vision transformers are robust learners
Paul, S.; and Chen, P.-Y. 2022 · 2022
Closest in time.
TeachAugment: Data Augmentation Optimization Using Teacher Knowledge
Suzuki, T. 2022 · 2022
Closest in time.
Simmim: A simple framework for masked image modeling
Xie, Z.; Zhang, Z.; Cao, Y.; Lin, Y.; Bao, J.; Yao, Z.; Dai, Q.; and Hu, H. 2022 · 2022
Closest in time.
Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual Understanding
Zhang, Z.; Zhang, H.; Zhao, L.; Chen, T.; and Pfister, T. 2022 · 2022
Closest in time.