Fetching the paper…
Reading the bibliography…
Vision transformers (ViTs) are usually considered to be less light-weight than convolutional neural networks (CNNs) due to the lack of inductive bias.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger · 2016
Earlier work this paper cites.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2017
Earlier work this paper cites.
Squeeze-and-excitation networks
J. Hu, L. Shen, and G. Sun · 2018
Earlier work this paper cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen · 2018
Earlier work this paper cites.
Cbam: Convolutional block attention module
S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon · 2018
Earlier work this paper cites.
Mmdetection: Open mmlab detection toolbox and benchmark
K. Chen, J. Wang, J. Pang, Y. Cao, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Xu, et al · 2019
Earlier work this paper cites.
Searching for mobilenetv3
A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan, W. Wang, Y. Zhu, R. Pang, V. Vasudevan, et al · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
M. Tan and Q. Le · 2019
Earlier work this paper cites.
Randaugment: Practical automated data augmentation with a reduced search space
E. D. Cubuk, B. Zoph, J. Shlens, and Q. V. Le · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Designing network design spaces
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollár · 2020
Cited alongside, same era.
Visual transformers: Token-based image representation and processing for computer vision
B. Wu, C. Xu, X. Dai, A. Wan, P. Zhang, Z. Yan, M. Tomizuka, J. Gonzalez, K. Keutzer, and P. Vajda · 2020
Cited alongside, same era.
Greedynas: Towards fast one-shot nas with greedy supernet
S. You, T. Huang, M. Yang, F. Wang, C. Qian, and C. Zhang · 2020
Cited alongside, same era.
Xcit: Cross-covariance image transformers
A. Ali, H. Touvron, M. Caron, P. Bojanowski, M. Douze, A. Joulin, I. Laptev, N. Neverova, G. Synnaeve, J. Verbeek, et al · 2021
Cited alongside, same era.
Twins: Revisiting the design of spatial attention in vision transformers
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen · 2021
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2021
Later among the works it cites.
Early convolutions help transformers see better
T. Xiao, M. Singh, E. Mintun, T. Darrell, P. Dollár, and R. Girshick · 2021
Later among the works it cites.
Segformer: Simple and efficient design for semantic segmentation with transformers
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo · 2021
Later among the works it cites.
Lite vision transformer with enhanced self-attention
C. Yang, Y. Wang, J. Zhang, H. Zhang, Z. Wei, Z. Lin, and A. Yuille · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Fang, L. Xie, X. Wang, X. Zhang, W. Liu, and Q. Tian · 2021
Cited alongside, same era.
Simvit: Exploring a simple vision transformer with sliding windows, 2021
G. Li, D. Xu, X. Cheng, L. Si, and C. Zheng · 2021
Cited alongside, same era.
Localvit: Bringing locality to vision transformers
Y. Li, K. Zhang, J. Cao, R. Timofte, and L. Van Gool · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Cited alongside, same era.
Mobilevit: light-weight, general-purpose, and mobile-friendly vision transformer
S. Mehta and M. Rastegari · 2021
Cited alongside, same era.
Bottleneck transformers for visual recognition
A. Srinivas, T.-Y. Lin, N. Parmar, J. Shlens, P. Abbeel, and A. Vaswani · 2021
Cited alongside, same era.
Segmenter: Transformer for semantic segmentation
R. Strudel, R. Garcia, I. Laptev, and C. Schmid · 2021
Cited alongside, same era.
J. Yang, C. Li, P. Zhang, X. Dai, B. Xiao, L. Yuan, and J. Gao · 2021
Later among the works it cites.
Incorporating convolution designs into visual transformers
K. Yuan, S. Guo, Z. Liu, A. Zhou, F. Yu, and W. Wu · 2021
Later among the works it cites.
Rest: An efficient transformer for visual recognition
Q. Zhang and Y.-B. Yang · 2021
Later among the works it cites.
Weakly supervised contrastive learning
M. Zheng, F. Wang, S. You, C. Qian, C. Zhang, X. Wang, and C. Xu · 2021
Later among the works it cites.
Ressl: Relational self-supervised learning with weak augmentation
M. Zheng, S. You, F. Wang, C. Qian, C. Zhang, X. Wang, and C. Xu · 2021
Later among the works it cites.
Greedynasv2: greedier search with a greedy path filter
T. Huang, S. You, F. Wang, C. Qian, C. Zhang, X. Wang, and C. Xu · 2022
Closest in time.
Dyrep: Bootstrapping training with dynamic re-parameterization
T. Huang, S. You, B. Zhang, Y. Du, F. Wang, C. Qian, and C. Xu · 2022
Closest in time.
Exploring plain vision transformer backbones for object detection
Y. Li, H. Mao, R. Girshick, and K. He · 2022
Closest in time.
Pvt v2: Improved baselines with pyramid vision transformer
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2022
Closest in time.