Hypercorrelation squeeze for few-shot segmentation
J. Min, D. Kang, and M. Cho · 2021
Later among the works it cites.
Do vision transformers see like convolutional neural networks?
M. Raghu, T. Unterthiner, S. Kornblith, C. Zhang, and A. Dosovitskiy · 2021
Later among the works it cites.
Bottleneck transformers for visual recognition
A. Srinivas, T.-Y. Lin, N. Parmar, J. Shlens, P. Abbeel, and A. Vaswani · 2021
Later among the works it cites.
Segmenter: Transformer for semantic segmentation
R. Strudel, R. Garcia, I. Laptev, and C. Schmid · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Original
I. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit, M. Lucic, and A. Dosovitskiy · 2021
Later among the works it cites.
Resmlp: Feedforward networks for image classification with data-efficient training
Original
H. Touvron, P. Bojanowski, M. Caron, M. Cord, A. El-Nouby, E. Grave, G. Izacard, A. Joulin, G. Synnaeve, J. Verbeek, and H. Jégou · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jegou · 2021
Later among the works it cites.
Augmenting convolutional networks with attention-based aggregation
Original
H. Touvron, M. Cord, A. El-Nouby, P. Bojanowski, A. Joulin, G. Synnaeve, J. Verbeek, and H. Jegou · 2021
Later among the works it cites.
Scaling local self-attention for parameter efficient visual backbones
A. Vaswani, P. Ramachandran, A. Srinivas, N. Parmar, B. Hechtman, and J. Shlens · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2021
Later among the works it cites.
Resnet strikes back: An improved training procedure in timm
Original
R. Wightman, H. Touvron, and H. Jégou · 2021
Later among the works it cites.
Cvt: Introducing convolutions to vision transformers
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang · 2021
Later among the works it cites.
Rethinking and improving relative position encoding for vision transformer
K. Wu, H. Peng, M. Chen, J. Fu, and H. Chao · 2021
Later among the works it cites.
Early convolutions help transformers see better
T. Xiao, P. Dollar, M. Singh, E. Mintun, T. Darrell, and R. Girshick · 2021
Later among the works it cites.
Co-scale conv-attentional image transformers
W. Xu, Y. Xu, T. Chang, and Z. Tu · 2021
Later among the works it cites.
Focal self-attention for local-global interactions in vision transformers
J. Yang, C. Li, P. Zhang, X. Dai, B. Xiao, L. Yuan, and J. Gao · 2021
Later among the works it cites.
Metaformer is actually what you need for vision
Original
W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, and S. Yan · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z.-H. Jiang, F. E. Tay, J. Feng, and S. Yan · 2021
Later among the works it cites.
A battle of network structures: An empirical study of cnn, transformer, and mlp
Original
Y. Zhao, G. Wang, C. Tang, C. Luo, W. Zeng, and Z.-J. Zha · 2021
Later among the works it cites.
Finding biological plausibility for adversarially robust features via metameric tasks
A. Harrington and A. Deza · 2022
Closest in time.
Pruning self-attentions into convolutional layers in single path
Original
H. He, J. Liu, Z. Pan, J. Cai, J. Zhang, D. Tao, and B. Zhuang · 2022
Closest in time.
Foveater: Foveated transformer for image classification
Original
A. Jonnalagadda, W. Y. Wang, B. S. Manjunath, and M. P. Eckstein · 2022
Closest in time.
Mpvit: Multi-path vision transformer for dense prediction
Y. Lee, J. Kim, J. Willette, and S. J. Hwang · 2022
Closest in time.
Uniformer: Unifying convolution and self-attention for visual recognition
Original
K. Li, Y. Wang, J. Zhang, P. Gao, G. Song, Y. Liu, H. Li, and Y. Qiao · 2022
Closest in time.
Nommer: Nominate synergistic context in vision transformer for visual recognition
H. Liu, X. Jiang, X. Li, Z. Bao, D. Jiang, and B. Ren · 2022
Closest in time.
A convnet for the 2020s
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie · 2022
Closest in time.
MaxViT: Multi-Axis Vision Transformer
Z. Tu, H. Talebi, H. Zhang, F. Yang, P. Milanfar, A. Bovik, and Y. Li · 2022
Closest in time.
Peripheral vision — Wikipedia, the free encyclopedia
Wikipedia contributors · 2022
Closest in time.