Fetching the paper…
Reading the bibliography…
Vision Transformers (ViTs) have triggered the most recent and significant breakthroughs in computer vision.
The fast fourier transform and its applications
J. W. Cooley, P. A. W. Lewis, and P. D. Welch · 1969
Earlier work this paper cites.
Low-pass filters for signal averaging
E. Voigtman and J. D. Winefordner · 1986
Earlier work this paper cites.
An adaptive gaussian filter for noise reduction and edge detection
G. Deng and L. Cahill · 1993
Earlier work this paper cites.
Microsoft COCO: common objects in context
T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
J. Ba, J. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Compressing convolutional neural networks in the frequency domain
W. Chen, J. T. Wilson, S. Tyree, K. Q. Weinberger, and Y. Chen · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Mask R-CNN
K. He, G. Gkioxari, P. Dollár, and R. B. Girshick · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T. Lin, P. Goyal, R. B. Girshick, K. He, and P. Dollár · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. B. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Earlier work this paper cites.
Shufflenet v2: Practical guidelines for efficient cnn architecture design
N. Ma, X. Zhang, H.-T. Zheng, and J. Sun · 2018
Earlier work this paper cites.
MMDetection: Open mmlab detection toolbox and benchmark
K. Chen, J. Wang, J. Pang, Y. Cao, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Xu, Z. Zhang, D. Cheng, C. Zhu, T. Cheng, Q. Zhao, B. Li, X. Lu, R. Zhu, Y. Wu, J. Dai, J. Wang, J. Shi, W. Ouyang, C. C. Loy, and D. Lin · 2019
Earlier work this paper cites.
Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution
Y. Chen, H. Fan, B. Xu, Z. Yan, Y. Kalantidis, M. Rohrbach, S. Yan, and J. Feng · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Earlier work this paper cites.
Frequency separation for real-world super-resolution
M. Fritsche, S. Gu, and R. Timofte · 2019
Earlier work this paper cites.
Panoptic feature pyramid networks
A. Kirillov, R. B. Girshick, K. He, and P. Dollár · 2019
Earlier work this paper cites.
Semantic understanding of scenes through the ADE20K dataset
B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba · 2019
Earlier work this paper cites.
On the relationship between self-attention and convolutional layers
J. Cordonnier, A. Loukas, and M. Jaggi · 2020
Earlier work this paper cites.
How much position information do convolutional neural networks encode?
M. A. Islam, S. Jia, and N. D. B. Bruce · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Cited alongside, same era.
Compressive transformers for long-range sequence modelling
J. W. Rae, A. Potapenko, S. M. Jayakumar, and T. P. Lillicrap · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
S. Wang, B. Li, M. Khabsa, H. Fang, and H. Ma · 2020
Cited alongside, same era.
Invertible image rescaling
M. Xiao, S. Zheng, C. Liu, Y. Wang, D. He, G. Ke, J. Bian, Z. Lin, and T. Liu · 2020
Cited alongside, same era.
Learning in the frequency domain
K. Xu, M. Qin, F. Sun, Y. Wang, Y. Chen, and F. Ren · 2020
Cited alongside, same era.
Guided frequency separation network for real-world super-resolution
Y. Zhou, W. Deng, T. Tong, and Q. Gao · 2020
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2021
Later among the works it cites.
Resnet strikes back: An improved training procedure in timm
R. Wightman, H. Touvron, and H. Jégou · 2021
Later among the works it cites.
Cvt: Introducing convolutions to vision transformers
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang · 2021
Later among the works it cites.
Early convolutions help transformers see better
T. Xiao, M. Singh, E. Mintun, T. Darrell, P. Dollár, and R. B. Girshick · 2021
Later among the works it cites.
Focal self-attention for local-global interactions in vision transformers
J. Yang, C. Li, P. Zhang, X. Dai, B. Xiao, L. Yuan, and J. Gao · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Xcit: Cross-covariance image transformers
A. Ali, H. Touvron, M. Caron, P. Bojanowski, M. Douze, A. Joulin, I. Laptev, N. Neverova, G. Synnaeve, J. Verbeek, and H. Jégou · 2021
Cited alongside, same era.
Glit: Neural architecture search for global and local image transformer
B. Chen, P. Li, C. Li, B. Li, L. Bai, C. Lin, M. Sun, J. Yan, and W. Ouyang · 2021
Cited alongside, same era.
Autoformer: Searching transformers for visual recognition
M. Chen, H. Peng, J. Fu, and H. Ling · 2021
Cited alongside, same era.
Rethinking attention with performers
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlós, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller · 2021
Cited alongside, same era.
Twins: Revisiting the design of spatial attention in vision transformers
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen · 2021
Cited alongside, same era.
Coatnet: Marrying convolution and attention for all data sizes
Z. Dai, H. Liu, Q. V. Le, and M. Tan · 2021
Cited alongside, same era.
K. Yuan, S. Guo, Z. Liu, A. Zhou, F. Yu, and W. Wu · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, F. E. Tay, J. Feng, and S. Yan · 2021
Later among the works it cites.
Multi-scale vision longformer: A new vision transformer for high-resolution image encoding
P. Zhang, X. Dai, J. Yang, B. Xiao, L. Yuan, L. Zhang, and J. Gao · 2021
Later among the works it cites.
Mixformer: Mixing features across windows and dimensions
Q. Chen, Q. Wu, J. Wang, Q. Hu, T. Hu, E. Ding, J. Cheng, and J. Wang · 2022
Closest in time.
Mobile-former: Bridging mobilenet and transformer
Y. Chen, X. Dai, D. Chen, M. Liu, X. Dong, L. Yuan, and Z. Liu · 2022
Closest in time.
Cswin transformer: A general vision transformer backbone with cross-shaped windows
X. Dong, J. Bao, D. Chen, W. Zhang, N. Yu, L. Yuan, D. Chen, and B. Guo · 2022
Closest in time.
M.-H. Guo, C.-Z. Lu, Z.-N. Liu, M.-M. Cheng, and S.-M. Hu · 2022
Closest in time.
A convnet for the 2020s
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie · 2022
Closest in time.
Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer
S. Mehta and M. Rastegari · 2022
Closest in time.
Less is more: Pay less attention in vision transformers
Z. Pan, B. Zhuang, H. He, J. Liu, and J. Cai · 2022
Closest in time.
How do vision transformers work?
N. Park and S. Kim · 2022
Closest in time.
Hornet: Efficient high-order spatial interactions with recursive gated convolutions
Y. Rao, W. Zhao, Y. Tang, J. Zhou, S.-L. Lim, and J. Lu · 2022
Closest in time.
Quadtree attention for vision transformers
S. Tang, J. Zhang, S. Zhu, and P. Tan · 2022
Closest in time.
Maxim: Multi-axis mlp for image processing
Z. Tu, H. Talebi, H. Zhang, F. Yang, P. Milanfar, A. Bovik, and Y. Li · 2022
Closest in time.
Pvtv2: Improved baselines with pyramid vision transformer
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2022
Closest in time.
Vision transformer with deformable attention
Z. Xia, X. Pan, S. Song, L. E. Li, and G. Huang · 2022
Closest in time.
Vsa: Learning varied-size window attention in vision transformers
Q. Zhang, Y. Xu, J. Zhang, and D. Tao · 2022
Closest in time.