Fetching the paper…
Reading the bibliography…
Transformers have shown great potential in various computer vision tasks owing to their strong capability in modeling long-range dependency using the self-attention mechanism.
Pyramid methods in image processing
E. H. Adelson, C. H. Anderson, J. R. Bergen, P. J. Burt, and J. M. Ogden · 1984
Earlier work this paper cites.
The laplacian pyramid as a compact image code
P. J. Burt and E. H. Adelson · 1987
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
Y. LeCun, Y. Bengio, et al · 1995
Earlier work this paper cites.
Gaussian pyramid wavelet transform for multiresolution analysis of images
H. Olkkonen and P. Pesola · 1996
Earlier work this paper cites.
Sift: Predicting amino acid changes that affect protein function
P. C. Ng and S. Henikoff · 2003
Earlier work this paper cites.
Pca-sift: A more distinctive representation for local image descriptors
Y. Ke and R. Sukthankar · 2004
Earlier work this paper cites.
Surf: Speeded up robust features
H. Bay, T. Tuytelaars, and L. Van Gool · 2006
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Image resolution enhancement by using discrete and stationary wavelet decomposition
H. Demirel and G. Anbarjafari · 2010
Earlier work this paper cites.
Orb: An efficient alternative to sift or surf
E. Rublee, V. Rabaud, K. Konolige, and G. Bradski · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Cats and dogs
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. V. Jawahar · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
J. Krause, M. Stark, J. Deng, and L. Fei-Fei · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
Spatial pyramid pooling in deep convolutional networks for visual recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Efficient piecewise training of deep structured models for semantic segmentation
G. Lin, C. Shen, A. Van Den Hengel, and I. Reid · 2016
Earlier work this paper cites.
Understanding the effective receptive field in deep convolutional neural networks
W. Luo, Y. Li, R. Urtasun, and R. S. Zemel · 2016
Earlier work this paper cites.
A benchmark dataset and evaluation methodology for video object segmentation
F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Earlier work this paper cites.
Multi-scale context aggregation by dilated convolutions
F. Yu and V. Koltun · 2016
Earlier work this paper cites.
Rethinking atrous convolution for semantic image segmentation
L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Earlier work this paper cites.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Earlier work this paper cites.
Deep laplacian pyramid networks for fast and accurate super-resolution
W.-S. Lai, J.-B. Huang, N. Ahuja, and M.-H. Yang · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Earlier work this paper cites.
Dual attention networks for multimodal reasoning and matching
H. Nam, J.-W. Ha, and J. Kim · 2017
Cited alongside, same era.
The 2017 davis challenge on video object segmentation
J. Pont-Tuset, F. Perazzi, S. Caelles, P. Arbeláez, A. Sorkine-Hornung, and L. Van Gool · 2017
Cited alongside, same era.
Dynamic routing between capsules
S. Sabour, N. Frosst, and G. E. Hinton · 2017
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra · 2017
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. Alemi · 2017
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2020
Later among the works it cites.
Grafit: Learning fine-grained image representations with coarse labels
H. Touvron, A. Sablayrolles, M. Douze, M. Cord, and H. Jégou · 2020
Later among the works it cites.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. S. Torr, and L. Zhang · 2020
Later among the works it cites.
Crossvit: Cross-attention multi-scale vision transformer for image classification
C.-F. Chen, Q. Fan, and R. Panda · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Cited alongside, same era.
Dilated residual networks
F. Yu, V. Koltun, and T. Funkhouser · 2017
Cited alongside, same era.
Pyramid scene parsing network
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia · 2017
Cited alongside, same era.
Scene parsing through ade20k dataset
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba · 2017
Cited alongside, same era.
Cascade r-cnn: Delving into high quality object detection
Z. Cai and N. Vasconcelos · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
X. Chen, S. Xie, and K. He · 2021
Closest in time.
Twins: Revisiting spatial attention design in vision transformers
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen · 2021
Closest in time.
Conditional positional encodings for vision transformers
X. Chu, Z. Tian, B. Zhang, X. Wang, X. Wei, H. Xia, and C. Shen · 2021
Closest in time.
Convit: Improving vision transformers with soft convolutional inductive biases
S. d’Ascoli, H. Touvron, M. Leavitt, A. Morcos, G. Biroli, and L. Sagun · 2021
Closest in time.
Repmlp: Re-parameterizing convolutions into fully-connected layers for image recognition
X. Ding, X. Zhang, J. Han, and G. Ding · 2021
Closest in time.
Xcit: Cross-covariance image transformers
A. El-Nouby, H. Touvron, M. Caron, P. Bojanowski, M. Douze, A. Joulin, I. Laptev, N. Neverova, G. Synnaeve, J. Verbeek, et al · 2021
Closest in time.
Levit: a vision transformer in convnet’s clothing for faster inference
B. Graham, A. El-Nouby, H. Touvron, P. Stock, A. Joulin, H. Jégou, and M. Douze · 2021
Closest in time.
Beyond self-attention: External attention using two linear layers for visual tasks
M.-H. Guo, Z.-N. Liu, T.-J. Mu, and S.-M. Hu · 2021
Closest in time.
K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang · 2021
Closest in time.
Gauge equivariant transformer
L. He, Y. Dong, Y. Wang, D. Tao, and Z. Lin · 2021
Closest in time.
Rethinking spatial dimensions of vision transformers
B. Heo, S. Yun, D. Han, S. Chun, J. Choe, and S. J. Oh · 2021
Closest in time.
Rethinking spatial dimensions of vision transformers
B. Heo, S. Yun, D. Han, S. Chun, J. Choe, and S. J. Oh · 2021
Closest in time.
Shuffle transformer: Rethinking spatial shuffle for vision transformer
Z. Huang, Y. Ben, G. Luo, P. Cheng, G. Yu, and B. Fu · 2021
Closest in time.
Localvit: Bringing locality to vision transformers
Y. Li, K. Zhang, J. Cao, R. Timofte, and L. Van Gool · 2021
Closest in time.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Closest in time.
Do you even need attention? a stack of feed-forward layers does surprisingly well on imagenet
L. Melas-Kyriazi · 2021
Closest in time.
Conformer: Local features coupling global representations for visual recognition
Z. Peng, W. Huang, S. Gu, L. Xie, Y. Wang, J. Jiao, and Q. Ye · 2021
Closest in time.
Bottleneck transformers for visual recognition
A. Srinivas, T.-Y. Lin, N. Parmar, J. Shlens, P. Abbeel, and A. Vaswani · 2021
Closest in time.
Mlp-mixer: An all-mlp architecture for vision
I. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, D. Keysers, J. Uszkoreit, M. Lucic, and A. Dosovitskiy · 2021
Closest in time.
Resmlp: Feedforward networks for image classification with data-efficient training
H. Touvron, P. Bojanowski, M. Caron, M. Cord, A. El-Nouby, E. Grave, A. Joulin, G. Synnaeve, J. Verbeek, and H. Jégou · 2021
Closest in time.
Going deeper with image transformers
H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jégou · 2021
Closest in time.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2021
Closest in time.
Cvt: Introducing convolutions to vision transformers
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang · 2021
Closest in time.
So-vit: Mind visual tokens for vision transformer
J. Xie, R. Zeng, Q. Wang, Z. Zhou, and P. Li · 2021
Closest in time.
Co-scale conv-attentional image transformers
W. Xu, Y. Xu, T. Chang, and Z. Tu · 2021
Closest in time.
Contnet: Why not use convolution and transformer at the same time?
H. Yan, Z. Li, W. Li, C. Wang, M. Wu, and C. Zhang · 2021
Closest in time.
Incorporating convolution designs into visual transformers
K. Yuan, S. Guo, Z. Liu, A. Zhou, F. Yu, and W. Wu · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z. Jiang, F. E. Tay, J. Feng, and S. Yan · 2021
Closest in time.
Deformable detr: Deformable transformers for end-to-end object detection
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai · 2021
Closest in time.