Fetching the paper…
Reading the bibliography…
Vision transformers have shown great potential in various computer vision tasks owing to their strong capability to model long-range dependency using the self-attention mechanism.
He K, Zhang X, Ren S, Sun J (2015) Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 37(9):1904–1916
1916
Earlier work this paper cites.
Adelson EH, Anderson CH, Bergen JR, Burt PJ, Ogden JM (1984) Pyramid methods in image processing. RCA engineer 29(6):33–41
1984
Earlier work this paper cites.
Burt PJ, Adelson EH (1987) The laplacian pyramid as a compact image code. In: Readings in computer vision, Elsevier, pp 671–679
1987
Earlier work this paper cites.
LeCun Y, Bengio Y, et al (1995) Convolutional networks for images, speech, and time series. The handbook of brain theory and neural networks 3361(10):1995
1995
Earlier work this paper cites.
Olkkonen H, Pesola P (1996) Gaussian pyramid wavelet transform for multiresolution analysis of images. Graphical Models and Image Processing 58(4):394–398
1996
Earlier work this paper cites.
Ng PC, Henikoff S (2003) Sift: Predicting amino acid changes that affect protein function. Nucleic acids research 31(13):3812–3814
2003
Earlier work this paper cites.
Ke Y, Sukthankar R (2004) Pca-sift: A more distinctive representation for local image descriptors. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, IEEE, vol 2, pp II–II
2004
Earlier work this paper cites.
Bay H, Tuytelaars T, Van Gool L (2006) Surf: Speeded up robust features. In: Proceedings of the European conference on computer vision, Springer, pp 404–417
2006
Earlier work this paper cites.
Nilsback ME, Zisserman A (2008) Automated flower classification over a large number of classes. In: Indian Conference on Computer Vision, Graphics and Image Processing
2008
Earlier work this paper cites.
Deng J, Dong W, Socher R, Li LJ, Li K, Fei-Fei L (2009) Imagenet: A large-scale hierarchical image database. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Ieee, pp 248–255
2009
Earlier work this paper cites.
Krizhevsky A, Hinton G, et al (2009) Learning multiple layers of features from tiny images
2009
Earlier work this paper cites.
Demirel H, Anbarjafari G (2010) Image resolution enhancement by using discrete and stationary wavelet decomposition. IEEE transactions on image Processing 20(5):1458–1460
2010
Earlier work this paper cites.
Rublee E, Rabaud V, Konolige K, Bradski G (2011) Orb: An efficient alternative to sift or surf. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, Ieee, pp 2564–2571
2011
Earlier work this paper cites.
Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems 25:1097–1105
2012
Earlier work this paper cites.
Parkhi OM, Vedaldi A, Zisserman A, Jawahar CV (2012) Cats and dogs. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2012
Earlier work this paper cites.
Krause J, Stark M, Deng J, Fei-Fei L (2013) 3d object representations for fine-grained categorization. In: 4th International IEEE Workshop on 3D Representation and Recognition (3dRR-13), Sydney, Australia
2013
Earlier work this paper cites.
Lin TY, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Dollár P, Zitnick CL (2014) Microsoft coco: Common objects in context. In: Proceedings of the European conference on computer vision, Springer, pp 740–755
2014
Earlier work this paper cites.
Zeiler MD, Fergus R (2014) Visualizing and understanding convolutional networks. In: Proceedings of the European conference on computer vision, Springer, pp 818–833
2014
Earlier work this paper cites.
LeCun Y, Bengio Y, Hinton G (2015) Deep learning. nature 521(7553):436–444
2015
Earlier work this paper cites.
Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition
2015
Earlier work this paper cites.
Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A (2015) Going deeper with convolutions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 1–9
2015
Earlier work this paper cites.
Ba JL, Kiros JR, Hinton GE (2016) Layer normalization. arXiv preprint arXiv:160706450
2016
Earlier work this paper cites.
He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 770–778
2016
Earlier work this paper cites.
Lin G, Shen C, Van Den Hengel A, Reid I (2016) Efficient piecewise training of deep structured models for semantic segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 3194–3203
2016
Earlier work this paper cites.
Luo W, Li Y, Urtasun R, Zemel RS (2016) Understanding the effective receptive field in deep convolutional neural networks. Advances in neural information processing systems 29:4898–4906
2016
Earlier work this paper cites.
Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z (2016) Rethinking the inception architecture for computer vision. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 2818–2826
2016
Earlier work this paper cites.
Yu F, Koltun V (2016) Multi-scale context aggregation by dilated convolutions. In: International Conference on Learning Representations
2016
Earlier work this paper cites.
Chen LC, Papandreou G, Schroff F, Adam H (2017) Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:170605587
2017
Earlier work this paper cites.
He K, Gkioxari G, Dollár P, Girshick R (2017) Mask r-cnn. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 2961–2969
2017
Earlier work this paper cites.
Howard AG, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Andreetto M, Adam H (2017) Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:170404861
2017
Earlier work this paper cites.
Huang G, Liu Z, Van Der Maaten L, Weinberger KQ (2017) Densely connected convolutional networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 4700–4708
2017
Earlier work this paper cites.
Lai WS, Huang JB, Ahuja N, Yang MH (2017) Deep laplacian pyramid networks for fast and accurate super-resolution. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 624–632
2017
Earlier work this paper cites.
Lin TY, Dollár P, Girshick R, He K, Hariharan B, Belongie S (2017) Feature pyramid networks for object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 2117–2125
2017
Cited alongside, same era.
Loshchilov I, Hutter F (2017) Decoupled weight decay regularization. arXiv preprint arXiv:171105101
2017
Cited alongside, same era.
Sabour S, Frosst N, Hinton GE (2017) Dynamic routing between capsules. Advances in neural information processing systems 30
2017
Cited alongside, same era.
Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D (2017) Grad-cam: Visual explanations from deep networks via gradient-based localization. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 618–626
2017
Cited alongside, same era.
Fan H, Xiong B, Mangalam K, Li Y, Yan Z, Malik J, Feichtenhofer C (2021) Multiscale vision transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 6824–6835
2021
Later among the works it cites.
Graham B, El-Nouby A, Touvron H, Stock P, Joulin A, Jégou H, Douze M (2021) Levit: a vision transformer in convnet’s clothing for faster inference. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 12259–12269
2021
Later among the works it cites.
Han K, Xiao A, Wu E, Guo J, Xu C, Wang Y (2021) Transformer in transformer. Advances in Neural Information Processing Systems 34:15908–15919
2021
Later among the works it cites.
He L, Dong Y, Wang Y, Tao D, Lin Z (2021) Gauge equivariant transformer. Advances in Neural Information Processing Systems 34:27331–27343
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I (2017) Attention is all you need. Advances in neural information processing systems 30
2017
Cited alongside, same era.
Yu F, Koltun V, Funkhouser T (2017) Dilated residual networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 472–480
2017
Cited alongside, same era.
Zhao H, Shi J, Qi X, Wang X, Jia J (2017) Pyramid scene parsing network. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 2881–2890
2017
Cited alongside, same era.
Zhou B, Zhao H, Puig X, Fidler S, Barriuso A, Torralba A (2017) Scene parsing through ade20k dataset. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 633–641
2017
Cited alongside, same era.
Cai Z, Vasconcelos N (2018) Cascade r-cnn: Delving into high quality object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 6154–6162
2018
Cited alongside, same era.
Loshchilov I, Hutter F (2018) Decoupled weight decay regularization. In: International Conference on Learning Representations
2018
Cited alongside, same era.
Sandler M, Howard A, Zhu M, Zhmoginov A, Chen LC (2018) Mobilenetv2: Inverted residuals and linear bottlenecks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 4510–4520
2018
Cited alongside, same era.
Heo B, Yun S, Han D, Chun S, Choe J, Oh SJ (2021) Rethinking spatial dimensions of vision transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 11936–11945
2021
Later among the works it cites.
Li Y, Zhang K, Cao J, Timofte R, Van Gool L (2021) Localvit: Bringing locality to vision transformers. arXiv preprint arXiv:210405707
2021
Later among the works it cites.
Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, Lin S, Guo B (2021) Swin transformer: Hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 10012–10022
2021
Later among the works it cites.
Peng Z, Huang W, Gu S, Xie L, Wang Y, Jiao J, Ye Q (2021) Conformer: Local features coupling global representations for visual recognition. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 367–376
2021
Later among the works it cites.
Pham H, Dai Z, Xie Q, Le QV (2021) Meta pseudo labels. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 11557–11568
2021
Later among the works it cites.
Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, et al (2021) Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, PMLR, pp 8748–8763
2021
Later among the works it cites.
Tang S, Gong R, Wang Y, Liu A, Wang J, Chen X, Yu F, Liu X, Song D, Yuille A, et al (2021) Robustart: Benchmarking robustness on architecture design and training techniques. arXiv preprint arXiv:210905211
2021
Later among the works it cites.
Tolstikhin IO, Houlsby N, Kolesnikov A, Beyer L, Zhai X, Unterthiner T, Yung J, Steiner A, Keysers D, Uszkoreit J, et al (2021) Mlp-mixer: An all-mlp architecture for vision. Advances in Neural Information Processing Systems 34:24261–24272
2021
Later among the works it cites.
Wu H, Xiao B, Codella N, Liu M, Dai X, Yuan L, Zhang L (2021) Cvt: Introducing convolutions to vision transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 22–31
2021
Later among the works it cites.
Xu Y, Zhang Q, Zhang J, Tao D (2021) Vitae: Vision transformer advanced by exploring intrinsic inductive bias. Advances in Neural Information Processing Systems 34:28522–28535
2021
Later among the works it cites.
Yan H, Li Z, Li W, Wang C, Wu M, Zhang C (2021) Contnet: Why not use convolution and transformer at the same time? arXiv preprint arXiv:210413497
2021
Later among the works it cites.
Yang J, Li C, Zhang P, Dai X, Xiao B, Yuan L, Gao J (2021) Focal self-attention for local-global interactions in vision transformers. Advances in Neural Information Processing Systems
2021
Later among the works it cites.
Yu H, Xu Y, Zhang J, Zhao W, Guan Z, Tao D (2021) Ap-10k: A benchmark for animal pose estimation in the wild. In: Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)
2021
Later among the works it cites.
Zhang P, Dai X, Yang J, Xiao B, Yuan L, Zhang L, Gao J (2021) Multi-scale vision longformer: A new vision transformer for high-resolution image encoding. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 2998–3008
2021
Later among the works it cites.
Zheng S, Lu J, Zhao H, Zhu X, Luo Z, Wang Y, Fu Y, Feng J, Xiang T, Torr PH, et al (2021) Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 6881–6890
2021
Later among the works it cites.
Chen Y, Dai X, Chen D, Liu M, Dong X, Yuan L, Liu Z (2022) Mobile-former: Bridging mobilenet and transformer. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 5270–5279
2022
Closest in time.
Guo MH, Lu CZ, Liu ZN, Cheng MM, Hu SM (2022) Visual attention network. arXiv preprint arXiv:220209741
2022
Closest in time.
He K, Chen X, Xie S, Li Y, Dollár P, Girshick R (2022) Masked autoencoders are scalable vision learners. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 16000–16009
2022
Closest in time.
Lee Y, Kim J, Willette J, Hwang SJ (2022) Mpvit: Multi-path vision transformer for dense prediction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 7287–7296
2022
Closest in time.
Liu Z, Hu H, Lin Y, Yao Z, Xie Z, Wei Y, Ning J, Cao Y, Zhang Z, Dong L, et al (2022) Swin transformer v2: Scaling up capacity and resolution. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 12009–12019
2022
Closest in time.
Wang W, Xie E, Li X, Fan DP, Song K, Liang D, Lu T, Luo P, Shao L (2022) Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media 8(3):415–424
2022
Closest in time.
Wei C, Fan H, Xie S, Wu CY, Yuille A, Feichtenhofer C (2022) Masked feature prediction for self-supervised visual pre-training. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 14668–14678
2022
Closest in time.
Xia Z, Pan X, Song S, Li LE, Huang G (2022) Vision transformer with deformable attention. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 4794–4803
2022
Closest in time.
Xie Z, Zhang Z, Cao Y, Lin Y, Bao J, Yao Z, Dai Q, Hu H (2022) Simmim: A simple framework for masked image modeling. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 9653–9663
2022
Closest in time.
Xu Y, Zhang J, Zhang Q, Tao D (2022) Vitpose: Simple vision transformer baselines for human pose estimation. Advances in Neural Information Processing Systems
2022
Closest in time.
Yu W, Luo M, Zhou P, Si C, Zhou Y, Wang X, Feng J, Yan S (2022) Metaformer is actually what you need for vision. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 10819–10829
2022
Closest in time.
Zhai X, Kolesnikov A, Houlsby N, Beyer L (2022) Scaling vision transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 12104–12113
2022
Closest in time.
Zhang Q, Xu Y, Zhang J, Tao D (2022) Vsa: Learning varied-size window attention in vision transformers. In: Proceedings of the European conference on computer vision
2022
Closest in time.