Fetching the paper…
Reading the bibliography…
The emergence of vision transformers (ViTs) in image classification has shifted the methodologies for visual representation learning.
Holschneider M, Kronland-Martinet R, Morlet J, Tchamitchian P (1990) A real-time algorithm for signal analysis with the help of the wavelet transform. In: Wavelets: Time-Frequency Methods and Phase Space Proceedings of the International Conference, Marseille, France, December 14–18, 1987, Springer, pp 286–297
1987
Earlier work this paper cites.
Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25
2012
Earlier work this paper cites.
Lin TY, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Dollár P, Zitnick CL (2014) Microsoft coco: Common objects in context. In: European Conference on Computer Vision, Springer, pp 740–755
2014
Earlier work this paper cites.
Mottaghi R, Chen X, Liu X, Cho NG, Lee SW, Fidler S, Urtasun R, Yuille A (2014) The role of context for object detection and semantic segmentation in the wild. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 891–898
2014
Earlier work this paper cites.
Chen L, Papandreou G, Kokkinos I, Murphy K, Yuille AL (2015) Semantic image segmentation with deep convolutional nets and fully connected CRFs. In: International Conference on Learning Representations
2015
Earlier work this paper cites.
Chen LC, Papandreou G, Kokkinos I, Murphy K, Yuille AL (2015) Semantic image segmentation with deep convolutional nets and fully connected CRFs. In: International Conference on Learning Representations
2015
Earlier work this paper cites.
Liu Z, Li X, Luo P, Loy CC, Tang X (2015) Semantic image segmentation via deep parsing network. In: IEEE International Conference on Computer Vision, pp 1377–1385
2015
Earlier work this paper cites.
Long J, Shelhamer E, Darrell T (2015) Fully convolutional networks for semantic segmentation. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 3431–3440
2015
Earlier work this paper cites.
Noh H, Hong S, Han B (2015) Learning deconvolution network for semantic segmentation. In: IEEE International Conference on Computer Vision, pp 1520–1528
2015
Earlier work this paper cites.
Ronneberger O, Fischer P, Brox T (2015) U-net: Convolutional networks for biomedical image segmentation. In: International Conference on Medical Image Computing and Computer Assisted Intervention, Springer, pp 234–241
2015
Earlier work this paper cites.
Russakovsky O, Deng J, Su H, Krause J, Satheesh S, Ma S, Huang Z, Karpathy A, Khosla A, Bernstein M, et al. (2015) Imagenet large scale visual recognition challenge. International Journal of Computer Vision 115:211–252
2015
Earlier work this paper cites.
Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition. In: International Conference on Learning Representations
2015
Earlier work this paper cites.
Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A (2015) Going deeper with convolutions. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 1–9
2015
Earlier work this paper cites.
Zheng S, Jayasumana S, Romera-Paredes B, Vineet V, Su Z, Du D, Huang C, Torr PHS (2015) Conditional random fields as recurrent neural networks. In: IEEE International Conference on Computer Vision, pp 1529–1537
2015
Earlier work this paper cites.
He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 770–778
2016
Earlier work this paper cites.
Larsson G, Maire M, Shakhnarovich G (2016) Fractalnet: Ultra-deep neural networks without residuals. In: International Conference on Learning Representations
2016
Earlier work this paper cites.
Yu F, Koltun V (2016) Multi-scale context aggregation by dilated convolutions. In: International Conference on Learning Representations
2016
Earlier work this paper cites.
Badrinarayanan V, Kendall A, Cipolla R (2017) Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence 39(12):2481–2495
2017
Earlier work this paper cites.
Chen LC, Papandreou G, Kokkinos I, Murphy K, Yuille AL (2017) Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Transactions on Pattern Analysis and Machine Intelligence 40(4):834–848
2017
Earlier work this paper cites.
Chollet F (2017) Xception: Deep learning with depthwise separable convolutions. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 1251–1258
2017
Earlier work this paper cites.
He K, Gkioxari G, Dollár P, Girshick R (2017) Mask r-cnn. In: IEEE International Conference on Computer Vision, pp 2961–2969
2017
Earlier work this paper cites.
Huang G, Liu Z, Van Der Maaten L, Weinberger KQ (2017) Densely connected convolutional networks. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 4700–4708
2017
Earlier work this paper cites.
Lin TY, Dollár P, Girshick RB, He K, Hariharan B, Belongie SJ (2017) Feature pyramid networks for object detection. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 2117–2125
2017
Earlier work this paper cites.
Lin TY, Goyal P, Girshick R, He K, Dollar P (2017) Focal loss for dense object detection. In: IEEE International Conference on Computer Vision, pp 2980–2988
2017
Earlier work this paper cites.
Peng C, Zhang X, Yu G, Luo G, Sun J (2017) Large kernel matters–improve semantic segmentation by global convolutional network. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 4353–4361
2017
Earlier work this paper cites.
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I (2017) Attention is all you need. In: Advances in Neural Information Processing Systems, pp 5998–6008
2017
Earlier work this paper cites.
Xie S, Girshick R, Dollár P, Tu Z, He K (2017) Aggregated residual transformations for deep neural networks. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 1492–1500
2017
Earlier work this paper cites.
Zhang H, Cisse M, Dauphin YN, Lopez-Paz D (2017) mixup: Beyond empirical risk minimization. In: International Conference on Learning Representations
2017
Earlier work this paper cites.
Zhao H, Shi J, Qi X, Wang X, Jia J (2017) Pyramid scene parsing network. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 2881–2890
2017
Cited alongside, same era.
Chen LC, Zhu Y, Papandreou G, Schroff F, Adam H (2018) Encoder-decoder with atrous separable convolution for semantic image segmentation. In: European Conference on Computer Vision, pp 801–818
2018
Cited alongside, same era.
Hu J, Shen L, Sun G (2018) Squeeze-and-excitation networks. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 7132–7141
2018
Cited alongside, same era.
Sandler M, Howard A, Zhu M, Zhmoginov A, Chen LC (2018) Mobilenetv2: Inverted residuals and linear bottlenecks. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 4510–4520
2018
Cited alongside, same era.
Veličković P, Cucurull G, Casanova A, Romero A, Lio P, Bengio Y (2018) Graph attention networks. In: International Conference on Learning Representations
OpenMMLab (2020) mmsegmentation. https://github.com/open-mmlab/mmsegmentation
2020
Later among the works it cites.
Wang H, Zhu Y, Green B, Adam H, Yuille A, Chen LC (2020) Axial-deeplab: Stand-alone axial-attention for panoptic segmentation. In: European Conference on Computer Vision, Springer, pp 108–126
2020
Later among the works it cites.
Yuan Y, Chen X, Wang J (2020) Object-contextual representations for semantic segmentation. In: European Conference on Computer Vision, Springer, pp 173–190
2020
Later among the works it cites.
Zhang L, Xu D, Arnab A, Torr PH (2020) Dynamic graph message passing networks. In: IEEE Conference on Computer Vision and Pattern Recognition
2020
Later among the works it cites.
Zhao H, Jia J, Koltun V (2020) Exploring self-attention for image recognition. In: IEEE Conference on Computer Vision and Pattern Recognition
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Wang X, Girshick R, Gupta A, He K (2018) Non-local neural networks. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 7794–7803
2018
Cited alongside, same era.
Yang M, Yu K, Zhang C, Li Z, Yang K (2018) Denseaspp for semantic segmentation in street scenes. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 3684–3692
2018
Cited alongside, same era.
Yu C, Wang J, Peng C, Gao C, Yu G, Sang N (2018) Bisenet: Bilateral segmentation network for real-time semantic segmentation. In: European Conference on Computer Vision, pp 325–341
2018
Cited alongside, same era.
Yuan Y, Wang J (2018) Ocnet: Object context network for scene parsing. arXiv preprint
2018
Cited alongside, same era.
Zhao H, Zhang Y, Liu S, Shi J, Change Loy C, Lin D, Jia J (2018) Psanet: Point-wise spatial attention network for scene parsing. In: European Conference on Computer Vision, pp 267–283
2018
Cited alongside, same era.
Bello I, Zoph B, Vaswani A, Shlens J, Le QV (2019) Attention augmented convolutional networks. In: IEEE International Conference on Computer Vision, pp 10076–10085
2019
Cited alongside, same era.
Cao Y, Xu J, Lin S, Wei F, Hu H (2019) Gcnet: Non-local networks meet squeeze-excitation networks and beyond. In: IEEE International Conference on Computer Vision workshops, pp 0–0
2019
Cited alongside, same era.
Later among the works it cites.
Bai S, Torr P, et al. (2021) Visual parser: Representing part-whole hierarchies with transformers. arXiv preprint
2021
Later among the works it cites.
Chen CFR, Fan Q, Panda R (2021) Crossvit: Cross-attention multi-scale vision transformer for image classification. In: IEEE International Conference on Computer Vision, pp 357–366
2021
Later among the works it cites.
Chu X, Tian Z, Wang Y, Zhang B, Ren H, Wei X, Xia H, Shen C (2021) Twins: Revisiting the design of spatial attention in vision transformers. In: Advances in Neural Information Processing Systems, pp 9355–9366
2021
Later among the works it cites.
Dai Z, Liu H, Le QV, Tan M (2021) Coatnet: Marrying convolution and attention for all data sizes. In: Advances in Neural Information Processing Systems, pp 3965–3977
2021
Later among the works it cites.
Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, et al. (2021) An image is worth 16x16 words: Transformers for image recognition at scale. In: International Conference on Learning Representations
2021
Later among the works it cites.
d’Ascoli S, Touvron H, Leavitt ML, Morcos AS, Biroli G, Sagun L (2021) Convit: Improving vision transformers with soft convolutional inductive biases. In: International Conference on Machine Learning, PMLR, pp 2286–2296
2021
Later among the works it cites.
Han K, Xiao A, Wu E, Guo J, Xu C, Wang Y (2021) Transformer in transformer. Advances in neural information processing systems 34:15908–15919
2021
Later among the works it cites.
Heo B, Yun S, Han D, Chun S, Choe J, Oh SJ (2021) Rethinking spatial dimensions of vision transformers. In: IEEE International Conference on Computer Vision, pp 11936–11945
2021
Later among the works it cites.
Li J, Yan Y, Liao S, Yang X, Shao L (????) Local-to-global self-attention in vision transformers. arxiv 2021. arXiv preprint arXiv:210704735
2021
Later among the works it cites.
Touvron H, Cord M, Douze M, Massa F, Sablayrolles A, Jégou H (2021) Training data-efficient image transformers & distillation through attention. In: International Conference on Machine Learning, PMLR, pp 10347–10357
2021
Later among the works it cites.
Wang W, Xie E, Li X, Fan DP, Song K, Liang D, Lu T, Luo P, Shao L (2021) Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In: IEEE International Conference on Computer Vision, pp 568–578
2021
Later among the works it cites.
Wu H, Xiao B, Codella N, Liu M, Dai X, Yuan L, Zhang L (2021) Cvt: Introducing convolutions to vision transformers. In: IEEE International Conference on Computer Vision, pp 22–31
2021
Later among the works it cites.
Yan H, Li Z, Li W, Wang C, Wu M, Zhang C (2021) Contnet: Why not use convolution and transformer at the same time? arXiv preprint
2021
Later among the works it cites.
Zheng S, Lu J, Zhao H, Zhu X, Luo Z, Wang Y, Fu Y, Feng J, Xiang T, Torr PH, Zhang L (2021) Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 6881–6890
2021
Later among the works it cites.
Bao H, Dong L, Piao S, Wei F (2022) Beit: Bert pre-training of image transformers. In: International Conference on Learning Representations
2022
Closest in time.
Chen CF, Panda R, Fan Q (2022) Regionvit: Regional-to-local attention for vision transformers. In: International Conference on Learning Representations
2022
Closest in time.
Dao T, Fu D, Ermon S, Rudra A, Ré C (2022) Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems 35:16344–16359
2022
Closest in time.
Song CH, Han HJ, Avrithis Y (2022) All the attention you need: Global-local, spatial-channel attention for image retrieval. In: IEEE Winter Conference on Applications of Computer Vision, pp 2754–2763
2022
Closest in time.
Touvron H, Cord M, Jégou H (2022) Deit iii: Revenge of the vit. In: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXIV, Springer, pp 516–533
2022
Closest in time.
Wang W, Xie E, Li X, Fan DP, Song K, Liang D, Lu T, Luo P, Shao L (2022) Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media 8(3):415–424
2022
Closest in time.
Yang C, Qiao S, Yu Q, Yuan X, Zhu Y, Yuille A, Adam H, Chen LC (2022) Moat: Alternating mobile convolution and attention brings strong vision models. In: International Conference on Learning Representations
2022
Closest in time.
Hassani A, Walton S, Li J, Li S, Shi H (2023) Neighborhood attention transformer. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 6185–6194
2023
Closest in time.