Fetching the paper…
Reading the bibliography…
We attempt to reduce the computational costs in vision transformers (ViTs), which increase quadratically in the token number.
Deng J, Dong W, Socher R, et al (2009) Imagenet: A large-scale hierarchical image database. In: Int. Conf. Comput. Vis., pp 248–255
2009
Earlier work this paper cites.
Krizhevsky A, Hinton G, et al (2009) Learning multiple layers of features from tiny images. Toronto, ON, Canada
2009
Earlier work this paper cites.
Hendrycks D, Gimpel K (2016) Gaussian error linear units (gelus). arXiv preprint arXiv:160608415
2016
Earlier work this paper cites.
Dai J, Qi H, Xiong Y, et al (2017) Deformable convolutional networks. In: IEEE Conf. Comput. Vis. Pattern Recog., pp 764–773
2017
Earlier work this paper cites.
Howard AG, Zhu M, Chen B, et al (2017) Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:170404861
2017
Earlier work this paper cites.
Sun C, Shrivastava A, Singh S, et al (2017) Revisiting unreasonable effectiveness of data in deep learning era. In: Int. Conf. Comput. Vis., pp 843–852
2017
Earlier work this paper cites.
Vaswani A, Shazeer N, Parmar N, et al (2017) Attention is all you need. In: Adv. Neural Inform. Process. Syst
2017
Earlier work this paper cites.
Zhou B, Zhao H, Puig X, et al (2017) Scene parsing through ade20k dataset. In: IEEE Conf. Comput. Vis. Pattern Recog., pp 633–641
2017
Earlier work this paper cites.
Huang G, Chen D, Li T, et al (2018) Multi-scale dense networks for resource efficient image classification. In: Int. Conf. Learn. Represent
2018
Earlier work this paper cites.
Xiao T, Liu Y, Zhou B, et al (2018) Unified perceptual parsing for scene understanding. In: Eur. Conf. Comput. Vis., pp 418–434
2018
Earlier work this paper cites.
Yu J, Yang L, Xu N, et al (2018) Slimmable neural networks. In: Int. Conf. Learn. Represent
2018
Earlier work this paper cites.
Zhang X, Zhou X, Lin M, et al (2018) Shufflenet: An extremely efficient convolutional neural network for mobile devices. In: IEEE Conf. Comput. Vis. Pattern Recog., pp 6848–6856
2018
Earlier work this paper cites.
Cai H, Gan C, Wang T, et al (2019) Once-for-all: Train one network and specialize it for efficient deployment. In: Int. Conf. Learn. Represent
2019
Earlier work this paper cites.
Carion N, Massa F, Synnaeve G, et al (2020) End-to-end object detection with transformers. In: Eur. Conf. Comput. Vis., pp 213–229
2020
Earlier work this paper cites.
Dosovitskiy A, Beyer L, Kolesnikov A, et al (2020) An image is worth 16x16 words: Transformers for image recognition at scale. In: Int. Conf. Learn. Represent
2020
Earlier work this paper cites.
Huang L, Tan J, Liu J, et al (2020) Hand-transformer: non-autoregressive structured modeling for 3d hand pose estimation. In: Eur. Conf. Comput. Vis., pp 17–33
2020
Cited alongside, same era.
Arnab A, Dehghani M, Heigold G, et al (2021) Vivit: A video vision transformer. In: Int. Conf. Comput. Vis., pp 6836–6846
2021
Cited alongside, same era.
Bertasius G, Wang H, Torresani L (2021) Is space-time attention all you need for video understanding? In: Int. Conf. Mach. Learn
2021
Cited alongside, same era.
Chen CFR, Fan Q, Panda R (2021) Crossvit: Cross-attention multi-scale vision transformer for image classification. In: IEEE Conf. Comput. Vis. Pattern Recog., pp 357–366
2021
Cited alongside, same era.
Graham B, El-Nouby A, Touvron H, et al (2021) Levit: a vision transformer in convnet’s clothing for faster inference. In: Int. Conf. Comput. Vis., pp 12,259–12,269
Yuan L, Chen Y, Wang T, et al (2021) Tokens-to-token vit: Training vision transformers from scratch on imagenet. In: Int. Conf. Comput. Vis., pp 558–567
2021
Later among the works it cites.
Zheng S, Lu J, Zhao H, et al (2021) Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In: IEEE Conf. Comput. Vis. Pattern Recog., pp 6881–6890
2021
Later among the works it cites.
2021
Later among the works it cites.
Chavan A, Shen Z, Liu Z, et al (2022) Vision transformer slimming: Multi-dimension searching in continuous optimization space. In: IEEE Conf. Comput. Vis. Pattern Recog., pp 4931–4941
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
Han K, Xiao A, Wu E, et al (2021) Transformer in transformer. In: Adv. Neural Inform. Process. Syst., pp 15,908–15,919
2021
Cited alongside, same era.
Huang Z, Ben Y, Luo G, et al (2021) Shuffle transformer: Rethinking spatial shuffle for vision transformer. arXiv preprint arXiv:210603650
2021
Cited alongside, same era.
Jiang ZH, Hou Q, Yuan L, et al (2021) All tokens matter: Token labeling for training better vision transformers. In: Adv. Neural Inform. Process. Syst., pp 18,590–18,602
2021
Cited alongside, same era.
Khan S, Naseer M, Hayat M, et al (2021) Transformers in vision: A survey. ACM Computing Surveys
2021
Cited alongside, same era.
Liang J, Cao J, Sun G, et al (2021) Swinir: Image restoration using swin transformer. In: Int. Conf. Comput. Vis., pp 1833–1844
2021
Cited alongside, same era.
Liu Z, Lin Y, Cao Y, et al (2021) Swin transformer: Hierarchical vision transformer using shifted windows. In: Int. Conf. Comput. Vis., pp 10,012–10,022
2021
Cited alongside, same era.
Pan B, Panda R, Jiang Y, et al (2021) Ia-red 2
2021
Cited alongside, same era.
2022
Closest in time.
Han K, Wang Y, Chen H, et al (2022) A survey on vision transformer. IEEE Trans Pattern Anal Mach Intell
2022
Closest in time.
Li W, Wang X, Xia X, et al (2022) Sepvit: Separable vision transformer. arXiv preprint arXiv:220315380
2022
Closest in time.
Liang Y, GE C, Tong Z, et al (2022) Not all patches are what you need: Expediting vision transformers via token reorganizations. In: Int. Conf. Learn. Represent
2022
Closest in time.
Ren S, Zhou D, He S, et al (2022) Shunted self-attention via multi-scale token aggregation. In: IEEE Conf. Comput. Vis. Pattern Recog., pp 10,853–10,862
2022
Closest in time.
Tang Y, Han K, Wang Y, et al (2022) Patch slimming for efficient vision transformers. In: IEEE Conf. Comput. Vis. Pattern Recog., pp 12,165–12,174
2022
Closest in time.
Xia Z, Pan X, Song S, et al (2022) Vision transformer with deformable attention. In: IEEE Conf. Comput. Vis. Pattern Recog., pp 4794–4803
2022
Closest in time.
Xu Y, Zhang Z, Zhang M, et al (2022) Evo-vit: Slow-fast token evolution for dynamic vision transformer. In: AAAI Conf. Artificial Intelli., pp 2964–2972
2022
Closest in time.
Yin H, Vahdat A, Alvarez J, et al (2022) A-ViT: Adaptive tokens for efficient vision transformer. In: IEEE Conf. Comput. Vis. Pattern Recog., pp 10,809–10,818
2022
Closest in time.
Zamir SW, Arora A, Khan S, et al (2022) Restormer: Efficient transformer for high-resolution image restoration. In: IEEE Conf. Comput. Vis. Pattern Recog., pp 5728–5739
2022
Closest in time.
Zhu X, Su W, Lu L, et al (2022) Deformable detr: Deformable transformers for end-to-end object detection. In: Int. Conf. Learn. Represent
2022
Closest in time.