Fetching the paper…
Reading the bibliography…
Vision Transformer (ViT) architectures are becoming increasingly popular and widely employed to tackle computer vision applications.
“Pruning versus clipping in neural networks”
Steven Janowsky · 1989
Earlier work this paper cites.
“A simple procedure for pruning back-propagation trained neural networks”
Ehud Karnin · 1990
Earlier work this paper cites.
“Network flows: Theory, algorithms, and applications”
Gary Waissi · 1994
Earlier work this paper cites.
“Stochastic local search: Foundations and applications”
Holger Hoos and Thomas Stützle · 2004
Earlier work this paper cites.
“Random features for large-scale kernel machines”
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate”
Dzmitry Bahdanau, Kyunghyun Cho and Yoshua Bengio · 2014
Earlier work this paper cites.
“Microsoft coco: Common objects in context”
Tsung-Yi Lin et al · 2014
Earlier work this paper cites.
“Distilling the knowledge in a neural network”
Geoffrey Hinton, Oriol Vinyals and Jeff Dean · 2015
Earlier work this paper cites.
“Stanford neural machine translation systems for spoken language domains”
Minh-Thang Luong and Christopher Manning · 2015
Earlier work this paper cites.
“ImageNet Large Scale Visual Recognition Challenge”
Olga Russakovsky et al · 2015
Earlier work this paper cites.
“Findings of the 2016 conference on machine translation”
Ondřej Bojar et al · 2016
Earlier work this paper cites.
“Deformable convolutional networks”
Jifeng Dai et al · 2017
Earlier work this paper cites.
“Mask r-cnn”
Kaiming He, Georgia Gkioxari, Piotr Dollár and Ross Girshick · 2017
Earlier work this paper cites.
“Focal loss for dense object detection”
Tsung-Yi Lin et al · 2017
Earlier work this paper cites.
“Learning efficient convolutional networks through network slimming”
Zhuang Liu et al · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Scene parsing through ade20k dataset”
Bolei Zhou et al · 2017
Earlier work this paper cites.
“Cascade r-cnn: Delving into high quality object detection”
Zhaowei Cai and Nuno Vasconcelos · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Earlier work this paper cites.
“Non-local neural networks”
Xiaolong Wang, Ross Girshick, Abhinav Gupta and Kaiming He · 2018
Earlier work this paper cites.
“Unified perceptual parsing for scene understanding”
Tete Xiao et al · 2018
Earlier work this paper cites.
“Panoptic feature pyramid networks”
Alexander Kirillov, Ross Girshick, Kaiming He and Piotr Dollár · 2019
Earlier work this paper cites.
“Roberta: A robustly optimized bert pretraining approach”
Yinhan Liu et al · 2019
Earlier work this paper cites.
“On the Relationship between Self-Attention and Convolutional Layers”
Jean-Baptiste Cordonnier, Andreas Loukas and Martin Jaggi · 2020
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale”
Alexey Dosovitskiy et al · 2020
Earlier work this paper cites.
“Transformers are rnns: Fast autoregressive transformers with linear attention”
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas and François Fleuret · 2020
Earlier work this paper cites.
“Dreaming to distill: Data-free knowledge transfer via deepinversion”
Hongxu Yin et al · 2020
Earlier work this paper cites.
“Dynamic graph message passing networks”
Li Zhang, Dan Xu, Anurag Arnab and Philip Torr · 2020
Earlier work this paper cites.
“Mix and match: A novel fpga-centric deep neural network quantization framework”
Sung-En Chang et al · 2021
Earlier work this paper cites.
“Multiscale vision transformers”
Haoqi Fan et al · 2021
Cited alongside, same era.
“All tokens matter: Token labeling for training better vision transformers”
Zi-Hang Jiang et al · 2021
Cited alongside, same era.
“Involution: Inverting the inherence of convolution for visual recognition”
Duo Li et al · 2021
Cited alongside, same era.
“Swin transformer: Hierarchical vision transformer using shifted windows”
Ze Liu et al · 2021
Cited alongside, same era.
“Post-training quantization for vision transformer”
Zhenhua Liu et al · 2021
Cited alongside, same era.
“Soft: Softmax-free transformer with linear complexity”
Jiachen Lu et al · 2021
Cited alongside, same era.
“Co-advise: Cross inductive bias distillation”
Sucheng Ren et al · 2022
Later among the works it cites.
“Patch slimming for efficient vision transformers”
Yehui Tang et al · 2022
Later among the works it cites.
“Efficient transformers: A survey”
Yi Tay, Mostafa Dehghani, Dara Bahri and Donald Metzler · 2022
Later among the works it cites.
“Three Things Everyone Should Know About Vision Transformers”
Hugo Touvron et al · 2022
Later among the works it cites.
“Pvt v2: Improved baselines with pyramid vision transformer”
Wenhai Wang et al · 2022
Later among the works it cites.
“Tinyvit: Fast pretraining distillation for small vision transformers”
Kan Wu et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Dynamicvit: Efficient vision transformers with dynamic token sparsification”
Yongming Rao et al · 2021
Cited alongside, same era.
“Mlp-mixer: An all-mlp architecture for vision”
Ilya Tolstikhin et al · 2021
Cited alongside, same era.
“Training data-efficient image transformers & distillation through attention”
Hugo Touvron et al · 2021
Cited alongside, same era.
“Pyramid vision transformer: A versatile backbone for dense prediction without convolutions”
Wenhai Wang et al · 2021
Cited alongside, same era.
Mingjian Zhu, Yehui Tang and Kai Han · 2021
Cited alongside, same era.
“DearKD: data-efficient early knowledge distillation for vision transformers”
Xianing Chen et al · 2022
Cited alongside, same era.
“A-vit: Adaptive tokens for efficient vision transformer”
Hongxu Yin et al · 2022
Later among the works it cites.
“Width & Depth Pruning for Vision Transformers”
Fang Yu et al · 2022
Later among the works it cites.
“Metaformer is actually what you need for vision”
Weihao Yu et al · 2022
Later among the works it cites.
“Neural window fully-connected crfs for monocular depth estimation”
Weihao Yuan et al · 2022
Later among the works it cites.
“Ptq4vit: Post-training quantization for vision transformers with twin uniform quantization”
Zhihang Yuan et al · 2022
Later among the works it cites.
“Scaling vision transformers”
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby and Lucas Beyer · 2022
Later among the works it cites.
“Minivit: Compressing vision transformers with weight multiplexing”
Jinnian Zhang et al · 2022
Later among the works it cites.
“Hydra Attention: Efficient Attention with Many Heads”
Daniel Bolya et al · 2023
Closest in time.
“Token Merging: Your ViT but Faster”
Daniel Bolya et al · 2023
Closest in time.
“EfficientViT: Lightweight Multi-Scale Attention for High-Resolution Dense Prediction”
Han Cai et al · 2023
Closest in time.
“Monarch Mixer: A simple sub-quadratic GEMM-based architecture”
Daniel Fu et al · 2023
Closest in time.
“Supervised masked knowledge distillation for few-shot transformers”
Han Lin et al · 2023
Closest in time.
“EcoFormer: Energy-Saving Attention with Linear Complexity”, 2023
Jing Liu et al · 2023
Closest in time.
“NoisyQuant: Noisy Bias-Enhanced Post-Training Activation Quantization for Vision Transformers”
Yijiang Liu et al · 2023
Closest in time.
“METER: a mobile vision transformer architecture for monocular depth estimation”
Lorenzo Papa, Paolo Russo and Irene Amerini · 2023
Closest in time.
“Dynamic spatial sparsification for efficient vision transformers and convolutional neural networks”
Yongming Rao et al · 2023
Closest in time.
“Image as a Foreign Language: BEiT Pretraining for Vision and Vision-Language Tasks”
Wenhui Wang et al · 2023
Closest in time.
“Global Vision Transformer Pruning With Hessian-Aware Saliency”
Huanrui Yang et al · 2023
Closest in time.
“Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference”
Haoran You et al · 2023
Closest in time.
“Boost Vision Transformer with GPU-Friendly Sparsity and Quantization”
Chong Yu, Tao Chen, Zhongxue Gan and Jiayuan Fan · 2023
Closest in time.
“X-Pruner: eXplainable Pruning for Vision Transformers”
Lu Yu and Wei Xiang · 2023
Closest in time.
“A Survey on Efficient Training of Transformers”, 2023, pp. 6823–6831
Bohan Zhuang et al · 2023
Closest in time.