Fetching the paper…
Reading the bibliography…
While vision transformers have achieved impressive results, effectively and efficiently accelerating these models can further boost performances.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization, 2014
Yunchao Gong, Liu Liu, Ming Yang, and Lubomir Bourdev · 2014
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding, 2015
Song Han, Huizi Mao, and William J. Dally · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era, 2017
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Earlier work this paper cites.
Pytorch image models
Ross Wightman · 2019
Earlier work this paper cites.
End-to-end object detection with transformers, 2020
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Earlier work this paper cites.
Designing network design spaces
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár · 2020
Cited alongside, same era.
High-performance large-scale image recognition without normalization, 2021
Andrew Brock, Soham De, Samuel L. Smith, and Karen Simonyan · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Cited alongside, same era.
Per-pixel classification is not all you need for semantic segmentation, 2021
Bowen Cheng, Alexander G. Schwing, and Alexander Kirillov · 2021
Cited alongside, same era.
Twins: Revisiting the design of spatial attention in vision transformers, 2021
Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
All tokens matter: Token labeling for training better vision transformers
Zihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Yujun Shi, Xiaojie Jin, Anran Wang, and Jiashi Feng · 2021
Later among the works it cites.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh · 2021
Later among the works it cites.
Bottleneck transformers for visual recognition, 2021
Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani · 2021
Later among the works it cites.
Co-scale conv-attentional image transformers, 2021
Weijian Xu, Yifan Xu, Tyler Chang, and Zhuowen Tu · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet, 2021
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Cited alongside, same era.
Levit: a vision transformer in convnet’s clothing for faster inference, 2021
Ben Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze · 2021
Cited alongside, same era.
Cmt: Convolutional neural networks meet vision transformers, 2021
Jianyuan Guo, Kai Han, Han Wu, Chang Xu, Yehui Tang, Chunjing Xu, and Yunhe Wang · 2021
Cited alongside, same era.
Masked autoencoders are scalable vision learners, 2021
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2021
Cited alongside, same era.
Perceiver: General perception with iterative attention, 2021
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira · 2021
Cited alongside, same era.
Crossvit: Cross-attention multi-scale vision transformer for image classification, 2021a
Chun-Fu Chen, Quanfu Fan, and Rameswar Panda
Cited in the paper.
Chasing sparsity in vision transformers: An end-to-end exploration
Tianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan, Lei Zhang, and Zhangyang Wang
Cited in the paper.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip H.S. Torr, and Li Zhang · 2021
Later among the works it cites.
Deepvit: Towards deeper vision transformer, 2021
Daquan Zhou, Bingyi Kang, Xiaojie Jin, Linjie Yang, Xiaochen Lian, Zihang Jiang, Qibin Hou, and Jiashi Feng · 2021
Later among the works it cites.
Exploring plain vision transformer backbones for object detection, 2022
Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He · 2022
Closest in time.
EVit: Expediting vision transformers via token reorganizations
Youwei Liang, Chongjian GE, Zhan Tong, Yibing Song, Jue Wang, and Pengtao Xie · 2022
Closest in time.