Fetching the paper…
Reading the bibliography…
Vision transformers (ViTs) have recently obtained success in many applications, but their intensive computation and heavy memory usage at both training and inference time limit their generalization.
Longformer: The long-document transformer
Beltagy, I.; Peters, M. E.; and Cohan, A. 2020 · 2004
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E.; and Zisserman, A. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A.; Hinton, G.; et al. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Deformable detr: Deformable transformers for end-to-end object detection
Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2020 · 2010
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Active bias: Training more accurate neural networks by emphasizing high variance samples
Chang, H.-S.; Learned-Miller, E.; and McCallum, A. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Not all samples are created equal: Deep learning with importance sampling
Katharopoulos, A.; and Fleuret, F. 2018 · 2018
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
Toneva, M.; Sordoni, A.; Combes, R. T. d.; Trischler, A.; Bengio, Y.; and Gordon, G. J. 2018 · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
Recht, B.; Roelofs, R.; Schmidt, L.; and Shankar, V. 2019 · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Earlier work this paper cites.
Learning texture transformer network for image super-resolution
Yang, F.; Yang, H.; Fu, J.; Lu, H.; and Guo, B. 2020 · 2020
Earlier work this paper cites.
T6D-Direct: Transformers for Multi-Object 6D Pose Direct Regression
Amini, A.; Periyasamy, A. S.; and Behnke, S. 2021 · 2021
Earlier work this paper cites.
BEiT: BERT Pre-Training of Image Transformers
Bao, H.; Dong, L.; and Wei, F. 2021 · 2021
Cited alongside, same era.
Crossvit: Cross-attention multi-scale vision transformer for image classification
Chen, C.-F. R.; Fan, Q.; and Panda, R. 2021 · 2021
Cited alongside, same era.
Up-detr: Unsupervised pre-training for object detection with transformers
Dai, Z.; Cai, B.; Lin, Y.; and Chen, J. 2021 · 2021
Cited alongside, same era.
TransVG: End-to-End Visual Grounding With Transformers
Deng, J.; Yang, Z.; Chen, T.; Zhou, W.; and Li, H. 2021 · 2021
Cited alongside, same era.
PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers
Dong, X.; Bao, J.; Zhang, T.; Chen, D.; Zhang, W.; Yuan, L.; Chen, D.; Wen, F.; and Yu, N. 2021 · 2021
Cited alongside, same era.
DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; and Hsieh, C.-J. 2021 · 2021
Later among the works it cites.
TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Ryoo, M. S.; Piergiovanni, A.; Arnab, A.; Dehghani, M.; and Angelova, A. 2021 · 2021
Later among the works it cites.
Bottleneck transformers for visual recognition
Srinivas, A.; Lin, T.-Y.; Parmar, N.; Shlens, J.; Abbeel, P.; and Vaswani, A. 2021 · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I.; Houlsby, N.; Kolesnikov, A.; Beyer, L.; Zhai, X.; Unterthiner, T.; Yung, J.; Keysers, D.; Uszkoreit, J.; Lucic, M.; et al. 2021 · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and J’egou, H. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
El-Nouby, A.; Neverova, N.; Laptev, I.; and Jégou, H. 2021 · 2021
Cited alongside, same era.
Masked autoencoders are scalable vision learners
He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; and Girshick, R. 2021 · 2021
Cited alongside, same era.
Rethinking Spatial Dimensions of Vision Transformers
Heo, B.; Yun, S.; Han, D.; Chun, S.; Choe, J.; and Oh, S. J. 2021 · 2021
Cited alongside, same era.
Generative Adversarial Transformers
Hudson, D. A.; and Zitnick, C. L. 2021 · 2021
Cited alongside, same era.
All Tokens Matter: Token Labeling for Training Better Vision Transformers
Jiang, Z.-H.; Hou, Q.; Yuan, L.; Zhou, D.; Shi, Y.; Jin, X.; Wang, A.; and Feng, J. 2021 · 2021
Cited alongside, same era.
HOTR: End-to-End Human-Object Interaction Detection with Transformers
Kim, B.; Lee, J.; Kang, J.; Kim, E.-S.; and Kim, H. J. 2021 · 2021
Cited alongside, same era.
SwinIR: Image Restoration Using Swin Transformer
Liang, J.; Cao, J.; Sun, G.; Zhang, K.; Van Gool, L.; and Timofte, R. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Wang, H.; Zhang, Z.; and Han, S. 2021 · 2021
Later among the works it cites.
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
Wang, W.; Xie, E.; Li, X.; Fan, D.-P.; Song, K.; Liang, D.; Lu, T.; Luo, P.; and Shao, L. 2021 · 2021
Later among the works it cites.
Evo-ViT: Slow-Fast Token Evolution for Dynamic Vision Transformer
Xu, Y.; Zhang, Z.; Zhang, M.; Sheng, K.; Li, K.; Dong, W.; Zhang, L.; Xu, C.; and Sun, X. 2021 · 2021
Later among the works it cites.
TransFER: Learning Relation-aware Facial Expression Representations with Transformers
Xue, F.; Wang, Q.; and Guo, G. 2021 · 2021
Later among the works it cites.
Point transformer
Zhao, H.; Jiang, L.; Jia, J.; Torr, P. H.; and Koltun, V. 2021 · 2021
Later among the works it cites.
SPViT: Enabling Faster Vision Transformers via Soft Token Pruning
Kong, Z.; Dong, P.; Ma, X.; Meng, X.; Niu, W.; Sun, M.; Ren, B.; Qin, M.; Tang, H.; and Wang, Y. 2022 · 2022
Closest in time.
EViT: Expediting Vision Transformers via Token Reorganizations
Liang, Y.; GE, C.; Tong, Z.; Song, Y.; Wang, J.; and Xie, P. 2022 · 2022
Closest in time.
Dynamic Spatial Sparsification for Efficient Vision Transformers and Convolutional Neural Networks
Rao, Y.; Liu, Z.; Zhao, W.; Zhou, J.; and Lu, J. 2022 · 2022
Closest in time.
Trockman, A.; and Kolter, J. Z. 2022 · 2022
Closest in time.
Vision Transformer with Deformable Attention
Xia, Z.; Pan, X.; Song, S.; Li, L. E.; and Huang, G. 2022 · 2022
Closest in time.
Width & Depth Pruning for Vision Transformers
Yu, F.; Huang, K.; Wang, M.; Cheng, Y.; Chu, W.; and Cui, L. 2022a · 2022
Closest in time.
PCT: Point cloud transformer
Guo, M.-H.; Cai, J.-X.; Liu, Z.-N.; Mu, T.-J.; Martin, R. R.; and Hu, S.-M. 2021 · 2096
Closest in time.