Fetching the paper…
Reading the bibliography…
Vision transformers (ViTs) have recently received explosive popularity, but the huge computational cost is still a severe issue.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I.; Peters, M. E.; and Cohan, A. 2020 · 2004
Earlier work this paper cites.
Quantifying attention flow in transformers
Abnar, S.; and Zuidema, W. 2020 · 2005
Earlier work this paper cites.
Synthesizer: Rethinking self-attention in transformer models
Tay, Y.; Bahri, D.; Metzler, D.; Juan, D.-C.; Zhao, Z.; and Zheng, C. 2020 · 2005
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S.; Li, B. Z.; Khabsa, M.; Fang, H.; and Ma, H. 2020 · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Deformable detr: Deformable transformers for end-to-end object detection
Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2020 · 2010
Earlier work this paper cites.
A survey on visual transformer
Han, K.; Wang, Y.; Chen, H.; Chen, X.; Guo, J.; Liu, Z.; Tang, Y.; Xiao, A.; Xu, C.; Xu, Y.; et al. 2020 · 2012
Earlier work this paper cites.
Realformer: Transformer likes residual attention
He, R.; Ravula, A.; Kanagal, B.; and Ainslie, J. 2020 · 2012
Earlier work this paper cites.
Similarity of neural network representations revisited
Kornblith, S.; Norouzi, M.; Lee, H.; and Hinton, G. 2019 · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Earlier work this paper cites.
Transformer interpretability beyond attention visualization
Chefer, H.; Gur, S.; and Wolf, L. 2021 · 2021
Cited alongside, same era.
Chasing Sparsity in Vision Transformers: An End-to-End Exploration
Chen, T.; Cheng, Y.; Gan, Z.; Yuan, L.; Zhang, L.; and Wang, Z. 2021 · 2021
Cited alongside, same era.
ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases
D’Ascoli, S.; Touvron, H.; Leavitt, M. L.; Morcos, A. S.; Biroli, G.; and Sagun, L. 2021 · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
Sstvos: Sparse spatiotemporal transformers for video object segmentation
Duke, B.; Ahmed, A.; Wolf, C.; Aarabi, P.; and Taylor, G. W. 2021 · 2021
Cited alongside, same era.
DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; and Hsieh, C.-J. 2021 · 2021
Closest in time.
Patch Slimming for Efficient Vision Transformers
Tang, Y.; Han, K.; Wang, Y.; Xu, C.; Guo, J.; Xu, C.; and Tao, D. 2021 · 2021
Closest in time.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I.; Houlsby, N.; Kolesnikov, A.; Beyer, L.; Zhai, X.; Unterthiner, T.; Yung, J.; Keysers, D.; Uszkoreit, J.; Lucic, M.; et al. 2021 · 2021
Closest in time.
Training data-efficient image transformers & distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021 · 2021
Closest in time.
Co-Scale Conv-Attentional Image Transformers
Xu, W.; Xu, Y.; Chang, T.; and Tu, Z. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
LeViT: a Vision Transformer in ConvNet’s Clothing for Faster Inference
Graham, B.; El-Nouby, A.; Touvron, H.; Stock, P.; Joulin, A.; Jégou, H.; and Douze, M. 2021 · 2021
Cited alongside, same era.
Rethinking Spatial Dimensions of Vision Transformers
Heo, B.; Yun, S.; Han, D.; Chun, S.; Choe, J.; and Oh, S. J. 2021 · 2021
Cited alongside, same era.
Learned Token Pruning for Transformers
Kim, S.; Shen, S.; Thorsley, D.; Gholami, A.; Hassoun, J.; and Keutzer, K. 2021 · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Cited alongside, same era.
Pan, B.; Jiang, Y.; Panda, R.; Wang, Z.; Feris, R.; and Oliva, A. 2021 · 2021
Cited alongside, same era.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wang, W.; Xie, E.; Li, X.; Fan, D.-P.; Song, K.; Liang, D.; Lu, T.; Luo, P.; and Shao, L. 2021a
Cited in the paper.
Evolving attention with residual convolutions
Wang, Y.; Yang, Y.; Bai, J.; Zhang, M.; Bai, J.; Yu, J.; Zhang, C.; Huang, G.; and Tong, Y. 2021b
Cited in the paper.
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L.; Chen, Y.; Wang, T.; Yu, W.; Shi, Y.; Jiang, Z.; Tay, F. E.; Feng, J.; and Yan, S. 2021 · 2021
Closest in time.
Vision transformer with progressive sampling
Yue, X.; Sun, S.; Kuang, Z.; Wei, M.; Torr, P. H.; Zhang, W.; and Lin, D. 2021 · 2021
Closest in time.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Zheng, S.; Lu, J.; Zhao, H.; Zhu, X.; Luo, Z.; Wang, Y.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P. H.; et al. 2021 · 2021
Closest in time.
Refiner: Refining Self-attention for Vision Transformers
Zhou, D.; Shi, Y.; Kang, B.; Yu, W.; Jiang, Z.; Li, Y.; Jin, X.; Hou, Q.; and Feng, J. 2021 · 2021
Closest in time.
Transformers in computational visual media: A survey
Xu, Y.; Wei, H.; Lin, M.; Deng, Y.; Sheng, K.; Zhang, M.; Tang, F.; Dong, W.; Huang, F.; and Xu, C. 2022 · 2022
Closest in time.