Fetching the paper…
Reading the bibliography…
Vision Transformers (ViTs) have emerged as powerful models in the field of computer vision, delivering superior performance across various vision tasks.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Clustering by fast search and find of density peaks
Rodriguez, A. and Laio, A · 2014
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Kenton, J. D. M.-W. C. and Toutanova, L. K · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Earlier work this paper cites.
Designing network design spaces
Radosavovic, I., Kosaraju, R. P., Girshick, R., He, K., and Dollár, P · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Earlier work this paper cites.
CrossViT: Cross-attention multi-scale vision transformer for image classification
Chen, C.-F. R., Fan, Q., and Panda, R · 2021
Earlier work this paper cites.
UP-DETR: Unsupervised pre-training for object detection with transformers
Dai, Z., Cai, B., Lin, Y., and Chen, J · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Earlier work this paper cites.
Transformer in transformer
Han, K., Xiao, A., Wu, E., Guo, J., Xu, C., and Wang, Y · 2021
Earlier work this paper cites.
Rethinking spatial dimensions of vision transformers
Heo, B., Yun, S., Han, D., Chun, S., Choe, J., and Oh, S. J · 2021
Cited alongside, same era.
All tokens matter: Token labeling for training better vision transformers
Jiang, Z.-H., Hou, Q., Yuan, L., Zhou, D., Shi, Y., Jin, X., Wang, A., and Feng, J · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Cited alongside, same era.
Token pooling in vision transformers
Marin, D., Chang, J.-H. R., Ranjan, A., Prabhu, A., Rastegari, M., and Tuzel, O · 2021
Cited alongside, same era.
IA-RED 2
Pan, B., Panda, R., Jiang, Y., Wang, Z., Feris, R., and Oliva, A · 2021
Cited alongside, same era.
DynamicViT: Efficient vision transformers with dynamic token sparsification
Vision transformer with progressive sampling
Yue, X., Sun, S., Kuang, Z., Wei, M., Torr, P. H., Zhang, W., and Lin, D · 2021
Later among the works it cites.
Deformable DETR: Deformable transformers for end-to-end object detection
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J · 2021
Later among the works it cites.
Adaptive token sampling for efficient vision transformers
Fayyaz, M., Koohpayegani, S. A., Jafari, F. R., Sengupta, S., Joze, H. R. V., Sommerlade, E., Pirsiavash, H., and Gall, J · 2022
Later among the works it cites.
Exploring plain vision transformer backbones for object detection
Li, Y., Mao, H., Girshick, R., and He, K · 2022
Later among the works it cites.
Not all patches are what you need: Expediting vision transformers via token reorganizations
Liang, Y., Ge, C., Tong, Z., Song, Y., Wang, J., and Xie, P · 2022
Later among the works it cites.
Patch slimming for efficient vision transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rao, Y., Zhao, W., Liu, B., Lu, J., Zhou, J., and Hsieh, C.-J · 2021
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Cited alongside, same era.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L · 2021
Cited alongside, same era.
Cvt: Introducing convolutions to vision transformers
Wu, H., Xiao, B., Codella, N., Liu, M., Dai, X., Yuan, L., and Zhang, L · 2021
Cited alongside, same era.
Co-scale conv-attentional image transformers
Xu, W., Xu, Y., Chang, T., and Tu, Z · 2021
Cited alongside, same era.
Tokens-to-Token ViT: Training vision transformers from scratch on imagenet
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z.-H., Tay, F. E., Feng, J., and Yan, S · 2021
Cited alongside, same era.
Tang, Y., Han, K., Wang, Y., Xu, C., Guo, J., Xu, C., and Tao, D · 2022
Later among the works it cites.
Evo-ViT: Slow-fast token evolution for dynamic vision transformer
Xu, Y., Zhang, Z., Zhang, M., Sheng, K., Li, K., Dong, W., Zhang, L., Xu, C., and Sun, X · 2022
Later among the works it cites.
A-ViT: Adaptive tokens for efficient vision transformer
Yin, H., Vahdat, A., Alvarez, J. M., Mallya, A., Kautz, J., and Molchanov, P · 2022
Later among the works it cites.
Token merging: Your vit but faster
Bolya, D., Fu, C.-Y., Dai, X., Zhang, P., Feichtenhofer, C., and Hoffman, J · 2023
Closest in time.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al · 2023
Closest in time.
Joint token pruning and squeezing towards more aggressive compression of vision transformers
Wei, S., Ye, T., Zhang, S., Tang, Y., and Liang, J · 2023
Closest in time.