Fetching the paper…
Reading the bibliography…
Recently, Vision Transformer (ViT) has achieved promising performance in image recognition and gradually serves as a powerful backbone in various vision tasks.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al., 2020 · 1901
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database, in: 2009 IEEE conference on computer vision and pattern recognition, Ieee. pp. 248–255
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L., 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al., 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al., 2020 · 2010
Earlier work this paper cites.
A survey on visual transformer
Han, K., Wang, Y., Chen, H., Chen, X., Guo, J., Liu, Z., Tang, Y., Xiao, A., Xu, C., Xu, Y., et al., 2020 · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., Hinton, G.E., 2012 · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context, in: Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, Springer. pp. 740–755
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L., 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., Dean, J., 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778
He, K., Zhang, X., Ren, S., Sun, J., 2016 · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., Poole, B., 2016 · 2016
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C.J., Mnih, A., Teh, Y.W., 2016 · 2016
Earlier work this paper cites.
A downsampled variant of imagenet as an alternative to the cifar datasets
Chrabaszcz, P., Loshchilov, I., Hutter, F., 2017 · 2017
Earlier work this paper cites.
Learning efficient convolutional networks through network slimming, in: Proceedings of the IEEE international conference on computer vision, pp. 2736–2744
Liu, Z., Li, J., Shen, Z., Huang, G., Yan, S., Zhang, C., 2017 · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era, in: Proceedings of the IEEE international conference on computer vision, pp. 843–852
Sun, C., Shrivastava, A., Singh, S., Gupta, A., 2017 · 2017
Earlier work this paper cites.
Attention is all you need, in: Advances in neural information processing systems, pp. 5998–6008
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I., 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2018 · 2018
Earlier work this paper cites.
Darts: Differentiable architecture search
Liu, H., Simonyan, K., Yang, Y., 2018 · 2018
Earlier work this paper cites.
Proxylessnas: Direct neural architecture search on target task and hardware, in: International Conference on Learning Representations
Cai, H., Zhu, L., Han, S., 2019 · 2019
Earlier work this paper cites.
Searching for mobilenetv3, in: Proceedings of the IEEE/CVF international conference on computer vision, pp. 1314–1324
Howard, A., Sandler, M., Chu, G., Chen, L.C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., et al., 2019 · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks, in: International Conference on Machine Learning, PMLR. pp. 6105–6114
Tan, M., Le, Q., 2019 · 2019
Earlier work this paper cites.
Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10734–10742
Wu, B., Dai, X., Zhang, P., Wang, Y., Sun, F., Wu, Y., Tian, Y., Vajda, P., Jia, Y., Keutzer, K., 2019 · 2019
Earlier work this paper cites.
End-to-end object detection with transformers, in: European Conference on Computer Vision, Springer. pp. 213–229
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S., 2020 · 2020
Earlier work this paper cites.
Calibration data-based cnn filter pruning for efficient layer fusion, in: 2020 IEEE 22nd International Conference on High Performance Computing and Communications; IEEE 18th International Conference on Smart City; IEEE 6th International Conference on Data Science and Systems (HPCC/SmartCity/DSS), IEEE. pp. 1300–1307
Chitty-Venkata, K.T., Somani, A.K., 2020 · 2020
Earlier work this paper cites.
Power-bert: Accelerating bert inference via progressive word-vector elimination, in: International Conference on Machine Learning, PMLR. pp. 3690–3699
Goyal, S., Choudhury, A.R., Raje, S., Chakaravarthy, V., Sabharwal, Y., Verma, A., 2020 · 2020
Cited alongside, same era.
Scop: Scientific control for reliable neural network pruning
Tang, Y., Wang, Y., Xu, Y., Tao, D., Xu, C., Xu, C., Xu, C., 2020 · 2020
Cited alongside, same era.
Twins: Revisiting the design of spatial attention in vision transformers
Chu, X., Tian, Z., Wang, Y., Zhang, B., Ren, H., Wei, X., Xia, H., Shen, C., 2021 · 2021
Cited alongside, same era.
Nasvit: Neural architecture search for efficient vision transformers with gradient conflict aware supernet training, in: International Conference on Learning Representations
Gong, C., Wang, D., Li, M., Chen, X., Yan, Z., Tian, Y., Chandra, V., et al., 2021 · 2021
Cited alongside, same era.
Transformer in transformer
Han, K., Xiao, A., Wu, E., Guo, J., Xu, C., Wang, Y., 2021 · 2021
Cited alongside, same era.
Learning disentangled representation implicitly via transformer for occluded person re-identification
Jia, M., Cheng, X., Lu, S., Zhang, J., 2022 · 2022
Closest in time.
Mpvit: Multi-path vision transformer for dense prediction, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7287–7296
Lee, Y., Kim, J., Willette, J., Hwang, S.J., 2022 · 2022
Closest in time.
Exploring plain vision transformer backbones for object detection, in: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part IX, Springer. pp. 280–296
Li, Y., Mao, H., Girshick, R., He, K., 2022 · 2022
Closest in time.
Not all patches are what you need: Expediting vision transformers via token reorganizations
Liang, Y., Ge, C., Tong, Z., Song, Y., Wang, J., Xie, P., 2022 · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Transreid: Transformer-based object re-identification, in: Proceedings of the IEEE/CVF international conference on computer vision, pp. 15013–15022
He, S., Luo, H., Wang, P., Wang, F., Li, H., Jiang, W., 2021 · 2021
Cited alongside, same era.
All tokens matter: Token labeling for training better vision transformers
Jiang, Z.H., Hou, Q., Yuan, L., Zhou, D., Shi, Y., Jin, X., Wang, A., Feng, J., 2021 · 2021
Cited alongside, same era.
Bossnas: Exploring hybrid cnn-transformers with block-wisely self-supervised neural architecture search, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 12281–12291
Li, C., Tang, T., Wang, G., Peng, J., Wang, B., Liang, X., Chang, X., 2021 · 2021
Cited alongside, same era.
Evit: Expediting vision transformers via token reorganizations, in: International Conference on Learning Representations
Liang, Y., Chongjian, G., Tong, Z., Song, Y., Wang, J., Xie, P., 2021 · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10012–10022
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B., 2021 · 2021
Cited alongside, same era.
Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer, in: International Conference on Learning Representations
Mehta, S., Rastegari, M., 2021 · 2021
Cited alongside, same era.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Rao, Y., Zhao, W., Liu, B., Lu, J., Zhou, J., Hsieh, C.J., 2021 · 2021
Cited alongside, same era.
Su, X., You, S., Xie, J., Zheng, M., Wang, F., Qian, C., Zhang, C., Wang, X., Xu, C., 2022 · 2022
Closest in time.
Patch slimming for efficient vision transformers, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12165–12174
Tang, Y., Han, K., Wang, Y., Xu, C., Guo, J., Xu, C., Tao, D., 2022 · 2022
Closest in time.
Tinyvit: Fast pretraining distillation for small vision transformers, in: European Conference on Computer Vision, Springer. pp. 68–85
Wu, K., Zhang, J., Peng, H., Liu, M., Xiao, B., Fu, J., Yuan, L., 2022 · 2022
Closest in time.
A-vit: Adaptive tokens for efficient vision transformer, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10809–10818
Yin, H., Vahdat, A., Alvarez, J.M., Mallya, A., Kautz, J., Molchanov, P., 2022 · 2022
Closest in time.
Minivit: Compressing vision transformers with weight multiplexing, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12145–12154
Zhang, J., Peng, H., Wu, K., Liu, M., Xiao, B., Fu, J., Yuan, L., 2022 · 2022
Closest in time.
Spatial-channel enhanced transformer for visible-infrared person re-identification
Zhao, J., Wang, H., Zhou, Y., Yao, R., Chen, S., El Saddik, A., 2022 · 2022
Closest in time.
Interaction transformer for human reaction generation
Chopin, B., Tang, H., Otberdout, N., Daoudi, M., Sebe, N., 2023 · 2023
Closest in time.
A practical survey on faster and lighter transformers
Fournier, Q., Caron, G.M., Aloise, D., 2023 · 2023
Closest in time.
Generalized image outpainting with u-transformer
Gao, P., Yang, X., Zhang, R., Goulermas, J.Y., Geng, Y., Yan, Y., Huang, K., 2023 · 2023
Closest in time.
Improving structural mri preprocessing with hybrid transformer gans
Grigas, O., Maskeliūnas, R., Damaševičius, R., 2023 · 2023
Closest in time.
Ultra-high resolution svbrdf recovery from a single image
Guo, J., Lai, S., Tu, Q., Tao, C., Zou, C., Guo, Y., 2023 · 2023
Closest in time.
Dilateformer: Multi-scale dilated transformer for visual recognition
Jiao, J., Tang, Y.M., Lin, K.Y., Gao, Y., Ma, J., Wang, Y., Zheng, W.S., 2023 · 2023
Closest in time.
Full stack optimization of transformer inference: a survey
Kim, S., Hooper, C., Wattanawong, T., Kang, M., Yan, R., Genc, H., Dinh, G., Huang, Q., Keutzer, K., Mahoney, M.W., et al., 2023 · 2023
Closest in time.
Token pooling in vision transformers for image classification, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 12–21
Marin, D., Chang, J.H.R., Ranjan, A., Prabhu, A., Rastegari, M., Tuzel, O., 2023 · 2023
Closest in time.
Pixel-level fusion approach with vision transformer for early detection of alzheimer’s disease
Odusami, M., Maskeliūnas, R., Damaševičius, R., 2023 · 2023
Closest in time.
Crimenet: Neural structured learning using vision transformer for violence detection
Rendón-Segador, F.J., Álvarez-García, J.A., Salazar-González, J.L., Tommasi, T., 2023 · 2023
Closest in time.
A multi-information fusion vit model and its application to the fault diagnosis of bearing with small data samples
Xu, Z., Tang, X., Wang, Z., 2023 · 2023
Closest in time.
Vit-llmr: Vision transformer-based lower limb motion recognition from fusion signals of mmg and imu
Zhang, H., Yang, K., Cao, G., Xia, C., 2023 · 2023
Closest in time.