Fetching the paper…
Reading the bibliography…
We propose global context vision transformer (GC ViT), a novel architecture that enhances parameter and compute utilization for computer vision.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., and Girshick, R · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G., Zhu, M., and Chen, B · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., and Torralba, A · 2017
Cited alongside, same era.
Squeeze-and-excitation networks
Hu, J., Shen, L., and Sun, G · 2018
Cited alongside, same era.
Unified perceptual parsing for scene understanding
Xiao, T., Liu, Y., Zhou, B., Jiang, Y., and Sun, J · 2018
Cited alongside, same era.
Mmdetection: Open mmlab detection toolbox and benchmark
Chen, K., Wang, J., Pang, J., Cao, Y., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., Xu, J., et al · 2019
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2019
Cited alongside, same era.
Pytorch image models
Wightman, R · 2019
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Later among the works it cites.
Do vision transformers see like convolutional neural networks?
Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C., and Dosovitskiy, A · 2021
Later among the works it cites.
Efficientnetv2: Smaller models and faster training
Tan, M. and Le, Q · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L · 2021
Later among the works it cites.
Cvt: Introducing convolutions to vision transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark
Contributors, M · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Designing network design spaces
Radosavovic, I., Kosaraju, R. P., Girshick, R., He, K., and Dollár, P · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Cited alongside, same era.
Xcit: Cross-covariance image transformers
Ali, A., Touvron, H., Caron, M., Bojanowski, P., Douze, M., Joulin, A., Laptev, I., Neverova, N., Synnaeve, G., Verbeek, J., et al · 2021
Cited alongside, same era.
Crossvit: Cross-attention multi-scale vision transformer for image classification, 2021
Chen, C.-F., Fan, Q., and Panda, R · 2021
Cited alongside, same era.
Wu, H., Xiao, B., Codella, N., Liu, M., Dai, X., Yuan, L., and Zhang, L · 2021
Later among the works it cites.
A-ViT: Adaptive tokens for efficient vision transformer
Yin, H., Vahdat, A., Alvarez, J., Mallya, A., Kautz, J., and Molchanov, P · 2021
Later among the works it cites.
Tokens-to-token ViT: Training vision transformers from scratch on imagenet
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z., Tay, F. E., Feng, J., and Yan, S · 2021
Later among the works it cites.
Cswin transformer: A general vision transformer backbone with cross-shaped windows
Dong, X., Bao, J., Chen, D., Zhang, W., Yu, N., Yuan, L., Chen, D., and Guo, B · 2022
Closest in time.
Edgevits: Competing light-weight cnns on mobile devices with vision transformers
Pan, J., Bulat, A., Tan, F., Zhu, X., Dudziak, L., Li, H., Tzimiropoulos, G., and Martinez, B · 2022
Closest in time.
Maxvit: Multi-axis vision transformer
Tu, Z., Talebi, H., Zhang, H., Yang, F., Milanfar, P., Bovik, A., and Li, Y · 2022
Closest in time.
Pvt v2: Improved baselines with pyramid vision transformer
Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L · 2022
Closest in time.
Metaformer is actually what you need for vision
Yu, W., Luo, M., Zhou, P., Si, C., Zhou, Y., Wang, X., Feng, J., and Yan, S · 2022
Closest in time.
Dino: Detr with improved denoising anchor boxes for end-to-end object detection
Zhang, H., Li, F., Liu, S., Zhang, L., Su, H., Zhu, J., Ni, L. M., and Shum, H.-Y · 2022
Closest in time.