Fetching the paper…
Reading the bibliography…
Vision Transformers have witnessed prevailing success in a series of vision tasks.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Deformable detr: Deformable transformers for end-to-end object detection
Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2020 · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012 · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G.; Sun, Y.; Liu, Z.; Sedra, D.; and Weinberger, K. Q. 2016 · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Zoph, B.; and Le, Q. V. 2016 · 2016
Earlier work this paper cites.
Mask r-cnn
He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. 2017 · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; and Adam, H. 2017 · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; and Dollár, P. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; and Batra, D. 2017 · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Xie, S.; Girshick, R.; Dollár, P.; Tu, Z.; and He, K. 2017 · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Zhou, B.; Zhao, H.; Puig, X.; Fidler, S.; Barriuso, A.; and Torralba, A. 2017 · 2017
Earlier work this paper cites.
Shufflenet v2: Practical guidelines for efficient cnn architecture design
Ma, N.; Zhang, X.; Zheng, H.-T.; and Sun, J. 2018 · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; and Chen, L.-C. 2018 · 2018
Cited alongside, same era.
Unified perceptual parsing for scene understanding
Xiao, T.; Liu, Y.; Zhou, B.; Jiang, Y.; and Sun, J. 2018 · 2018
Cited alongside, same era.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
Zhang, X.; Zhou, X.; Lin, M.; and Sun, J. 2018 · 2018
Cited alongside, same era.
Searching for mobilenetv3
Howard, A.; Sandler, M.; Chu, G.; Chen, L.-C.; Chen, B.; Tan, M.; Wang, W.; Zhu, Y.; Pang, R.; Vasudevan, V.; et al. 2019 · 2019
Cited alongside, same era.
Panoptic feature pyramid networks
Kirillov, A.; Girshick, R.; He, K.; and Dollár, P. 2019 · 2019
Transformer in transformer
Han, K.; Xiao, A.; Wu, E.; Guo, J.; Xu, C.; and Wang, Y. 2021 · 2021
Later among the works it cites.
Shuffle transformer: Rethinking spatial shuffle for vision transformer
Huang, Z.; Ben, Y.; Luo, G.; Cheng, P.; Yu, G.; and Fu, B. 2021 · 2021
Later among the works it cites.
Bossnas: Exploring hybrid cnn-transformers with block-wisely self-supervised neural architecture search
Li, C.; Tang, T.; Wang, G.; Peng, J.; Wang, B.; Liang, X.; and Chang, X. 2021 · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Later among the works it cites.
Mobilevit: light-weight, general-purpose, and mobile-friendly vision transformer
Mehta, S.; and Rastegari, M. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The evolved transformer
So, D.; Le, Q.; and Liang, C. 2019 · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M.; and Le, Q. 2019 · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Cited alongside, same era.
Ghostnet: More features from cheap operations
Han, K.; Wang, Y.; Tian, Q.; Guo, J.; Xu, C.; and Xu, C. 2020 · 2020
Cited alongside, same era.
Up-detr: Unsupervised pre-training for object detection with transformers
Dai, Z.; Cai, B.; Lin, Y.; and Chen, J. 2021 · 2021
Cited alongside, same era.
Cswin transformer: A general vision transformer backbone with cross-shaped windows
Dong, X.; Bao, J.; Chen, D.; Zhang, W.; Yu, N.; Yuan, L.; Chen, D.; and Guo, B. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021 · 2021
Later among the works it cites.
Cvt: Introducing convolutions to vision transformers
Wu, H.; Xiao, B.; Codella, N.; Liu, M.; Dai, X.; Yuan, L.; and Zhang, L. 2021 · 2021
Later among the works it cites.
SegFormer: Simple and efficient design for semantic segmentation with transformers
Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J. M.; and Luo, P. 2021 · 2021
Later among the works it cites.
Co-scale conv-attentional image transformers
Xu, W.; Xu, Y.; Chang, T.; and Tu, Z. 2021 · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L.; Chen, Y.; Wang, T.; Yu, W.; Shi, Y.; Jiang, Z.-H.; Tay, F. E.; Feng, J.; and Yan, S. 2021 · 2021
Later among the works it cites.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Zheng, S.; Lu, J.; Zhao, H.; Zhu, X.; Luo, Z.; Wang, Y.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P. H.; et al. 2021 · 2021
Later among the works it cites.
Neural architecture search with a lightweight transformer for text-to-image synthesis
Li, W.; Wen, S.; Shi, K.; Yang, Y.; and Huang, T. 2022 · 2022
Closest in time.
Liu, Z.; Mao, H.; Wu, C.-Y.; Feichtenhofer, C.; Darrell, T.; and Xie, S. 2022 · 2022
Closest in time.
Metaformer is actually what you need for vision
Yu, W.; Luo, M.; Zhou, P.; Si, C.; Zhou, Y.; Wang, X.; Feng, J.; and Yan, S. 2022 · 2022
Closest in time.