Fetching the paper…
Reading the bibliography…
While features of different scales are perceptually important to visual inputs, existing vision transformers do not yet take advantage of them explicitly.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in International Conference on Machine Learning, ICML , J. Fürnkranz and T. Joachims, Eds., 2010, pp. 807–814
2010
Earlier work this paper cites.
T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: common objects in context,” in European Conference on Computer Vision, ECCV , vol. 8693, 2014, pp. 740–755
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations, ICLR , 2015
2015
Earlier work this paper cites.
L. J. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” CoRR , vol. abs/1607.06450, 2016
2016
Earlier work this paper cites.
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger, “Deep networks with stochastic depth,” in European Conference on Computer Vision, ECCV , B. Leibe, J. Matas, N. Sebe, and M. Welling, Eds., vol. 9908, 2016, pp. 646–661
2016
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. B. Girshick, “Mask R-CNN,” in International Conference on Computer Vision, ICCV , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ADE20K dataset,” in Conference on Computer Vision and Pattern Recognition, CVPR , 2017, pp. 5122–5130
2017
Earlier work this paper cites.
P. Shaw, J. Uszkoreit, and A. Vaswani, “Self-attention with relative position representations,” in Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL , 2018, pp. 464–468
2018
Earlier work this paper cites.
H. Zhang, M. Cissé, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in International Conference on Learning Representations, ICLR , 2018
2018
Earlier work this paper cites.
T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun, “Unified perceptual parsing for scene understanding,” in European Conference on Computer Vision, ECCV , vol. 11209, 2018, pp. 432–448
2018
Earlier work this paper cites.
S. Yun, D. Han, S. Chun, S. J. Oh, Y. Yoo, and J. Choe, “Cutmix: Regularization strategy to train strong classifiers with localizable features,” in International Conference on Computer Vision, ICCV , 2019, pp. 6022–6031
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Feichtenhofer, “X3d: Expanding architectures for efficient video recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 203–213
2020
Earlier work this paper cites.
E. D. Cubuk, B. Zoph, J. Shlens, and Q. Le, “Randaugment: Practical automated data augmentation with a reduced search space,” in Neural Information Processing Systems, NeurIPS , 2020
2020
Earlier work this paper cites.
Z. Zhong, L. Zheng, G. Kang, S. Li, and Y. Yang, “Random erasing data augmentation,” in Association for the Advancement of Artificial Intelligence, AAAI , 2020, pp. 13 001–13 008
2020
Earlier work this paper cites.
T. Lin, P. Goyal, R. B. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” Transactions on Pattern Analysis and Machine Intelligence, PAMI , vol. 42, no. 2, pp. 318–327, 2020
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Computer Vision European Conference, ECCV , vol. 12346, 2020, pp. 213–229
2020
Earlier work this paper cites.
M. Contributors, “MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark,” 2020
2020
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations, ICLR , 2021
2021
Earlier work this paper cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning, ICML , vol. 139, 2021, pp. 10 347–10 357
2021
Earlier work this paper cites.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Srinivas, T. Lin, N. Parmar, J. Shlens, P. Abbeel, and A. Vaswani, “Bottleneck transformers for visual recognition,” in Conference on Computer Vision and Pattern Recognition, CVPR , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Jiang, Q. Hou, L. Yuan, D. Zhou, Y. Shi, X. Jin, A. Wang, and J. Feng, “All tokens matter: Token labeling for training better vision transformers,” in Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual , M. Ranzato, A. Beygelzimer, Y. N. Dauphin, P. Liang, and J. W. Vaughan, Eds., 2021, pp. 18 590–18 602
2021
Cited alongside, same era.
Y. Li, C. Wu, H. Fan, K. Mangalam, B. Xiong, J. Malik, and C. Feichtenhofer, “Mvitv2: Improved multiscale vision transformers for classification and detection,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 4794–4804, 2021
2021
Cited alongside, same era.
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021 . IEEE, 2021, pp. 9630–9640
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
M. Chen, H. Peng, J. Fu, and H. Ling, “Autoformer: Searching transformers for visual recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2021, pp. 12 270–12 280
2021
Cited alongside, same era.
M. Chen, K. Wu, B. Ni, H. Peng, B. Liu, J. Fu, H. Chao, and H. Ling, “Searching the search space of vision transformer,” Advances in Neural Information Processing Systems , vol. 34, pp. 8714–8726, 2021
2021
Cited alongside, same era.
2021
Later among the works it cites.
Z. Dai, H. Liu, Q. V. Le, and M. Tan, “Coatnet: Marrying convolution and attention for all data sizes,” in Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual , M. Ranzato, A. Beygelzimer, Y. N. Dauphin, P. Liang, and J. W. Vaughan, Eds., 2021, pp. 3965–3977
2021
Later among the works it cites.
W. Wang, L. Yao, L. Chen, B. Lin, D. Cai, X. He, and W. Liu, “Crossformer: A versatile vision transformer hinging on cross-scale attention,” in International Conference on Learning Representations, ICLR , 2022
2022
Later among the works it cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. B. Girshick, “Masked autoencoders are scalable vision learners,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022 . IEEE, 2022, pp. 15 979–15 988
2022
Later among the works it cites.
A. Baevski, W. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “data2vec: A general framework for self-supervised learning in speech, vision and language,” in International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA , ser. Proceedings of Machine Learning Research, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvári, G. Niu, and S. Sabato, Eds., vol. 162. PMLR, 2022, pp. 1298–1312
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Dong, J. Bao, D. Chen, W. Zhang, N. Yu, L. Yuan, D. Chen, and B. Guo, “Cswin transformer: A general vision transformer backbone with cross-shaped windows,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 12 124–12 134
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Lee, J. Kim, J. Willette, and S. J. Hwang, “Mpvit: Multi-path vision transformer for dense prediction,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022 . IEEE, 2022, pp. 7277–7286
2022
Later among the works it cites.
R. Yang, H. Ma, J. Wu, Y. Tang, X. Xiao, M. Zheng, and X. Li, “Scalablevit: Rethinking the context-oriented generalization of vision transformer,” in Computer Vision - ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part XXIV , ser. Lecture Notes in Computer Science, S. Avidan, G. J. Brostow, M. Cissé, G. M. Farinella, and T. Hassner, Eds., vol. 13684. Springer, 2022, pp. 480–496
2022
Later among the works it cites.
M. Ding, B. Xiao, N. Codella, P. Luo, J. Wang, and L. Yuan, “Davit: Dual attention vision transformers,” in Computer Vision - ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part XXIV , ser. Lecture Notes in Computer Science, S. Avidan, G. J. Brostow, M. Cissé, G. M. Farinella, and T. Hassner, Eds., vol. 13684. Springer, 2022, pp. 74–92
2022
Later among the works it cites.
A. Kirillov, R. B. Girshick, K. He, and P. Dollár, “Panoptic feature pyramid networks,” in Conference on Computer Vision and Pattern Recognition, CVPR , 2019, pp. 6399–6408
2022
Later among the works it cites.
L. Yuan, Q. Hou, Z. Jiang, J. Feng, and S. Yan, “VOLO: vision outlooker for visual recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 5, pp. 6575–6586, 2023
2023
Closest in time.
2023
Closest in time.