Fetching the paper…
Reading the bibliography…
Vision transformers have shown great success on numerous computer vision tasks.
Teubner, 1910
H. Minkowski, Geometrie der zahlen · 1910
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in IEEE Conf. Comput. Vis. Pattern Recog
2009
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” tech. rep., University of Toronto, 2009
2009
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Adv. in Neural Inf. Process. Syst
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Int. Conf. Learn. Represent
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conf. Comput. Vis. Pattern Recog
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. of the IEEE Conf. Comput. Vis. Pattern Recognit
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Adv. in Neural Inf. Process. Syst
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-CAM: Visual explanations from deep networks via gradient-based localization,” in Proc. of the IEEE Int. Conf. on Comput. Vis
2017
Earlier work this paper cites.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in CVPR
2017
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ADE20K dataset,” in IEEE Conf. Comput. Vis. Pattern Recog
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “SGDR: Stochastic gradient descent with warm restarts,” in Int. Conf. Learn. Represent
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proc. of the IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR)
2018
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proc. of the IEEE Conf. Comput. Vis. Pattern Recognit
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
H. Hu, Z. Zhang, Z. Xie, and S. Lin, “Local relation networks for image recognition,” in Proc. of the IEEE/CVF Int. Conf. on Comput. Vis
2019
Earlier work this paper cites.
P. Ramachandran, N. Parmar, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens, “Stand-alone self-attention in vision models,” Adv. in Neural Inf. Process. Syst
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
X. Cheng, Y. Zhong, Y. Dai, P. Ji, and H. Li, “Noise-aware unsupervised deep lidar-stereo fusion,” in CVPR
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
Y. Cao, J. Xu, S. Lin, F. Wei, and H. Hu, “Gcnet: Non-local networks meet squeeze-excitation networks and beyond,” in Proc. of the IEEE/CVF Int. Conf. on Comput. Vis. (ICCV) Workshops
2019
Cited alongside, same era.
A. Kirillov, R. Girshick, K. He, and P. Dollár, “Panoptic feature pyramid networks,” in Proc. of the IEEE/CVF Conf. Comput. Vis. Pattern Recognit
2019
Cited alongside, same era.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al
2020
Cited alongside, same era.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are RNNs: Fast autoregressive transformers with linear attention,” in Int. Conf. on Machine Learning
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Int. Conf. Comput. Vis
2021
Later among the works it cites.
A. Vaswani, P. Ramachandran, A. Srinivas, N. Parmar, B. Hechtman, and J. Shlens, “Scaling local self-attention for parameter efficient visual backbones,” in Proc. of the IEEE/CVF Conf. Comput. Vis. Pattern Recognit
2021
Later among the works it cites.
P. Zhang, X. Dai, J. Yang, B. Xiao, L. Yuan, L. Zhang, and J. Gao, “Multi-scale vision longformer: A new vision transformer for high-resolution image encoding,” 2021
2021
Later among the works it cites.
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable DETR: Deformable transformers for end-to-end object detection,” in Int. Conf. Learn. Represent
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
H. Zhao, J. Jia, and V. Koltun, “Exploring self-attention for image recognition,” in Proc. of the IEEE/CVF Conf. Comput. Vis. Pattern Recognit
2020
Cited alongside, same era.
2020
Cited alongside, same era.
N. Kitaev, Ł. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” in Int. Conf. Learn. Represent
2020
Cited alongside, same era.
G. Daras, N. Kitaev, A. Odena, and A. G. Dimakis, “SMYRF: Efficient attention using asymmetric clustering,” Adv. in Neural Inf. Process. Syst
2020
Cited alongside, same era.
2020
Cited alongside, same era.
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollár, “Designing network design spaces,” in IEEE Conf. Comput. Vis. Pattern Recog
2020
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in Int. Conf. Learn. Represent
2021
Cited alongside, same era.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Proc. of the IEEE/CVF Int. Conf. on Comput. Vis. (ICCV)
2021
Later among the works it cites.
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen, “Twins: Revisiting the design of spatial attention in vision transformers,” Adv. in Neural Inf. Process. Syst
2021
Later among the works it cites.
Z. Shen, M. Zhang, H. Zhao, S. Yi, and H. Li, “Efficient attention: Attention with linear complexities,” in Proc. of the IEEE/CVF Winter Conf. on Appl. of Comput. Vis
2021
Later among the works it cites.
A. Ali, H. Touvron, M. Caron, P. Bojanowski, M. Douze, A. Joulin, I. Laptev, N. Neverova, G. Synnaeve, J. Verbeek, et al
2021
Later among the works it cites.
J. Wang, Y. Zhong, Y. Dai, S. Birchfield, K. Zhang, N. Smolyanskiy, and H. Li, “Deep two-view structure-from-motion revisited,” in Proceedings of the IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)
2021
Later among the works it cites.
X. Chu, B. Zhang, Z. Tian, X. Wei, and H. Xia, “Do we really need explicit position encodings for vision transformers?,” arXiv preprint 2102.10882v1
2021
Later among the works it cites.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Int. Conf. Comput. Vis
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proc. of the IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)
2022
Closest in time.
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler, “Efficient transformers: A survey,” 2022
2022
Closest in time.
Q. Zhen, W. Sun, H. Deng, D. Li, Y. Wei, B. Lv, J. Yan, L. Kong, and Y. Zhong, “cosformer: Rethinking softmax in attention,” in Int. Conf. on Learn. Representations
2022
Closest in time.
Z. Xia, X. Pan, S. Song, L. E. Li, and G. Huang, “Vision transformer with deformable attention,” in Proc. of the IEEE/CVF Conf. Comput. Vis. Pattern Recognit
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “PVT v2: Improved baselines with pyramid vision transformer,” Computational Visual Media
2022
Closest in time.