Fetching the paper…
Reading the bibliography…
Vision Transformer (ViT) has emerged as a competitive alternative to convolutional neural networks for various computer vision applications.
2006
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
B. Cui, Y. Li, M. Chen, and Z. Zhang, “Fine-tune bert with sparse self-attention mechanism,” in Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) , 2019, pp. 3548–3553
2019
Earlier work this paper cites.
R. Wightman, “Pytorch image models,” https://github.com/rwightman/pytorch-image-models , 2019
2019
Earlier work this paper cites.
L. Ye, M. Rochan, Z. Liu, and Y. Wang, “Cross-modal self-attention network for referring image segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 10 502–10 511
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in European conference on computer vision . Springer, 2020, pp. 213–229
2020
Earlier work this paper cites.
Google LLC., “Pixel 3,” https://g.co/kgs/pVRc1Y , accessed 2020-09-01
2020
Earlier work this paper cites.
T. J. Ham, S. J. Jung, S. Kim, Y. H. Oh, Y. Park, Y. Song, J.-H. Park, S. Lee, K. Park, J. W. Lee et al. , “Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2020, pp. 328–341
2020
Earlier work this paper cites.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are rnns: Fast autoregressive transformers with linear attention,” in International Conference on Machine Learning . PMLR, 2020, pp. 5156–5165
2020
Earlier work this paper cites.
N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=rkgNKkHtvB
2020
Cited alongside, same era.
NVIDIA Inc., “NVIDIA Jetson TX2,” https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-tx2/ , accessed 2020-09-01
2020
Cited alongside, same era.
NVIDIA LLC., “GeForce RTX 2080 TI Graphics Card — NVIDIA,” 2021, https://www.nvidia.com/en-me/geforce/graphics-cards/rtx-2080-ti/ , accessed 2020-09-01
2020
Cited alongside, same era.
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang et al. , “Big bird: Transformers for longer sequences,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 283–17 297, 2020
2020
Cited alongside, same era.
L. Lu, Y. Jin, H. Bi, Z. Luo, P. Li, T. Wang, and Y. Liang, “Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,” in MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , 2021, pp. 977–991
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Shen, M. Zhang, H. Zhao, S. Yi, and H. Li, “Efficient attention: Attention with linear complexities,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision , 2021, pp. 3531–3539
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jegou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning , vol. 139, July 2021, pp. 10 347–10 357
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
B. Chen, T. Dao, E. Winsor, Z. Song, A. Rudra, and C. Ré, “Scatterbrain: Unifying sparse and low-rank attention,” in Advances in Neural Information Processing Systems , A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., 2021. [Online]. Available: https://openreview.net/forum?id=SehIKudiIo1
2021
Cited alongside, same era.
C.-F. R. Chen, Q. Fan, and R. Panda, “Crossvit: Cross-attention multi-scale vision transformer for image classification,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 357–366
2021
Cited alongside, same era.
H. Chen, Z. Luo, J. Zhang, L. Zhou, X. Bai, Z. Hu, C.-L. Tai, and L. Quan, “Learning to match features with seeded graph matching network,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2021, pp. 6301–6310
2021
Cited alongside, same era.
T. Chen, Y. Cheng, Z. Gan, L. Yuan, L. Zhang, and Z. Wang, “Chasing sparsity in vision transformers: An end-to-end exploration,” Advances in Neural Information Processing Systems , vol. 34, pp. 19 974–19 988, 2021
2021
Cited alongside, same era.
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller, “Rethinking attention with performers,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=Ua6zuk0WRH
2021
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
B. Graham, A. El-Nouby, H. Touvron, P. Stock, A. Joulin, H. Jegou, and M. Douze, “Levit: A vision transformer in convnet’s clothing for faster inference,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2021, pp. 12 259–12 269
2021
Cited alongside, same era.
Later among the works it cites.
H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2021, pp. 97–110
2021
Later among the works it cites.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 568–578
2021
Later among the works it cites.
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang, “Cvt: Introducing convolutions to vision transformers,” 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Zhu, K. Han, Y. Tang, and Y. Wang, “Visual transformer pruning,” arXiv e-prints , pp. arXiv–2104, 2021
2021
Later among the works it cites.
2022
Closest in time.
H. Germain, V. Lepetit, and G. Bourmaud, “Visual correspondence hallucination,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=jaLDP8Hp_gc
2022
Closest in time.
C. Gong, D. Wang, M. Li, X. Chen, Z. Yan, Y. Tian, qiang liu, and V. Chandra, “NASVit: Neural architecture search for efficient vision transformers with gradient conflict aware supernet training,” in International Conference on Learning Representations , 2022
2022
Closest in time.
Z. Qu, L. Liu, F. Tu, Z. Chen, Y. Ding, and Y. Xie, “Dota: detect and omit weak attentions for scalable transformer acceleration,” in Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , 2022, pp. 14–26
2022
Closest in time.
G. Shen, J. Zhao, Q. Chen, J. Leng, C. Li, and M. Guo, “Salo: an efficient spatial accelerator enabling hybrid sparse attention mechanisms for long sequences,” Proceedings of the 59th ACM/IEEE Design Automation Conference , 2022
2022
Closest in time.
S. Tang, J. Zhang, S. Zhu, and P. Tan, “Quadtree attention for vision transformers,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=fR-EnKWL_Zb
2022
Closest in time.