Fetching the paper…
Reading the bibliography…
Attention layers are the core component of transformers, the current state-of-the-art neural network architecture.
The behavior of eigenvalues and singular values under perturbations of restricted rank
Robert C Thompson · 1976
Earlier work this paper cites.
A Limit Theorem for the Norm of Random Matrices
Stuart Geman · 1980
Earlier work this paper cites.
Circular law theorem for random markov matrices
Charles Bordenave, Pietro Caputo, and Djalil Chafaï · 2011
Earlier work this paper cites.
Topics in random matrix theory
Terence Tao · 2012
Earlier work this paper cites.
Products of rectangular random matrices: Singular values and progressive scattering
Gernot Akemann, Jesper R. Ipsen, and Mario Kieburg · 2013
Earlier work this paper cites.
Spectral density of generalized wishart matrices and free multiplicative convolution
Wojciech Młotkowski, Maciej A Nowak, Karol A Penson, and Karol Życzkowski · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate, 2016
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2016
Earlier work this paper cites.
Free probability and random matrices
James A Mingo and Roland Speicher · 2017
Earlier work this paper cites.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, and Aidan N Gomez · 2017
Earlier work this paper cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Boris Hanin · 2018
Cited alongside, same era.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2018
Cited alongside, same era.
Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz, and Jeffrey Pennington · 2018
Cited alongside, same era.
Infinite attention: Nngp and ntk for deep attention networks
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, and Roman Novak · 2020
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas · 2021
Cited alongside, same era.
Revisiting over-smoothing in bert from the perspective of graph
Han Shi, Jiahui Gao, Hang Xu, Xiaodan Liang, Zhenguo Li, Lingpeng Kong, Stephen Lee, and James T Kwok · 2022
Later among the works it cites.
Peihao Wang, Wenqing Zheng, Tianlong Chen, and Zhangyang Wang · 2022
Later among the works it cites.
Centered self-attention layers
Ameen Ali, Tomer Galanti, and Lior Wolf · 2023
Later among the works it cites.
Heejong Bong and Arun Kumar Kuchibhotla · 2023
Later among the works it cites.
Replacing softmax with relu in vision transformers, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bobby He, James Martens, Guodong Zhang, Aleksandar Botev, Andrew Brock, Samuel L Smith, and Yee Whye Teh · 2022
Cited alongside, same era.
Moving beyond sub-gaussianity in high-dimensional statistics: Applications in covariance estimation and linear regression
Arun Kumar Kuchibhotla and Abhishek Chakrabortty · 2022
Cited alongside, same era.
Activation function design for deep networks: linearity and effective initialisation
Michael Murray, Vinayak Abrol, and Jared Tanner · 2022
Cited alongside, same era.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto, Sidak Pal Singh, and Aurelien Lucchi · 2022
Cited alongside, same era.
Mitchell Wortsman, Jaehoon Lee, Justin Gilmer, and Simon Kornblith · 2023
Later among the works it cites.
Setting the record straight on transformer oversmoothing
Gbetondji JS Dovonon, Michael M Bronstein, and Matt J Kusner · 2024
Closest in time.
Thiziri Nait Saada and Alireza Naderi · 2024
Closest in time.
Theory, analysis, and best practices for sigmoid self-attention
Jason Ramapuram, Federico Danieli, Eeshan Dhekane, Floris Weers, Dan Busbridge, Pierre Ablin, Tatiana Likhomanenko, Jagrit Digani, Zijin Gu, Amitis Shidani, et al · 2024
Closest in time.
Tianzhu Ye, Li Dong, Yuqing Xia, Yutao Sun, Yi Zhu, Gao Huang, and Furu Wei · 2024
Closest in time.