Fetching the paper…
Reading the bibliography…
Training strategies for modern deep neural networks (NNs) tend to induce a heavy-tailed (HT) empirical spectral density (ESD) in the layer weights.
A simple general approach to inference about the tail of a distribution
Bruce M Hill · 1975
Earlier work this paper cites.
Estimation of the mean of a multivariate normal distribution
Charles M Stein · 1981
Earlier work this paper cites.
An introduction to probability theory
P.A.P. Moran · 1984
Earlier work this paper cites.
0. patashnik, concrete mathematics
Ronald L Graham and Donald E Knuth · 1989
Earlier work this paper cites.
On kernel-target alignment
Nello Cristianini, John Shawe-Taylor, Andre Elisseeff, and Jaz Kandola · 2001
Earlier work this paper cites.
Relations between communication complexity, linear arrangements, and computational complexity
Jürgen Forster, Matthias Krause, Satyanarayana V Lokam, Rustam Mubarakzjanov, Niels Schmitt, and Hans Ulrich Simon · 2001
Earlier work this paper cites.
Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices
Jinho Baik, Gérard Ben Arous, and Sandrine Péché · 2005
Earlier work this paper cites.
The elements of statistical learning: data mining, inference and prediction
Trevor Hastie, Robert Tibshirani, Jerome Friedman, and James Franklin · 2005
Earlier work this paper cites.
The spectrum of heavy tailed random matrices
Gérard Ben Arous and Alice Guionnet · 2008
Earlier work this paper cites.
Power-law distributions in empirical data
Aaron Clauset, Cosma Rohilla Shalizi, and Mark EJ Newman · 2009
Earlier work this paper cites.
An introduction to random matrices
Greg W Anderson, Alice Guionnet, and Ofer Zeitouni · 2010
Earlier work this paper cites.
Reconstruction of a low-rank matrix in the presence of gaussian noise
Andrey A Shabalin and Andrew B Nobel · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Mathematical analysis II , volume 220
Vladimir Antonovich Zorich and Octavio Paniagua · 2016
Earlier work this paper cites.
Optimal shrinkage of singular values
Matan Gavish and David L Donoho · 2017
Earlier work this paper cites.
Free probability and random matrices , volume 35
James A Mingo and Roland Speicher · 2017
Earlier work this paper cites.
signsgd: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Earlier work this paper cites.
Detection limits in the high-dimensional spiked rectangular model
Ahmed El Alaoui and Michael I Jordan · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
A random matrix approach to neural networks
Cosme Louart, Zhenyu Liao, and Romain Couillet · 2018
Earlier work this paper cites.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Jiri Hron, Mark Rowland, Richard E Turner, and Zoubin Ghahramani · 2018
Earlier work this paper cites.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares
Rong Ge, Sham M Kakade, Rahul Kidambi, and Praneeth Netrapalli · 2019
Cited alongside, same era.
First exit time analysis of stochastic gradient descent under heavy-tailed gradient noise
Thanh Huy Nguyen, Umut Simsekli, Mert Gurbuzbalaban, and Gaël Richard · 2019
Cited alongside, same era.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
Cited alongside, same era.
Spectra of the conjugate kernel and neural tangent kernel for linear-width neural networks
Zhou Fan and Zhichao Wang · 2020
Cited alongside, same era.
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2020
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2022
Later among the works it cites.
Optimal denoising of rotationally invariant rectangular matrices
Emanuele Troiani, Vittorio Erba, Florent Krzakala, Antoine Maillard, and Lenka Zdeborová · 2022
Later among the works it cites.
High-Dimensional Data Analysis with Low-Dimensional Models: Principles, Computation, and Applications
John Wright and Yi Ma · 2022
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz · 2023
Later among the works it cites.
Learning in the presence of low-dimensional structure: A spiked random matrix perspective
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, and Denny Wu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Heavy-tailed universality predicts trends in test accuracies for very large pre-trained deep neural networks
Charles H Martin and Michael W Mahoney · 2020
Cited alongside, same era.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2021
Cited alongside, same era.
Heavy tails in sgd and compressibility of overparametrized neural networks
Melih Barsbey, Milad Sefidgaran, Murat A Erdogdu, Gael Richard, and Umut Simsekli · 2021
Cited alongside, same era.
Spiked separable covariance matrices and principal components
Xiucai Ding and Fan Yang · 2021
Cited alongside, same era.
The heavy-tail phenomenon in sgd
Mert Gurbuzbalaban, Umut Simsekli, and Lingjiong Zhu · 2021
Cited alongside, same era.
Multiplicative noise and heavy tails in stochastic optimization
Liam Hodgkinson and Michael Mahoney · 2021
Cited alongside, same era.
Random features for kernel approximation: A survey on algorithms, theory, and beyond
Fanghui Liu, Xiaolin Huang, Yudong Chen, and Johan AK Suykens · 2021
Cited alongside, same era.
Smoothing the landscape boosts the signal for sgd: Optimal sample complexity for learning single index models
Alex Damian, Eshaan Nichani, Rong Ge, and Jason D Lee · 2023
Later among the works it cites.
How two-layer neural networks learn, one (giant) step at a time
Yatin Dandi, Florent Krzakala, Bruno Loureiro, Luca Pesce, and Ludovic Stephan · 2023
Later among the works it cites.
Normalization techniques in training dnns: Methodology, analysis and application
Lei Huang, Jie Qin, Yi Zhou, Fan Zhu, Li Liu, and Ling Shao · 2023
Later among the works it cites.
Singular vectors of sums of rectangular random matrices and optimal estimation of high-rank signals: The extensive spike model
Itamar D. Landau, Gabriel C. Mel, and Surya Ganguli · 2023
Later among the works it cites.
A theory of non-linear feature learning with one gradient step in two-layer neural networks
Behrad Moniri, Donghwan Lee, Hamed Hassani, and Edgar Dobriban · 2023
Later among the works it cites.
Provable guarantees for nonlinear feature learning in three-layer neural networks
Eshaan Nichani, Alex Damian, and Jason D Lee · 2023
Later among the works it cites.
Algorithmic stability of heavy-tailed sgd with general loss functions
Anant Raj, Lingjiong Zhu, Mert Gurbuzbalaban, and Umut Simsekli · 2023
Later among the works it cites.
Spectral evolution and invariance in linear-width neural networks
Zhichao Wang, Andrew William Engel, Anand Sarwate, Ioana Dumitriu, and Tony Chiang · 2023
Later among the works it cites.
Test accuracy vs. generalization gap: Model selection in nlp without accessing training or testing data
Yaoqing Yang, Ryan Theisen, Liam Hodgkinson, Joseph E Gonzalez, Kannan Ramchandran, Charles H Martin, and Michael W Mahoney · 2023
Later among the works it cites.
Temperature balancing, layer-wise weight analysis, and neural network training
Yefan Zhou, Tianyu Pang, Keqin Liu, Charles H. Martin, Michael W. Mahoney, and Yaoqing Yang · 2023
Later among the works it cites.
On single-index models beyond gaussian data
Aaron Zweig, Loucas Pillaud-Vivien, and Joan Bruna · 2023
Later among the works it cites.
Asymptotics of feature learning in two-layer networks after one gradient-step
Hugo Cui, Luca Pesce, Yatin Dandi, Florent Krzakala, Yue M Lu, Lenka Zdeborová, and Bruno Loureiro · 2024
Closest in time.
The computational complexity of learning gaussian single-index models
Alex Damian, Loucas Pillaud-Vivien, Jason D Lee, and Joan Bruna · 2024
Closest in time.
Neural network learns low-dimensional polynomials with sgd near the information-theoretic limit
Jason D Lee, Kazusato Oko, Taiji Suzuki, and Denny Wu · 2024
Closest in time.
Model balancing helps low-data training and fine-tuning
Zihang Liu, Yuanzhe Hu, Tianyu Pang, Yefan Zhou, Pu Ren, and Yaoqing Yang · 2024
Closest in time.
A random matrix theory perspective on the spectrum of learned features and asymptotic generalization capabilities
Yatin Dandi, Luca Pesce, Hugo Cui, Florent Krzakala, Yue Lu, and Bruno Loureiro · 2025
Closest in time.