Fetching the paper…
Reading the bibliography…
Recently, there has been growing evidence that if the width and depth of a neural network are scaled toward the so-called rich feature learning limit (\mup and its depth extension), then some hyperparameters -- such as the learning rate -- exhibit transfer from small to very large models.
Über die eigenwerte bei den differentialgleichungen der mathematischen physik
R. Courant · 1920
Earlier work this paper cites.
Bayesian Learning for Neural Networks
R. M. Neal · 1995
Earlier work this paper cites.
Information geometry
S. Amari · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Earlier work this paper cites.
The symmetric eigenvalue problem
B. N. Parlett · 1998
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 2002
Earlier work this paper cites.
Chaos in dynamical systems
E. Ott · 2002
Earlier work this paper cites.
Random search for hyper-parameter optimization
J. Bergstra and Y. Bengio · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Y. Nesterov · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Efficient transfer learning method for automatic hyperparameter tuning
D. Yogatama and G. Mann · 2014
Earlier work this paper cites.
Hyperparameter search in machine learning, 2015
M. Claesen and B. D. Moor · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
D. Maclaurin, D. Duvenaud, and R. Adams · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Earlier work this paper cites.
Scalable bayesian optimization using deep neural networks
J. Snoek, O. Rippel, K. Swersky, R. Kiros, N. Satish, N. Sundaram, M. Patwary, M. Prabhat, and R. Adams · 2015
Earlier work this paper cites.
Understanding edge-of-stability training dynamics with a minimalist example
X. Zhu, Z. Wang, X. Wang, M. Zhou, and R. Ge · 2015
Earlier work this paper cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Earlier work this paper cites.
Non-stochastic best arm identification and hyperparameter optimization
K. Jamieson and A. Talwalkar · 2016
Earlier work this paper cites.
Second-order optimization for neural networks
J. Martens · 2016
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
L. Sagun, L. Bottou, and Y. LeCun · 2016
Earlier work this paper cites.
Practical gauss-newton optimisation for deep learning
A. Botev, H. Ritter, and D. Barber · 2017
Earlier work this paper cites.
Forward and reverse gradient-based hyperparameter optimization
L. Franceschi, M. Donini, P. Frasconi, and M. Pontil · 2017
Earlier work this paper cites.
Deep neural networks as gaussian processes
J. Lee, Y. Bahri, R. Novak, S. S. Schoenholz, J. Pennington, and J. Sohl-Dickstein · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
L. Sagun, U. Evci, V. U. Guney, Y. Dauphin, and L. Bottou · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Earlier work this paper cites.
Deep convolutional networks as shallow gaussian processes
A. Garriga-Alonso, C. E. Rasmussen, and L. Aitchison · 2018
Earlier work this paper cites.
Gradient descent happens in a tiny subspace
G. Gur-Ari, D. A. Roberts, and E. Dyer · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Earlier work this paper cites.
On the relation between the sharpest directions of dnn loss and the sgd step length
S. Jastrzkebski, Z. Kenton, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2018
Earlier work this paper cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
L. Li, K. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar · 2018
Cited alongside, same era.
Scalable hyperparameter transfer learning
V. Perrone, R. Jenatton, M. W. Seeger, and C. Archambeau · 2018
Cited alongside, same era.
On exact computation with an infinitely wide neural net
S. Arora, S. S. Du, W. Hu, Z. Li, R. R. Salakhutdinov, and R. Wang · 2019
Cited alongside, same era.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2019
Cited alongside, same era.
An investigation into neural net optimization via hessian eigenvalue density
B. Ghorbani, S. Krishnan, and Y. Xiao · 2019
Cited alongside, same era.
Sgd: General analysis and improved rates
R. M. Gower, N. Loizou, X. Qian, A. Sailanbayev, E. Shulgin, and P. Richtárik · 2019
Cited alongside, same era.
Tensor programs iv: Feature learning in infinite-width neural networks
G. Yang and E. J. Hu · 2021
Later among the works it cites.
Asymptotics of representation learning in finite bayesian neural networks
J. Zavatone-Veth, A. Canatar, B. Ruben, and C. Pehlevan · 2021
Later among the works it cites.
Second-order regression models exhibit progressive sharpening to the edge of stability
A. Agarwala, F. Pedregosa, and J. Pennington · 2022
Later among the works it cites.
Understanding the unstable convergence of gradient descent
K. Ahn, J. Zhang, and S. Sra · 2022
Later among the works it cites.
Understanding gradient descent on the edge of stability in deep learning
S. Arora, Z. Li, and A. Panigrahi · 2022
Later among the works it cites.
Self-consistent dynamical field theory of kernel evolution in wide neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finite depth and width corrections to the neural tangent kernel
B. Hanin and M. Nica · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington · 2019
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Y. Li, C. Wei, and T. Ma · 2019
Cited alongside, same era.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Truncated back-propagation for bilevel optimization
A. Shaban, C.-A. Cheng, N. Hatch, and B. Boots · 2019
Cited alongside, same era.
B. Bordelon and C. Pehlevan · 2022
Later among the works it cites.
Adaptive gradient methods at the edge of stability
J. M. Cohen, B. Ghorbani, S. Krishnan, N. Agarwal, S. Medapati, M. Badura, D. Suo, D. Cardoze, Z. Nado, G. E. Dahl, et al · 2022
Later among the works it cites.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
A. Damian, E. Nichani, and J. D. Lee · 2022
Later among the works it cites.
On the infinite-depth limit of finite-width neural networks
S. Hayou · 2022
Later among the works it cites.
The neural covariance sde: Shaped infinite depth-and-width networks at initialization
M. Li, M. Nica, and D. Roy · 2022
Later among the works it cites.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
L. Noci, S. Anagnostidis, L. Biggio, A. Orvieto, S. P. Singh, and A. Lucchi · 2022
Later among the works it cites.
Meta-principled family of hyperparameter scaling strategies
S. Yaida · 2022
Later among the works it cites.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
G. Yang, E. J. Hu, I. Babuschkin, S. Sidor, X. Liu, D. Farhi, N. Ryder, J. Pachocki, W. Chen, and J. Gao · 2022
Later among the works it cites.
Dynamics of finite width kernel and prediction fluctuations in mean field neural networks
B. Bordelon and C. Pehlevan · 2023
Later among the works it cites.
Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit
B. Bordelon, L. Noci, M. B. Li, B. Hanin, and C. Pehlevan · 2023
Later among the works it cites.
Steering deep feature learning with backward aligned feature updates
L. Chizat and P. Netrapalli · 2023
Later among the works it cites.
Handbook of convergence theorems for (stochastic) gradient methods
G. Garrigos and R. M. Gower · 2023
Later among the works it cites.
Width and depth limits commute in residual networks
S. Hayou and G. Yang · 2023
Later among the works it cites.
Maximal initial learning rates in deep relu networks
G. Iyer, B. Hanin, and D. Rolnick · 2023
Later among the works it cites.
Depth dependence of mup learning rates in relu mlps
S. Jelassi, B. Hanin, Z. Ji, S. J. Reddi, S. Bhojanapalli, and S. Kumar · 2023
Later among the works it cites.
Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and width
D. S. Kalra and M. Barkeshli · 2023
Later among the works it cites.
D. S. Kalra, T. He, and M. Barkeshli · 2023
Later among the works it cites.
Differential equation scaling limits of shaped and unshaped neural networks
M. B. Li and M. Nica · 2023
Later among the works it cites.
The shaped transformer: Attention models in the infinite depth-and-width limit
L. Noci, C. Li, M. B. Li, B. He, T. Hofmann, C. Maddison, and D. M. Roy · 2023
Later among the works it cites.
Asdl: A unified interface for gradient preconditioning in pytorch
K. Osawa, S. Ishikawa, R. Yokota, S. Li, and T. Hoefler · 2023
Later among the works it cites.
On the interplay between stepsize tuning and progressive sharpening
V. Roulet, A. Agarwala, and F. Pedregosa · 2023
Later among the works it cites.
Trajectory alignment: understanding the edge of stability phenomenon via bifurcation theory
M. Song and C. Yun · 2023
Later among the works it cites.
Feature-learning networks are consistent across widths at realistic scales
N. Vyas, A. Atanasov, B. Bordelon, D. Morwani, S. Sainathan, and C. Pehlevan · 2023
Later among the works it cites.
Tensor programs ivb: Adaptive optimization in the infinite-width limit
G. Yang and E. Littwin · 2023
Later among the works it cites.
Tensor programs vi: Feature learning in infinite-depth neural networks
G. Yang, D. Yu, C. Zhu, and S. Hayou · 2023
Later among the works it cites.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, Y. Du, C. Yang, Y. Chen, Z. Chen, J. Jiang, R. Ren, Y. Li, X. Tang, Z. Liu, P. Liu, J.-Y. Nie, and J.-R. Wen · 2023
Later among the works it cites.
Infinite limits of multi-head transformer dynamics
B. Bordelon, H. T. Chaudhry, and C. Pehlevan · 2024
Closest in time.