Fetching the paper…
Reading the bibliography…
Deep learning succeeds by doing hierarchical feature learning, yet tuning hyper-parameters (HP) such as initialization scales, learning rates etc., only give indirect control over this behavior.
Introduction to optimization
Boris T. Polyak · 1987
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Earlier work this paper cites.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Earlier work this paper cites.
L2 regularization versus batch and weight normalization
Twan Van Laarhoven · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis Bach · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Boris Hanin · 2018
Earlier work this paper cites.
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick · 2018
Earlier work this paper cites.
Neural Tangent Kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2018
Earlier work this paper cites.
Neural networks as interacting particle systems: Asymptotic convexity of the loss landscape and universal scaling of the approximation error
Grant M. Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lénaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Revisiting the Polyak step size
Elad Hazan and Sham Kakade · 2019
Cited alongside, same era.
Spectrum concentration in deep residual learning: a free probability approach
Zenan Ling and Robert C. Qiu · 2019
Cited alongside, same era.
The future is log-Gaussian: ResNets and their infinite-depth-and-width limit at initialization
Mufan Li, Mihai Nica, and Dan Roy · 2021
Later among the works it cites.
Tensor programs IV: Feature learning in infinite-width neural networks
Greg Yang and Edward J Hu · 2021
Later among the works it cites.
Tuning large neural networks via zero-shot hyperparameter transfer
Greg Yang, Edward Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2021
Later among the works it cites.
Robust training of neural networks using scale invariant architectures
Zhiyuan Li, Srinadh Bhojanapalli, Manzil Zaheer, Sashank Reddi, and Sanjiv Kumar · 2022
Later among the works it cites.
Feature learning and signal propagation in deep neural networks
Yizhang Lou, Chris E Mingard, and Soufiane Hayou · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wojciech Tarnowski, Piotr Warchoł, Stanisław Jastrzȩbski, Jacek Tabor, and Maciej Nowak · 2019
Cited alongside, same era.
A mathematical model for automatic differentiation in machine learning
Jérôme Bolte and Edouard Pauwels · 2020
Cited alongside, same era.
Products of many large random matrices and gradients in deep neural networks
Boris Hanin and Mihai Nica · 2020
Cited alongside, same era.
Randomized numerical linear algebra: Foundations and algorithms
Per-Gunnar Martinsson and Joel A Tropp · 2020
Cited alongside, same era.
Ruosi Wan, Zhanxing Zhu, Xiangyu Zhang, and Jian Sun · 2020
Cited alongside, same era.
Tensor programs III: Neural matrix laws
Greg Yang · 2020
Cited alongside, same era.
Neural networks as kernel learners: The silent alignment effect
Alexander Atanasov, Blake Bordelon, and Cengiz Pehlevan · 2021
Cited alongside, same era.
Pierre Marion, Adeline Fermanian, Gérard Biau, and Jean-Philippe Vert · 2022
Later among the works it cites.
Spectral evolution and invariance in linear-width neural networks
Zhichao Wang, Andrew Engel, Anand Sarwate, Ioana Dumitriu, and Tony Chiang · 2022
Later among the works it cites.
Stabilize deep ResNet with a sharp scaling factor
Huishuai Zhang, Da Yu, Mingyang Yi, Wei Chen, and Tie-Yan Liu · 2022
Later among the works it cites.
Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit
Blake Bordelon, Lorenzo Noci, Mufan Bill Li, Boris Hanin, and Cengiz Pehlevan · 2023
Closest in time.
Samy Jelassi, Boris Hanin, Ziwei Ji, Sashank J Reddi, Srinadh Bhojanapalli, and Sanjiv Kumar · 2023
Closest in time.
Feature-learning networks are consistent across widths at realistic scales
Nikhil Vyas, Alexander Atanasov, Blake Bordelon, Depen Morwani, Sabarish Sainathan, and Cengiz Pehlevan · 2023
Closest in time.
Feature learning in infinite-depth neural networks
Greg Yang, Dingli Yu, Chen Zhu, and Soufiane Hayou · 2023
Closest in time.
Deep linear networks for regression are implicitly regularized towards flat minima
Pierre Marion and Lénaïc Chizat · 2024
Closest in time.