Fetching the paper…
Reading the bibliography…
The dynamics of Deep Linear Networks (DLNs) is dramatically affected by the variance $\sigma^2$ of the parameters at initialization $\theta_0$.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Mei, Song, Misiakiewicz, Theodor, & Montanari, Andrea. 2019 · 1902
Earlier work this paper cites.
Yang, Greg. 2019 · 1902
Earlier work this paper cites.
On Exact Computation with an Infinitely Wide Neural Net
Arora, Sanjeev, Du, Simon S, Hu, Wei, Li, Zhiyuan, Salakhutdinov, Ruslan, & Wang, Ruosong. 2019c · 1904
Earlier work this paper cites.
Disentangling feature and lazy learning in deep neural networks: an empirical study
Geiger, Mario, Spigler, Stefano, Jacot, Arthur, & Wyart, Matthieu. 2019 · 1906
Earlier work this paper cites.
Dynamics of deep neural networks and neural tangent hierarchy
Huang, Jiaoyang, & Yau, Horng-Tzer. 2019 · 1909
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, Pierre, & Hornik, Kurt. 1989 · 1989
Earlier work this paper cites.
Liu, Chaoyue, Zhu, Libin, & Belkin, Mikhail. 2020 · 2003
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, Roman. 2010 · 2010
Earlier work this paper cites.
Li, Zhiyuan, Luo, Yuping, & Lyu, Kaifeng. 2020 · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, Andrew M., McClelland, James L., & Ganguli, Surya. 2014 · 2014
Earlier work this paper cites.
Deep Learning without Poor Local Minima
Kawaguchi, Kenji. 2016 · 2016
Earlier work this paper cites.
Gradient Descent Only Converges to Minimizers
Lee, Jason D., Simchowitz, Max, Jordan, Michael I., & Recht, Benjamin. 2016 · 2016
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks
Advani, Madhu S., & Saxe, Andrew M. 2017 · 2017
Cited alongside, same era.
Gradient descent can take exponential time to escape saddle points
Du, Simon S., Jin, Chi, Lee, Jason D., Jordan, Michael I., Póczos, Barnabás, & Singh, Aarti. 2017 · 2017
Cited alongside, same era.
How to Escape Saddle Points Efficiently
Jin, Chi, Ge, Rong, Netrapalli, Praneeth, Kakade, Sham M., & Jordan, Michael I. 2017 · 2017
Cited alongside, same era.
Deep neural networks as gaussian processes
Lee, Jaehoon, Bahri, Yasaman, Novak, Roman, Schoenholz, Samuel S, Pennington, Jeffrey, & Sohl-Dickstein, Jascha. 2017 · 2017
Cited alongside, same era.
Gaussian Process Behaviour in Wide Deep Neural Networks
de G. Matthews, Alexander G., Hron, Jiri, Rowland, Mark, Turner, Richard E., & Ghahramani, Zoubin. 2018 · 2018
Cited alongside, same era.
A mathematical theory of semantic development in deep neural networks
Saxe, Andrew M., McClelland, James L., & Ganguli, Surya. 2019 · 2019
Later among the works it cites.
Implicit Bias of Gradient Descent for Wide Two-layer Neural Networks Trained with the Logistic Loss
Chizat, Lénaïc, & Bach, Francis. 2020 · 2020
Later among the works it cites.
Training Linear Neural Networks: Non-Local Convergence and Complexity Results
Eftekhari, Armin. 2020 · 2020
Later among the works it cites.
The Implicit Bias of Depth: How Incremental Learning Drives Generalization
Gissin, Daniel, Shalev-Shwartz, Shai, & Daniely, Amit. 2020 · 2020
Later among the works it cites.
Directional convergence and alignment in deep learning
Ji, Ziwei, & Telgarsky, Matus. 2020 · 2020
Later among the works it cites.
Gradient Descent Maximizes the Margin of Homogeneous Neural Networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Jacot, Arthur, Gabriel, Franck, & Hongler, Clément. 2018 · 2018
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ji, Ziwei, & Telgarsky, Matus. 2018 · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Mei, Song, Montanari, Andrea, & Nguyen, Phan-Minh. 2018 · 2018
Cited alongside, same era.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
Rotskoff, Grant, & Vanden-Eijnden, Eric. 2018 · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, Daniel, Hoffer, Elad, Nacson, Mor Shpigel, Gunasekar, Suriya, & Srebro, Nathan. 2018 · 2018
Cited alongside, same era.
Gradient Descent Provably Optimizes Over-parameterized Neural Networks
Du, Simon S., Zhai, Xiyu, Poczos, Barnabas, & Singh, Aarti. 2019 · 2019
Cited alongside, same era.
Implicit Regularization of Discrete Gradient Dynamics in Linear Neural Networks
Gidel, Gauthier, Bach, Francis, & Lacoste-Julien, Simon. 2019 · 2019
Cited alongside, same era.
Lyu, Kaifeng, & Li, Jian. 2020 · 2020
Later among the works it cites.
Implicit Bias in Deep Linear Classification: Initialization Scale vs Training Accuracy
Moroshko, Edward, Woodworth, Blake E, Gunasekar, Suriya, Lee, Jason D, Srebro, Nati, & Soudry, Daniel. 2020 · 2020
Later among the works it cites.
Kernel and Rich Regimes in Overparametrized Models
Woodworth, Blake, Gunasekar, Suriya, Savarese, Pedro, Moroshko, Edward, Golan, Itay, Lee, Jason, Soudry, Daniel, & Srebro, Nathan. 2020 · 2020
Later among the works it cites.
Feature Learning in Infinite-Width Neural Networks
Yang, Greg, & Hu, Edward J. 2020 · 2020
Later among the works it cites.
Learning deep models: Critical points and local openness
Nouiehed, Maher, & Razaviyayn, Meisam. 2021 · 2021
Closest in time.
Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances
Simsek, Berfin, Ged, François, Jacot, Arthur, Spadaro, Francesco, Hongler, Clement, Gerstner, Wulfram, & Brea, Johanni. 2021 · 2021
Closest in time.
A unifying view on implicit bias in training linear neural networks
Yun, Chulhee, Krishnan, Shankar, & Mobahi, Hossein. 2021 · 2021
Closest in time.