Fetching the paper…
Reading the bibliography…
Linear networks provide valuable insights into the workings of neural networks in general.
Arora, S., Du, S. S., Hu, W., Li, Z., and Wang, R · 1901
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Arora, S., Cohen, N., Hu, W., and Luo, Y · 1905
Earlier work this paper cites.
The approximation of one matrix by another of lower rank
Eckart, C. and Young, G · 1936
Earlier work this paper cites.
Symmetric gauge functions and unitarily invariant norms
Mirsky, L · 1966
Earlier work this paper cites.
On the trajectories of the gradient of an analytical function
Lojasiewicz, S · 1982
Earlier work this paper cites.
A generalization of the eckart-young-mirsky matrix approximation theorem
Golub, G. H., Hoffman, A., and Stewart, G. W · 1987
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K · 1989
Earlier work this paper cites.
A primer of real analytic functions
Parks, H. R. and Krantz, S · 1992
Earlier work this paper cites.
Global analysis of oja’s flow for neural networks
Yan, W.-Y., Helmke, U., and Moore, J. B · 1994
Earlier work this paper cites.
Critical points of matrix least squares distance functions
Helmke, U. and Shayman, M. A · 1995
Earlier work this paper cites.
On the distribution of the largest eigenvalue in principal components analysis
Johnstone, I. M. et al · 2001
Earlier work this paper cites.
Convergence of the iterates of descent methods for analytic cost functions
Absil, P.-A., Mahony, R., and Andrews, B · 2005
Earlier work this paper cites.
Lectures on analytic differential equations , volume 86
Illashenko and Yakovenko · 2008
Earlier work this paper cites.
Deflation methods for sparse pca
Mackey, L. W · 2009
Earlier work this paper cites.
Large-scale sparse principal component analysis with application to text data
Zhang, Y. and Ghaoui, L. E · 2011
Earlier work this paper cites.
Machine learning: a probabilistic perspective
Murphy, K. P · 2012
Earlier work this paper cites.
How close is the sample covariance matrix to the actual covariance matrix?
Vershynin, R · 2012
Earlier work this paper cites.
Optimal detection of sparse principal components in high dimension
Berthet, Q., Rigollet, P., et al · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
Information-theoretically optimal sparse pca
Deshpande, Y. and Montanari, A · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Earlier work this paper cites.
Identity matters in deep learning
Hardt, M. and Ma, T · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K · 2016
Cited alongside, same era.
Gradient descent converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Cited alongside, same era.
Differentiating the singular value decomposition
Townsend, J · 2016
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Brutzkus, A. and Globerson, A · 2017
On the spectral bias of neural networks
Rahaman, N., Baratin, A., Arpit, D., Draxler, F., Lin, M., Hamprecht, F. A., Bengio, Y., and Courville, A · 2018
Later among the works it cites.
Exponential convergence time of gradient descent for one-dimensional deep linear neural networks
Shamir, O · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Later among the works it cites.
Learning one-hidden-layer relu networks via gradient descent
Zhang, X., Yu, Y., Wang, L., and Gu, Q · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B. E., Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2017
Cited alongside, same era.
First-order methods almost always avoid saddle points
Lee, J. D., Panageas, I., Piliouras, G., Simchowitz, M., Jordan, M. I., and Recht, B · 2017
Cited alongside, same era.
Depth creates no bad local minima
Lu, H. and Kawaguchi, K · 2017
Cited alongside, same era.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Tian, Y · 2017
Cited alongside, same era.
Global optimality conditions for deep neural networks
Yun, C., Sra, S., and Jadbabaie, A · 2017
Cited alongside, same era.
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Later among the works it cites.
Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers
Bah, B., Rauhut, H., Terstiege, U., and Westdickenberg, M · 2019
Later among the works it cites.
Gradient descent with identity initialization efficiently learns positive-definite linear transformations by deep residual networks
Bartlett, P. L., Helmbold, D. P., and Long, P. M · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q · 2019
Later among the works it cites.
How much over-parameterization is sufficient to learn deep relu networks?
Chen, Z., Cao, Y., Zou, D., and Gu, Q · 2019
Later among the works it cites.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Later among the works it cites.
Width provably matters in optimization for deep linear neural networks
Du, S. S. and Hu, W · 2019
Later among the works it cites.
Moses: A streaming algorithm for linear dimensionality reduction
Eftekhari, A., Hauser, R., and Grammenos, A · 2019
Later among the works it cites.
On connected sublevel sets in deep learning
Nguyen, Q · 2019
Later among the works it cites.
Trainability and data-dependent initialization of over-parameterized relu neural networks
Shin, Y. and Karniadakis, G. E · 2019
Later among the works it cites.
On learning over-parameterized neural networks: A functional approximation prospective
Su, L. and Yang, P · 2019
Later among the works it cites.
Pure and spurious critical points: a geometric study of linear networks
Trager, M., Kohn, K., and Bruna, J · 2019
Later among the works it cites.
Global convergence of adaptive gradient methods for an over-parameterized neural network
Wu, X., Du, S. S., and Ward, R · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Yehudai, G. and Shamir, O · 2019
Later among the works it cites.
Fast convergence of natural gradient descent for overparameterized neural networks
Zhang, G., Martens, J., and Grosse, R · 2019
Later among the works it cites.
The global optimization geometry of shallow linear neural networks
Zhu, Z., Soudry, D., Eldar, Y. C., and Wakin, M. B · 2019
Later among the works it cites.
Provable benefit of orthogonal initialization in optimizing deep linear networks
Hu, W., Xiao, L., and Pennington, J · 2020
Closest in time.