Fetching the paper…
Reading the bibliography…
We analyze algorithms for approximating a function $f(x) = \Phi x$ mapping $\Re^d$ to $\Re^d$ using deep linear neural networks, i.e.
On the existence and uniqueness of the real logarithm of a matrix
Culver, Walter J · 1966
Earlier work this paper cites.
Topics in Matrix Analysis
Horn, R. A · 1986
Earlier work this paper cites.
Efficient agnostic learning of neural networks with bounded fan-in
Lee, Wee Sun, Bartlett, Peter L, and Williamson, Robert C · 1996
Earlier work this paper cites.
Matrix algebra from a statistician’s perspective , volume 1
Harville, D. A · 1997
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
Zhong, Kai, Song, Zhao, Jain, Prateek, Bartlett, Peter L., and Dhillon, Inderjit S · 1997
Earlier work this paper cites.
Convex Optimization
Boyd, S. P. and Vandenberghe, L · 2004
Earlier work this paper cites.
Matrix analysis
Horn, Roger A and Johnson, Charles R · 2013
Earlier work this paper cites.
Hessian schatten-norm regularization for linear inverse problems
Lefkimmiatis, Stamatios, Ward, John Paul, and Unser, Michael · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2013
Earlier work this paper cites.
Learning polynomials with neural networks
Andoni, A., Panigrahy, R., Valiant, G., and Zhang, L · 2014
Earlier work this paper cites.
Provable bounds for learning some deep representations
Arora, Sanjeev, Bhaskara, Aditya, Ge, Rong, and Ma, Tengyu · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Livni, Roi, Shalev-Shwartz, Shai, and Shamir, Ohad · 2014
Cited alongside, same era.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Janzamin, Majid, Sedghi, Hanie, and Anandkumar, Anima · 2015
Cited alongside, same era.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Daniely, A., Frostig, R., and Singer, Y · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Identity matters in deep learning
Hardt, M. and Ma, T · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with relu activation
Li, Yuanzhi and Yuan, Yang · 2017
Later among the works it cites.
The loss surface of deep and wide neural networks
Nguyen, Quynh and Hein, Matthias · 2017
Later among the works it cites.
How regularization affects the critical points in linear networks
Taghvaei, A., Kim, J. W., and Mehta, P · 2017
Later among the works it cites.
On the learnability of fully-connected neural networks
Zhang, Yuchen, Lee, Jason, Wainwright, Martin, and Jordan, Michael · 2017
Later among the works it cites.
Bartlett, Peter L., Evans, Steven N., and Long, Philip M · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Cited alongside, same era.
On the quality of the initial basin in overspecified neural networks
Safran, Itay and Shamir, Ohad · 2016
Cited alongside, same era.
l1-regularized neural networks are improperly learnable in polynomial time
Zhang, Yuchen, Lee, Jason D, and Jordan, Michael I · 2016
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Brutzkus, A. and Globerson, A · 2017
Cited alongside, same era.
SGD learns the conjugate kernel class of the network
Daniely, A · 2017
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Ge, R., Jin, C., and Zheng, Y
Cited in the paper.
Learning one-hidden-layer neural networks with landscape design
Ge, Rong, Lee, Jason D, and Ma, Tengyu
Cited in the paper.
Closest in time.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, Alon, Globerson, Amir, Malach, Eran, and Shalev-Shwartz, Shai · 2018
Closest in time.
Learning one-hidden-layer neural networks with landscape design
Ge, Rong, Lee, Jason D, and Ma, Tengyu · 2018
Closest in time.
Skip connections eliminate singularities
Orhan, A Emin and Pitkow, Xaq · 2018
Closest in time.
Electron-proton dynamics in deep learning
Zhang, Q., Panigrahy, R., and Sachdeva, S · 2018
Closest in time.