Fetching the paper…
Reading the bibliography…
We derive a nonlinear integro-differential transport equation describing collective evolution of weights under gradient descent in large-width neural-network-like models.
Computing with infinite networks
Williams, C. K. (1997) · 1997
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Bayesian learning for neural networks
Neal, R. M. (2012) · 2012
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Nesterov, Y. (2013) · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S. (2013) · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y. (2015) · 2015
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K. (2016) · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Poole, B., Lahiri, S., Raghu, M., Sohl-Dickstein, J., and Ganguli, S. (2016) · 2016
Cited alongside, same era.
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J. (2016) · 2016
Cited alongside, same era.
Deep linear neural networks with arbitrary loss: All local minima are global
Laurent, T. and von Brecht, J. (2017) · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
Pennington, J. and Bahri, Y. (2017) · 2017
Later among the works it cites.
Spurious local minima are common in two-layer relu neural networks
Safran, I. and Shamir, O. (2017) · 2017
Later among the works it cites.
Opening the black box of deep neural networks via information
Shwartz-Ziv, R. and Tishby, N. (2017) · 2017
Later among the works it cites.
Tian, Y. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nguyen, Q. and Hein, M. (2017) · 2017
Cited alongside, same era.
Bartlett, P. L., Helmbold, D. P., and Long, P. M. (2018) · 2018
Closest in time.
Deep relaxation: partial differential equations for optimizing deep neural networks
Chaudhari, P., Oberman, A., Osher, S., Soatto, S., and Carlier, G. (2018) · 2018
Closest in time.