Fetching the paper…
Reading the bibliography…
We analyze numerically the training dynamics of deep neural networks (DNN) by using methods developed in statistical physics of glassy systems.
The mathematics of diffusion
Crank, J · 1979
Earlier work this paper cites.
Analytical solution of the off-equilibrium dynamics of a long-range spin-glass model
Cugliandolo, L. F. and Kurchan, J · 1993
Earlier work this paper cites.
Phase space geometry and slow dynamics
Kurchan, J. and Laloux, L · 1996
Earlier work this paper cites.
Out of equilibrium dynamics in spin-glasses and other glassy systems
Bouchaud, J.-P., Cugliandolo, L. F., Kurchan, J., and Mezard, M · 1998
Earlier work this paper cites.
Efficient backprop
LeCun, Y., Bottou, L., Orr, G., and Müller, K.-R · 1998
Earlier work this paper cites.
Determining computational complexity from characteristic ‘phase transitions’
Monasson, R., Zecchina, R., Kirkpatrick, S., Selman, B., and Troyansky, L · 1999
Earlier work this paper cites.
Theory of phase-ordering kinetics
Bray, A. J · 2002
Earlier work this paper cites.
Analytic and algorithmic solution of random satisfiability problems
Mézard, M., Parisi, G., and Zecchina, R · 2002
Earlier work this paper cites.
Course 7: Dynamics of glassy systems
Cugliandolo, L. F · 2003
Earlier work this paper cites.
Spin-glass theory for pedestrians
Castellani, T. and Cavagna, A · 2005
Earlier work this paper cites.
On the rigidity of amorphous solids
Wyart, M · 2005
Earlier work this paper cites.
Cugliandolo-kurchan equations for dynamics of spin-glasses
Ben Arous, G., Dembo, A., and Guionnet, A · 2006
Earlier work this paper cites.
Rigorous inequalities between length and time scales in glassy systems
Montanari, A. and Semerjian, G · 2006
Earlier work this paper cites.
Gibbs states and the set of solutions of random constraint satisfaction problems
Krz̧akała, F., Montanari, A., Ricci-Tersenghi, F., Semerjian, G., and Zdeborová, L · 2007
Earlier work this paper cites.
Algorithmic barriers from phase transitions
Achlioptas, D. and Coja-Oghlan, A · 2008
Cited alongside, same era.
The jamming scenario: an introduction and outlook
Liu, A. J., Nagel, S. R., van Saarloos, W., and Wyart, M · 2010
Cited alongside, same era.
Theoretical perspective on the glass transition and amorphous materials
Berthier, L. and Biroli, G · 2011
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Cited alongside, same era.
Explorations on high dimensional landscapes
Sagun, L., Güney, V. U., Ben Arous, G., and LeCun, Y · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Ben Arous, G., and LeCun, Y · 2015
Gradient descent converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Later among the works it cites.
Stuck in a what? adventures in weight space
Lipton, Z. C · 2016
Later among the works it cites.
Singularity of the hessian in deep learning
Sagun, L., Bottou, L., and LeCun, Y · 2016
Later among the works it cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Soudry, D. and Carmon, Y · 2016
Later among the works it cites.
Statistical physics of inference: Thresholds and algorithms
Zdeborová, L. and Krzakala, F · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dynamics of stochastic gradient algorithms
Li, Q., Tai, C., and Weinan, E · 2015
Cited alongside, same era.
Unreasonable effectiveness of learning neural networks: From accessible states and robust ensembles to basic algorithmic schemes
Baldassi, C., Borgs, C., Chayes, J. T., Ingrosso, A., Lucibello, C., Saglietti, L., and Zecchina, R · 2016
Cited alongside, same era.
Slow relaxations and non-equilibrium dynamics in classical and quantum systems
Biroli, G · 2016
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2016
Cited alongside, same era.
Topology and geometry of deep rectified network optimization landscapes
Freeman, C. D. and Bruna, J · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Later among the works it cites.
Spectral gap estimates in mean field spin glasses
Ben Arous, G. and Jagannath, A · 2017
Later among the works it cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Later among the works it cites.
Three factors influencing minima in sgd
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2017
Later among the works it cites.
Models and algorithms for the next generation of glass transition studies
Ninarello, A., Berthier, L., and Coslovich, D · 2017
Later among the works it cites.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Güney, V. U., Dauphin, Y., and Bottou, L · 2017
Later among the works it cites.
Activated aging dynamics and effective trap model description in the random energy model
Baity-Jesi, M., Biroli, G., and Cammarota, C · 2018
Closest in time.
Theory for swap acceleration near the glass and jamming transitions
Brito, C., Lerner, E., and Wyart, M · 2018
Closest in time.