Fetching the paper…
Reading the bibliography…
In this paper we establish a connection between non-convex optimization methods for training deep neural networks and nonlinear partial differential equations (PDEs).
Uber eine Klasse von Mittelbildungen mit Anwendungen auf die Determinantentheorie
Schur, I. (1923) · 1923
Earlier work this paper cites.
Brownian motion in a field of force and the diffusion model of chemical reactions
Kramers, H. A. (1940) · 1940
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Proximité et dualité dans un espace hilbertien
Moreau, J.-J. (1965) · 1965
Earlier work this paper cites.
Monotone operators and the proximal point algorithm
Rockafellar, R. T. (1976) · 1976
Earlier work this paper cites.
Inequalities: theory of majorization and its applications
Marshall, A. W., Olkin, I., and Arnold, B. C. (1979) · 1979
Earlier work this paper cites.
Diffusions hypercontractives
Bakry, D. and Émery, M. (1985) · 1983
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o (1/k2)
Nesterov, Y. (1983) · 1983
Earlier work this paper cites.
Fokker-planck equation
Risken, H. (1984) · 1984
Earlier work this paper cites.
Diffusions for global optimization
Geman, S. and Hwang, C.-R. (1986) · 1986
Earlier work this paper cites.
Diffusion for global optimization in r
Chiang, T.-S., Hwang, C.-R., and Sheu, S. (1987) · 1987
Earlier work this paper cites.
Asymptotic global behavior for stochastic approximation and diffusions with slowly decreasing noise effects: global minimization via Monte Carlo
Kushner, H. (1987) · 1987
Earlier work this paper cites.
Spin glass theory and beyond
Mezard, M., Parisi, G., and Virasoo, M. A. (1987) · 1987
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1988) · 1988
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K. (1989) · 1989
Earlier work this paper cites.
Partial differential equations
Evans, L. C. (1998) · 1998
Earlier work this paper cites.
The variational formulation of the Fokker–Planck equation
Jordan, R., Kinderlehrer, D., and Otto, F. (1998) · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998) · 1998
Earlier work this paper cites.
Convex analysis and optimization
Bertsekas, D., Nedi, A., Ozdaglar, A., et al. (2003) · 2003
Earlier work this paper cites.
Semiconcave functions, Hamilton-Jacobi equations, and optimal control
Cannarsa, P. and Sinestrari, C. (2004) · 2004
Earlier work this paper cites.
Metastability: A potential theoretic approach
Bovier, A. and den Hollander, F. (2006) · 2006
Earlier work this paper cites.
Contractions in the 2-wasserstein length space and thermalization of granular media
Carrillo, J. A., McCann, R. J., and Villani, C. (2006) · 2006
Cited alongside, same era.
Controlled Markov processes and viscosity solutions
Fleming, W. H. and Soner, H. M. (2006) · 2006
Cited alongside, same era.
Large population stochastic dynamic games: closed-loop McKean-Vlasov systems and the Nash certainty equivalence principle
Huang, M., Malhamé, R. P., Caines, P. E., et al. (2006) · 2006
Cited alongside, same era.
Convergent difference schemes for degenerate elliptic and parabolic equations: Hamilton-Jacobi equations and free boundary problems
Oberman, A. M. (2006) · 2006
Cited alongside, same era.
The statistics of critical points of Gaussian fields on large-dimensional spaces
Bray, A. and Dean, D. (2007) · 2007
Cited alongside, same era.
Stochastic processes and applications
Pavliotis, G. A. (2014) · 2014
Later among the works it cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A., McClelland, J., and Ganguli, S. (2014) · 2014
Later among the works it cites.
Striving for simplicity: The all convolutional net
Springenberg, J., Dosovitskiy, A., Brox, T., and Riedmiller, M. (2014) · 2014
Later among the works it cites.
Subdominant dense clusters allow for simple learning and high computational performance in neural networks with discrete synapses
Baldassi, C., Ingrosso, A., Lucibello, C., Saglietti, L., and Zecchina, R. (2015) · 2015
Later among the works it cites.
On the energy landscape of deep networks
Chaudhari, P. and Soatto, S. (2015) · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Replica symmetry breaking condition exposed by random matrix calculation of landscape complexity
Fyodorov, Y. and Williams, I. (2007) · 2007
Cited alongside, same era.
Mean field games
Lasry, J.-M. and Lions, P.-L. (2007) · 2007
Cited alongside, same era.
Multiscale methods: averaging and homogenization
Pavliotis, G. A. and Stuart, A. (2008) · 2008
Cited alongside, same era.
Learning multiple layers of features from tiny images
Krizhevsky, A. (2009) · 2009
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A. (2009) · 2009
Cited alongside, same era.
An analysis of single-layer networks in unsupervised feature learning
Coates, A., Lee, H., and Ng, A. Y. (2010) · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Cited alongside, same era.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Ben Arous, G., and LeCun, Y. (2015) · 2015
Later among the works it cites.
Global optimality in tensor factorization, deep learning, and beyond
Haeffele, B. and Vidal, R. (2015) · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Later among the works it cites.
Variational dropout and the local reparameterization trick
Kingma, D. P., Salimans, T., and Welling, M. (2015) · 2015
Later among the works it cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G. (2015) · 2015
Later among the works it cites.
Dynamics of stochastic gradient algorithms
Li, Q., Tai, C., et al. (2015) · 2015
Later among the works it cites.
Optimal transport for applied mathematicians
Santambrogio, F. (2015) · 2015
Later among the works it cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015) · 2015
Later among the works it cites.
Deep learning with elastic averaging SGD
Zhang, S., Choromanska, A., and LeCun, Y. (2015) · 2015
Later among the works it cites.
Information dropout: learning optimal representations through noise
Achille, A. and Soatto, S. (2016) · 2016
Later among the works it cites.
Local entropy as a measure for sampling solutions in constraint satisfaction problems
Baldassi, C., Ingrosso, A., Lucibello, C., Saglietti, L., and Zecchina, R. (2016b) · 2016
Later among the works it cites.
Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R. (2016) · 2016
Later among the works it cites.
Noisy activation functions
Gulcehre, C., Moczulski, M., Denil, M., and Bengio, Y. (2016) · 2016
Later among the works it cites.
Training Recurrent Neural Networks by Diffusion
Mobahi, H. (2016) · 2016
Later among the works it cites.
Singularity of the Hessian in Deep Learning
Sagun, L., Bottou, L., and LeCun, Y. (2016) · 2016
Later among the works it cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Soudry, D. and Carmon, Y. (2016) · 2016
Later among the works it cites.