Fetching the paper…
Reading the bibliography…
Traditionally, neural networks are parameterized using optimization procedures such as stochastic gradient descent, RMSProp and ADAM.
Ridge regression: Biased estimation for nonorthogonal problems
A. Hoerl and R. Kennard · 1970
Earlier work this paper cites.
Optimization by simulated annealing
S. Kirkpatrick, C.D. Gelatt, and M.P. Vecchi · 1983
Earlier work this paper cites.
A unified formulation of the constant temperature molecular dynamics methods
S. Nosé · 1984
Earlier work this paper cites.
Canonical dynamics: Equilibrium phase-space distributions
W. Hoover · 1985
Earlier work this paper cites.
Diffusion in a rough potential
R. Zwanzig · 1988
Earlier work this paper cites.
Markov Chain Monte Carlo maximum likelihood
C.J. Geyer · 1991
Earlier work this paper cites.
Simulated tempering: a new Monte Carlo scheme
E. Marinari and G. Parisi · 1992
Earlier work this paper cites.
Stability of Markovian processes II: Continuous-time processes and sampled chains
S.P. Meyn and R.L. Tweedie · 1993
Earlier work this paper cites.
Bayesian regularization and pruning using a Laplace prior
P. Williams · 1995
Earlier work this paper cites.
Exponential convergence of Langevin distributions and their discrete approximations
G.O. Roberts and R.L. Tweedie · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
Ergodicity for SDEs and approximations: locally Lipschitz vector fields and degenerate noise
J.C. Mattingly, A.M.Stuart, and D.J. Higham · 2002
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
H. Kushner and G.G. Yin · 2003
Earlier work this paper cites.
Handbook of Stochastic Methods for Physics, Chemistry, and the Natural Sciences
C. Gardiner · 2004
Earlier work this paper cites.
Observations on rate theory for rugged energy landscapes
E. Pollak, A. Auerbach, and P. Talkner · 2008
Earlier work this paper cites.
Hypocoercivity for kinetic equations with linear relaxation terms
J. Dolbeault, C. Mouhot, and C. Schmeiser · 2009
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
K. Jarrett, K. Kavukcuoglu, M. Ranzato, and Y. LeCun · 2009
Earlier work this paper cites.
Dlib-ml: A machine learning toolkit
Davis E. King · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Deep sparse rectifier networks
X. Glorot, A. Bordes, and Y. Bengio · 2011
Cited alongside, same era.
Adaptive stochastic methods for sampling driven molecular systems
A. Jones and B. Leimkuhler · 2011
Cited alongside, same era.
Bayesian learning via stochastic gradient Langevin dynamics
M. Welling and Y.W. Teh · 2011
Cited alongside, same era.
Machine learning: A probabilistic perspective
K.P. Murphy · 2012
Cited alongside, same era.
Bayesian learning for neural networks , volume 118
R.M. Neal · 2012
Cited alongside, same era.
Empirical evaluation of rectified activations in convolutional network
B. Xu, N. Wang, T. Chen, and M. Li · 2015
Later among the works it cites.
An empirical analysis of deep network loss surfaces
D.J. Im, M. Tao, and K. Branson · 2016
Later among the works it cites.
Adaptive thermostats for noisy gradient systems
B. Leimkuhler and X. Shang · 2016
Later among the works it cites.
Energy landscapes for machine learning
A.J. Ballard, R. Das, S. Martiniani, D. Mehta, L. Sagun, J.D. Stevenson, and D.J. Wales · 2017
Later among the works it cites.
Non-asymptotic convergence analysis for the unadjusted Langevin algorithm
A. Durmus and E. Moulines · 2017
Later among the works it cites.
Three factors influencing minima in sgd
S. Jastrzȩbski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A.J. Storkey · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lecture 6.5 - RMSprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Adadelta: An adaptive learning rate method
M. Zeiler · 2012
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. Dauphin, R. Pascanu, C. Gülçehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Cited alongside, same era.
Bayesian sampling using stochastic gradient thermostats
N. Ding, Y. Fang, R. Babbush, C. Chen, R.D. Skeel, and H. Neven · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. Arous, and Y. LeCun · 2015
Cited alongside, same era.
Qualitatively characterizing neural network optimization problems
I.J. Goodfellow, O. Vinyals, and A.M. Saxe · 2015
Cited alongside, same era.
Later among the works it cites.
Automatic differentiation in PyTorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Later among the works it cites.
Langevin dynamics with variable coefficients and nonconservative forces: from stationary states to numerical methods
M. Sachs, B. Leimkuhler, and V. Danos · 2017
Later among the works it cites.
Quantum-chemical insights from deep tensor neural networks
K.T. Schütt, F. Arbabzadah, K.R. Müller S. Chmiela, and A. Tkatchenko · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
A.C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Later among the works it cites.
Improving palliative care with deep learning
A. Avati, K. Jung, S. Harman S, L. Downing, A. Ng, and N. Shah · 2018
Later among the works it cites.
The promises and pitfalls of stochastic gradient langevin dynamics
N. Brosse, A. Durmus, and E. Moulines · 2018
Later among the works it cites.
Exponential relaxation of the Nosé-Hoover equation under Brownian heating
D.P. Herzog · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, , M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis · 2018
Later among the works it cites.
Understanding generalization through visualizations
W.R. Huang, Z. Emam, M. Goldblum, L. Fowl, J.K. Terry, F. Huang, and T. Goldstein · 2019
Closest in time.
Lca: Loss change allocation for neural network training
J. Lan, R. Liu, H. Zhou, and J. Yosinski · 2019
Closest in time.
Hypocoercivity properties of adaptive langevin dynamics
B. Leimkuhler, M. Sachs, and G. Stoltz · 2019
Closest in time.