Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) is almost ubiquitously used for training non-convex optimization tasks.
Pólya, G. (1945) Remarks on computing the probability integral in one and two dimensions. In Proceedings of the 1st Berkeley Symposium on Mathematical Statistics and Probability
1945
Earlier work this paper cites.
LeCun, Y., Bottou, L., Orr, G.B. & Müller, K.R. (1998) Efficient backprop. In Neural networks: Tricks of the trade
1998
Earlier work this paper cites.
Anton, A., Markowich, P., Toscani, G. & Unterreiter, A. (2001) On convex Sobolev inequalities and the rate of convergence to equilibrium for Fokker-Planck type equations. Communication in Partial Differential Equations
2001
Earlier work this paper cites.
Bovier, A., Eckhoff, M., Gayrard, V. & Klein, M. (2004) Metastability in reversible diffusion processes I: Sharp asymptotics for capacities and exit times. In Journal of the European Mathematical Society
2004
Earlier work this paper cites.
Bovier, A. Gayrard, V. & Klein, M. (2004) Metastability in reversible diffusion processes II: Precise asymptotics for small eigenvalues. In Journal of the European Mathematical Society
2004
Earlier work this paper cites.
Kolpas, A., Moehlis, J. & Kevrekidis, I.G. (2007) Coarse-grained analysis of stochasticity-induced switching between collective motion states. Proceedings of the National Academy of Sciences
2007
Earlier work this paper cites.
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Senior, A., Tucker, P., Yang, K., Le, Q.V. & Ng, A.Y. (2012) Large scale distributed deep networks. In Advances in Neural Information Processing Systems
2012
Earlier work this paper cites.
Berglund, N. (2013) Kramers’ law: Validity, derivations and generalisations. In Markov Processes Relat. Fields
2013
Earlier work this paper cites.
Dauphin, Y.N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S. & Bengio, Y. (2014) Identifying and attacking the saddle point problem in high-dimensional non-convex optimization. In Advances in Neural Information Processing Systems
2014
Earlier work this paper cites.
Pavliotis, G.A. (2014) Stochastic processes and applications: diffusion processes, the Fokker-Planck and Langevin equations. Springer
2014
Cited alongside, same era.
Pavliotis, G.A. (2014) Stochastic processes and applications: Diffusion processes, the Fokker-Planck and Langevin equations
2014
Cited alongside, same era.
Choromanska, A., Henaff, M., Mathieu, M., Arous, G.B. & LeCun, Y. (2015) The loss surfaces of multilayer networks. In Res Math Sci
2015
Cited alongside, same era.
Ge, R., Huang, R., Jin, C. & Yuan, Y. (2015) Escaping from saddle points–online stochastic gradient for tensor decomposition. In Conference on Learning Theory
2015
Cited alongside, same era.
Chaudhari, P., Oberman, A., Osher, S., Soatto, S. & Carlier, G. (2017) Deep Relaxation: partial differential equations for optimizing deep neural networks. In International Conference on Learning Representations
2017
Later among the works it cites.
Keskar, N.S., Mudigere, D., Nocedal, J., Smelyanskiy, M. & Tang, P.T.P. (2017) On large-batch training for deep learning: Generalization gap and sharp minima. In International Conference on Learning Representations
2017
Later among the works it cites.
2017
Later among the works it cites.
Mandt, S., Hoffman, M.D. & Blei, D.M. (2017) Stochastic gradient descent as approximate bayesian inference In Journal of Machine Learning Research
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Dinh, L., Pascanu, R., Bengio, S. & Bengio, Y. (2017) Sharp minima can generalize for deep nets. In International Conference on Machine Learning
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Hoffer, E., Hubara, I. & Soudry, D. (2017) Train longer, generalize better: closing the generalization gap in large batch training of neural networks. In Advances in Neural Information Processing Systems
2017
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.