Fetching the paper…
Reading the bibliography…
In several experimental reports on nonconvex optimization problems in machine learning, stochastic gradient descent (SGD) was observed to prefer minimizers with flat basins in comparison to more deterministic methods, yet there is very little rigorous understanding of this phenomenon.
M. Styblinski and T.S. Tang, Experiments in nonconvex optimization: stochastic approximation with function smoothing and simulated annealing , Neural Networks 3 (1990), pp. 467–483
1990
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, Simplifying neural nets by discovering flat minima , in Advances in neural information processing systems . 1995, pp. 529–536
1995
Earlier work this paper cites.
D.P. Bertsekas, Nonlinear programming , Athena scientific Belmont, 1999
1999
Earlier work this paper cites.
R. Ge, F. Huang, C. Jin, and Y. Yuan, Escaping from saddle points—online stochastic gradient for tensor decomposition , in Conference on Learning Theory . 2015, pp. 797–842
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
V. Patel, Kalman-based stochastic gradient method with stop condition and insensitivity to conditioning , SIAM Journal on Optimization 26 (2016), pp. 2620–2648
2016
Cited alongside, same era.
P. Balaprakash, Private Communication (2017)
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2018
Closest in time.