Fetching the paper…
Reading the bibliography…
We analyze the variance of stochastic gradients along negative curvature directions in certain non-convex machine learning models and show that stochastic gradients exhibit a strong component along these directions.
Fast exact multiplication by the hessian
Pearlmutter, B. A · 1994
Earlier work this paper cites.
Trust region methods
Conn, A. R., Gould, N. I., and Toint, P. L · 2000
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Nesterov, Y. and Polyak, B. T · 2006
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L · 2010
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E. and Bach, F. R · 2011
Earlier work this paper cites.
How Much Patience to You Have?: A Worst-case Perspective on Smooth Noncovex Optimization
Cartis, C., Gould, N. I., and Toint, P. L · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Most tensor problems are np-hard
Hillar, C. J. and Lim, L.-H · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Nesterov, Y · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2015
Cited alongside, same era.
Escaping from saddle points-online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Cited alongside, same era.
Learning halfspaces and neural networks with random initialization
Zhang, Y., Lee, J. D., Wainwright, M. J., and Jordan, M. I · 2015
Cited alongside, same era.
Gradient descent converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Cited alongside, same era.
Chaudhari, P. and Soatto, S · 2017
Later among the works it cites.
Exploiting negative curvature in deterministic and stochastic optimization
Curtis, F. E. and Robinson, D. P · 2017
Later among the works it cites.
Accurate, large minibatch sgd: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Later among the works it cites.
Sub-sampled cubic regularization for non-convex optimization
Kohler, J. M. and Lucchi, A · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Levy, K. Y · 2016
Cited alongside, same era.
Natasha 2: Faster non-convex optimization than sgd
Allen-Zhu, Z · 2017
Cited alongside, same era.
Neon2: Finding local minima via first-order oracles
Allen-Zhu, Z. and Li, Y · 2017
Cited alongside, same era.
Carmon, Y., Hinder, O., Duchi, J. C., and Sidford, A · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I
Cited in the paper.
Accelerated gradient descent escapes saddle points faster than gradient descent
Jin, C., Netrapalli, P., and Jordan, M. I
Cited in the paper.
Reddi, S. J., Zaheer, M., Sra, S., Poczos, B., Bach, F., Salakhutdinov, R., and Smola, A. J · 2017
Later among the works it cites.
Simchowitz, M., Alaoui, A. E., and Recht, B · 2017
Later among the works it cites.
Newton-type methods for non-convex optimization under inexact hessian information
Xu, P., Roosta-Khorasani, F., and Mahoney, M. W · 2017
Later among the works it cites.
First-order stochastic algorithms for escaping from saddle points in almost linear time
Xu, Y. and Yang, T · 2017
Later among the works it cites.
A hitting time analysis of stochastic gradient langevin dynamics
Zhang, Y., Liang, P., and Charikar, M · 2017
Later among the works it cites.