Fetching the paper…
Reading the bibliography…
We focus on minimizing nonconvex finite-sum functions that typically arise in machine learning problems.
The modification of newton’s method for unconstrained optimization by bounding cubic terms
Griewank, A · 1981
Earlier work this paper cites.
Simplified neuron model as a principal component analyzer
Oja, E · 1982
Earlier work this paper cites.
Estimating the largest eigenvalue by the power and lanczos algorithms with a random start
Kuczyński, J. and Woźniakowski, H · 1992
Earlier work this paper cites.
Fast exact multiplication by the hessian
Pearlmutter, B. A · 1994
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Nesterov, Y. and Polyak, B. T · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
Martens, J · 2010
Earlier work this paper cites.
Libsvm: a library for support vector machines
Chang, C.-C. and Lin, C.-J · 2011
Earlier work this paper cites.
Recovering low-rank matrices from few coefficients in any basis
Gross, D · 2011
Cited alongside, same era.
Learning recurrent neural networks with hessian-free optimization
Martens, J. and Sutskever, I · 2011
Cited alongside, same era.
Krylov subspace descent for deep learning
Vinyals, O. and Povey, D · 2012
Cited alongside, same era.
A fast divide-and-conquer algorithm for computing the spectra of real symmetric tridiagonal matrices
Coakley, E. S. and Rokhlin, V · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Cited alongside, same era.
A stochastic quasi-newton method for large-scale optimization
Byrd, R. H., Hansen, S. L., Nocedal, J., and Singer, Y · 2016
Cited alongside, same era.
Exploiting negative curvature in deterministic and stochastic optimization
Curtis, F. E. and Robinson, D. P · 2017
Later among the works it cites.
Sub-sampled cubic regularization for non-convex optimization
Kohler, J. M. and Lucchi, A · 2017
Later among the works it cites.
A generic approach for escaping saddle points
Reddi, S. J., Zaheer, M., Sra, S., Poczos, B., Bach, F., Salakhutdinov, R., and Smola, A. J · 2017
Later among the works it cites.
A line-search algorithm inspired by the adaptive cubic regularization framework and complexity analysis
Bergou, E. H., Diouane, Y., and Gratton, S · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient descent efficiently finds the cubic-regularized non-convex newton step
Carmon, Y. and Duchi, J. C · 2016
Cited alongside, same era.
Gradient descent converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Cited alongside, same era.
Using improved directions of negative curvature for the solution of bound-constrained nonconvex problems
Cano, J., Moguerza, J. M., and Prieto, F. J · 2017
Cited alongside, same era.
Finding approximate local minima faster than gradient descent
Agarwal, N., Allen-Zhu, Z., Bullins, B., Hazan, E., and Ma, T
Cited in the paper.
Second-order stochastic optimization for machine learning in linear time
Agarwal, N., Bullins, B., and Hazan, E
Cited in the paper.
Adaptive cubic regularisation methods for unconstrained optimization. part i: motivation, convergence and numerical results
Cartis, C., Gould, N. I., and Toint, P. L
Cited in the paper.
Later among the works it cites.
Accelerated methods for nonconvex optimization
Carmon, Y., Duchi, J. C., Hinder, O., and Sidford, A · 2018
Later among the works it cites.
Adaptive negative curvature descent with applications in non-convex optimization
Liu, M., Li, Z., Wang, X., Yi, J., and Yang, T · 2018
Later among the works it cites.
Cubic regularization with momentum for nonconvex optimization
Wang, Z., Zhou, Y., Liang, Y., and Lan, G · 2018
Later among the works it cites.
First-order stochastic algorithms for escaping from saddle points in almost linear time
Xu, Y., Rong, J., and Yang, T · 2018
Later among the works it cites.