Fetching the paper…
Reading the bibliography…
While first-order optimization methods such as stochastic gradient descent (SGD) are popular in machine learning (ML), they come with well-known deficiencies, including relatively-slow convergence, sensitivity to the settings of hyper-parameters such as learning rate, stagnation at high training errors, and difficulty in escaping flat regions and saddle points.
“The modification of Newton’s method for unconstrained optimization by bounding cubic terms”
Andreas Griewank · 1981
Earlier work this paper cites.
“Newton’s method with a model trust region modification”
Danny Sorensen · 1982
Earlier work this paper cites.
“The conjugate gradient method and trust regions in large scale optimization”
Trond Steihaug · 1983
Earlier work this paper cites.
“On the limited memory BFGS method for large scale optimization”, 1989, pp. 503–528
Dong. Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
“Fast exact multiplication by the Hessian”
Barak Pearlmutter · 1994
Earlier work this paper cites.
“Regression shrinkage and selection via the Lasso”
Robert Tibshirani · 1996
Earlier work this paper cites.
“Trust region methods”
Andrew Conn, Nicholas Gould and Philippe Toint · 2000
Earlier work this paper cites.
“The Elements of Statistical Learning”
Jerome Friedman, Trevor Hastie and Robert Tibshirani · 2001
Earlier work this paper cites.
“GALAHAD, a library of thread-safe Fortran 90 packages for large-scale nonlinear optimization”
Nicholas Gould, Dominique Orban and Philippe Toint · 2003
Earlier work this paper cites.
“Cubic regularization of Newton method and its global performance”
Yurii Nesterov and Boris Polyak · 2006
Earlier work this paper cites.
“Numerical optimization”
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
“Learning multiple layers of features from tiny images” Data available at https://www.cs.toronto.edu/~kriz/cifar.html
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
“Deep learning via Hessian-free optimization”
James Martens · 2010
Earlier work this paper cites.
“Adaptive cubic regularisation methods for unconstrained optimization. Part I: motivation, convergence and numerical results”
Coralia Cartis, Nicholas Gould and Philippe Toint · 2011
Earlier work this paper cites.
“Adaptive cubic regularisation methods for unconstrained optimization. Part II: worst-case function-and derivative-evaluation complexity”
Coralia Cartis, Nicholas Gould and Philippe Toint · 2011
Cited alongside, same era.
“LIBSVM: A library for support vector machines” Software available at http://www.csie.ntu.edu.tw/~cjlin/libsvm
Chih-Chung Chang and Chih-Jen Lin · 2011
Cited alongside, same era.
“Adaptive and stochastic algorithms for EIT and DC resistivity problems with piecewise constant solutions and many measurements”
Kees van Doel and Uri Ascher · 2012
Cited alongside, same era.
“Metric learning: a survey”
Brian Kulis · 2012
Cited alongside, same era.
“Efficient backprop”
Yann LeCun, L“’eon Bottou, Genevieve Orr and Klaus-Robert M“”uller · 2012
Cited alongside, same era.
“Training deep and recurrent networks with Hessian-free optimization”
“Escaping From Saddle Points-Online Stochastic Gradient for Tensor Decomposition.”
Rong Ge, Furong Huang, Chi Jin and Yang Yuan · 2015
Later among the works it cites.
“Deep learning”
Yann LeCun, Yoshua Bengio and Geoffrey Hinton · 2015
Later among the works it cites.
“Assessing stochastic algorithms for large scale nonlinear least squares problems using extremal probabilities of linear combinations of gamma random variables”
Farbod Roosta-Khorasani, G“’abor. Sz“’ekely and Uri Ascher · 2015
Later among the works it cites.
“Exact and Inexact Subsampled Newton Methods for Optimization”
Raghu Bollapragada, Richard Byrd and Jorge Nocedal · 2016
Later among the works it cites.
“Optimization methods for large-scale machine learning”
L“’eon Bottou, Frank Curtis and Jorge Nocedal · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
James Martens and Ilya Sutskever · 2012
Cited alongside, same era.
“Optimization for machine learning”
Suvrit Sra, Sebastian Nowozin and Stephen Wright · 2012
Cited alongside, same era.
“Exact solutions to the nonlinear dynamics of learning in deep linear neural networks”
Andrew Saxe, James McClelland and Surya Ganguli · 2013
Cited alongside, same era.
“On the importance of initialization and momentum in deep learning”
Ilya Sutskever, James Martens, George Dahl and Geoffrey Hinton · 2013
Cited alongside, same era.
“Identifying and attacking the saddle point problem in high-dimensional non-convex optimization”
Yann Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli and Yoshua Bengio · 2014
Cited alongside, same era.
“Data completion and stochastic algorithms for PDE inversion problems with many measurements”
Farbod Roosta-Khorasani, Kees van Doel and Uri Ascher · 2014
Cited alongside, same era.
“Stochastic algorithms for inverse problems involving PDEs and many measurements”
Farbod Roosta-Khorasani, Kees van Doel and Uri Ascher · 2014
Cited alongside, same era.
“Deep learning”
Ian Goodfellow, Yoshua Bengio and Aaron Courville · 2016
Later among the works it cites.
“Deep residual learning for image recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Later among the works it cites.
“The Power of Normalization: Faster Evasion of Saddle Points”
Kfir Levy · 2016
Later among the works it cites.
“Sub-sampled Newton methods I: globally convergent algorithms”
Farbod Roosta-Khorasani and Michael. Mahoney · 2016
Later among the works it cites.
“Sub-sampled Newton methods II: local convergence rates”
Farbod Roosta-Khorasani and Michael Mahoney · 2016
Later among the works it cites.
“Sub-sampled newton methods with non-uniform sampling”
Peng Xu, Jiyan Yang, Farbod Roosta-Khorasani, Christopher R“’e and Michael Mahoney · 2016
Later among the works it cites.
“An Investigation of Newton-Sketch and Subsampled Newton Methods”
Albert Berahas, Raghu Bollapragada and Jorge Nocedal · 2017
Closest in time.
“How to Escape Saddle Points Efficiently”
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham Kakade and Michael Jordan · 2017
Closest in time.
“Newton-Type Methods for Non-Convex Optimization Under Inexact Hessian Information”
Peng Xu, Farbod Roosta-Khorasani and Michael. Mahoney · 2017
Closest in time.