Fetching the paper…
Reading the bibliography…
Second-order optimizers hold intriguing potential for deep learning, but suffer from increased cost and sensitivity to the non-convexity of the loss surface as compared to gradient-based approaches.
Minimization of functions having Lipschitz continuous first partial derivatives
L. Armijo · 1966
Earlier work this paper cites.
NIST special database 19 handprinted forms and characters database
P. J. Grother · 1995
Earlier work this paper cites.
Numerical methods for unconstrained optimization and nonlinear equations , volume 16
J. E. Dennis Jr and R. B. Schnabel · 1996
Earlier work this paper cites.
Penalized regressions: the bridge versus the lasso
W. J. Fu · 1998
Earlier work this paper cites.
A hybrid linear/nonlinear training algorithm for feedforward neural networks
S. McLoone, M. D. Brown, G. Irwin, and A. Lightbody · 1998
Earlier work this paper cites.
Real analysis: modern techniques and their applications , volume 40
G. B. Folland · 1999
Earlier work this paper cites.
A simple and efficient algorithm for gene selection using sparse logistic regression
S. K. Shevade and S. S. Keerthi · 2003
Earlier work this paper cites.
Convex optimization
S. Boyd, S. P. Boyd, and L. Vandenberghe · 2004
Earlier work this paper cites.
Variable projections neural network training
V. Pereyra, G. Scherer, and F. Wong · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Earlier work this paper cites.
Deep learning via Hessian-free optimization
J. Martens · 2010
Cited alongside, same era.
Sublinear optimization for machine learning
K. L. Clarkson, E. Hazan, and D. P. Woodruff · 2012
Cited alongside, same era.
The MNIST database of handwritten digit images for machine learning research [best of the web]
L. Deng · 2012
Cited alongside, same era.
Efficiency of coordinate descent methods on huge-scale optimization problems
Y. Nesterov · 2012
Cited alongside, same era.
Block coordinate descent algorithms for large-scale sparse multiclass classification
M. Blondel, K. Seki, and K. Uehara · 2013
Cited alongside, same era.
Convex optimization: Algorithms and complexity
S. Bubeck · 2014
Error bounds for approximations with deep ReLU networks
D. Yarotsky · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Later among the works it cites.
scikit-optimize/scikit-optimize: v0.5.2, Mar. 2018
T. Head, MechCoder, G. Louppe, I. Shcherbatyi, fcharras, Z. Vinícius, cmmalone, C. Schröder, nel215, N. Campos, T. Young, S. Cereda, T. Fan, rene rex, K. K. Shi, J. Schwabedal, carlosdanielcsantos, Hvass-Labs, M. Pak, SoManyUsernamesTaken, F. Callaway, L. Estève, L. Besson, M. Cherti, K. Pfannschmidt, F. Linzberger, C. Cauet, A. Gut, A. Mueller, and A. Fabisch · 2018
Later among the works it cites.
Robust training and initialization of deep neural networks: An adaptive basis viewpoint
E. C. Cyr, M. A. Gulian, R. G. Patel, M. Perego, and N. A. Trask · 2019
Later among the works it cites.
Nonlinear approximation and (deep) relu networks
I. Daubechies, R. DeVore, S. Foucart, B. Hanin, and G. Petrova · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Coordinate descent algorithms
S. J. Wright · 2015
Cited alongside, same era.
Practical Gauss-Newton optimisation for deep learning
A. Botev, H. Ritter, and D. Barber · 2017
Cited alongside, same era.
Stable architectures for deep neural networks
E. Haber and L. Ruthotto · 2017
Cited alongside, same era.
Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms
H. Xiao, K. Rasul, and R. Vollgraf · 2017
Cited alongside, same era.
Later among the works it cites.
Deep ReLU networks and high-order finite element methods
J. A. Opschoor, P. Petersen, and C. Schwab · 2019
Later among the works it cites.
Large-scale distributed second-order optimization using Kronecker-factored approximate curvature for deep convolutional neural networks
K. Osawa, Y. Tsuji, Y. Ueno, A. Naruse, R. Yokota, and S. Matsuoka · 2019
Later among the works it cites.
Newton-type methods for non-convex optimization under inexact hessian information
P. Xu, F. Roosta, and M. W. Mahoney · 2019
Later among the works it cites.
Least squares auto-tuning
S. T. Barratt and S. P. Boyd · 2020
Closest in time.
Scalable and practical natural gradient for large-scale deep learning
K. Osawa, Y. Tsuji, Y. Ueno, A. Naruse, C.-S. Foo, and R. Yokota · 2020
Closest in time.