Fetching the paper…
Reading the bibliography…
We investigate the use of regularized Newton methods with adaptive norms for optimizing neural networks.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
The conjugate gradient method and trust regions in large scale optimization
Trond Steihaug · 1983
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Sue Becker, Yann Le Cun, et al · 1988
Earlier work this paper cites.
Characteristic vectors of bordered matrices with infinite dimensions i
Eugene P Wigner · 1993
Earlier work this paper cites.
Training feedforward networks with the marquardt algorithm
Martin T Hagan and Mohammad B Menhaj · 1994
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Solving the ill-conditioning in neural network learning
Patrick Van Der Smagt and Gerd Hirzinger · 1998
Earlier work this paper cites.
Trust region methods
Andrew R Conn, Nicholas IM Gould, and Philippe L Toint · 2000
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Earlier work this paper cites.
Numerical optimization, 2nd Edition
Jorge Nocedal and Stephen J Wright · 2006
Earlier work this paper cites.
An estimator for the diagonal of a matrix
Costas Bekas, Effrosyni Kokiopoulou, and Yousef Saad · 2007
Earlier work this paper cites.
Second-order stagewise backpropagation for hessian-matrix analyses and investigation of negative curvature
Eiji Mizutani and Stuart E Dreyfus · 2008
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Leon Bottou · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Adaptive cubic regularisation methods for unconstrained optimization. part i: motivation, convergence and numerical results
Coralia Cartis, Nicholas IM Gould, and Philippe L Toint · 2011
Earlier work this paper cites.
Improved preconditioner for hessian free optimization
Olivier Chapelle and Dumitru Erhan · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Complexity bounds for second-order optimality in unconstrained optimization
Coralia Cartis, Nicholas IM Gould, and Ph L Toint · 2012
Earlier work this paper cites.
How Much Patience to You Have?: A Worst-case Perspective on Smooth Noncovex Optimization
Coralia Cartis, Nicholas IM Gould, and Philippe L Toint · 2012
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2013
Cited alongside, same era.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
Sub-sampled cubic regularization for non-convex optimization
Jonas Moritz Kohler and Aurelien Lucchi · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Luca Desmaison, Alban aComplexity bounds for second-order optimality in unconstrained optimizationnd Antiga, and Adam Lerer · 2017
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2017
Later among the works it cites.
Failures of gradient-based deep learning
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
Escaping from saddle points-online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Convergence rate analysis of a stochastic trust region method for nonconvex optimization
Jose Blanchet, Coralia Cartis, Matt Menickelly, and Katya Scheinberg · 2016
Cited alongside, same era.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Cited alongside, same era.
Nilesh Tripuraneni, Mitchell Stern, Chi Jin, Jeffrey Regier, and Michael I Jordan · 2017
Later among the works it cites.
Second-order optimization for non-convex machine learning: An empirical study
Peng Xu, Farbod Roosta-Khorasan, and Michael W Mahoney · 2017
Later among the works it cites.
Newton-type methods for non-convex optimization under inexact hessian information
Peng Xu, Farbod Roosta-Khorasani, and Michael W Mahoney · 2017
Later among the works it cites.
The case for full-matrix adaptive regularization
Naman Agarwal, Brian Bullins, Xinyi Chen, Elad Hazan, Karan Singh, Cyril Zhang, and Yi Zhang · 2018
Later among the works it cites.
Negative eigenvalues of the hessian in deep neural networks
Guillaume Alain, Nicolas Le Roux, and Pierre-Antoine Manzagol · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
Stochastic optimization using a trust-region method and random models
Ruobing Chen, Matt Menickelly, and Katya Scheinberg · 2018
Later among the works it cites.
Escaping saddles with stochastic gradients
Hadi Daneshmand, Jonas Kohler, Aurelien Lucchi, and Thomas Hofmann · 2018
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
Simon S Du and Jason D Lee · 2018
Later among the works it cites.
A distributed second-order algorithm you can trust
Celestine Dünner, Aurelien Lucchi, Matilde Gargiani, An Bian, Thomas Hofmann, and Martin Jaggi · 2018
Later among the works it cites.
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2018
Later among the works it cites.
Stochastic second-order methods for non-convex optimization with inexact hessian and gradient
Liu Liu, Xuanqing Liu, Cho-Jui Hsieh, and Dacheng Tao · 2018
Later among the works it cites.
Revisiting small batch training for deep neural networks
Dominic Masters and Carlo Luschi · 2018
Later among the works it cites.
Second-order optimization method for large mini-batch: Training resnet-50 on imagenet in 35 epochs
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2018
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2018
Later among the works it cites.
Inexact non-convex newton-type methods
Zhewei Yao, Peng Xu, Farbod Roosta-Khorasani, and Michael W Mahoney · 2018
Later among the works it cites.
Inefficiency of k-fac for large batch size training
Linjian Ma, Gabe Montague, Jiayu Ye, Zhewei Yao, Amir Gholami, Kurt Keutzer, and Michael W Mahoney · 2019
Closest in time.