Fetching the paper…
Reading the bibliography…
In this paper, we consider stochastic second-order methods for minimizing a finite summation of nonconvex functions.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Local convergence analysis for partitioned quasi-newton updates
Andreas Griewank and Ph L Toint · 1982
Earlier work this paper cites.
Representations of quasi-newton matrices and their use in limited memory methods
Richard H Byrd, Jorge Nocedal, and Robert B Schnabel · 1994
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2001
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen J Wright · 2006
Earlier work this paper cites.
Optimization theory and methods: nonlinear programming
Wenyu Sun and Ya-Xiang Yuan · 2006
Earlier work this paper cites.
Deep learning via Hessian free optimization
James Martens · 2010
Earlier work this paper cites.
Weijun Zhou and Xiaojun Chen · 2010
Earlier work this paper cites.
On the use of stochastic Hessian information in optimization methods for machine learning
Richard H Byrd, Gillian M Chin, Will Neveitt, and Jorge Nocedal · 2011
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
The nature of statistical learning theory
Vladimir Vapnik · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Cited alongside, same era.
A multi-batch l-bfgs method for machine learning
Albert S Berahas, Jorge Nocedal, and Martin Takác · 2016
Cited alongside, same era.
A stochastic quasi-Newton method for large-scale optimization
R. H. Byrd, S. L. Hansen, Jorge Nocedal, and Y. Singer · 2016
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Stochastic block bfgs: Squeezing more curvature out of data
Robert Gower, Donald Goldfarb, and Peter Richtárik · 2016
Cited alongside, same era.
A Kronecker-factored approximate Fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Cited alongside, same era.
On the generalization ability of online gradient descent algorithm under the quadratic growth condition
Daqing Chang, Ming Lin, and Changshui Zhang · 2018
Later among the works it cites.
Uniform convergence of gradients for non-convex learning and optimization
Dylan J Foster, Ayush Sekhari, and Karthik Sridharan · 2018
Later among the works it cites.
A fully stochastic second-order trust region method
Frank E Curtis and Rui Shi · 2019
Later among the works it cites.
Structured quasi-newton methods for optimization with orthogonality constraints
Jiang Hu, Bo Jiang, Lin Lin, Zaiwen Wen, and Ya-xiang Yuan · 2019
Later among the works it cites.
Stochastic gradient descent for nonconvex learning without bounded gradient assumptions
Yunwen Lei, Ting Hu, Guiying Li, and Ke Tang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
Practical Gauss-Newton optimisation for deep learning
Aleksandar Botev, Hippolyt Ritter, and David Barber · 2017
Cited alongside, same era.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Lam M Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Cited alongside, same era.
Newton sketch: A near linear-time optimization algorithm with linear-quadratic convergence
Mert Pilanci and Martin J Wainwright · 2017
Cited alongside, same era.
Stochastic Quasi-Newton Methods for Nonconvex Stochastic Optimization
Xiao Wang, Shiqian Ma, Donald Goldfarb, and Wei Liu · 2017
Cited alongside, same era.
A stochastic semismooth newton method for nonsmooth nonconvex optimization
Andre Milzarek, Xiantao Xiao, Shicong Cen, Zaiwen Wen, and Michael Ulbrich · 2019
Later among the works it cites.
Large-scale distributed second-order optimization using kronecker-factored approximate curvature for deep convolutional neural networks
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2019
Later among the works it cites.
Sub-sampled newton methods
Fred Roosta and Michael W. Mahoney · 2019
Later among the works it cites.
Newton-type methods for non-convex optimization under inexact Hessian information
Peng Xu, Fred Roosta, and Michael W. Mahoney · 2019
Later among the works it cites.
A stochastic extra-step quasi-newton method for nonsmooth nonconvex optimization
Minghan Yang, Andre Milzarek, Zaiwen Wen, and Tong Zhang · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Closest in time.
Practical quasi-newton methods for training deep neural networks
Donald Goldfarb, Yi Ren, and Achraf Bahamou · 2020
Closest in time.
Newton methods for convolutional neural networks
Chien-Chih Wang, Kent Loong Tan, and Chih-Jen Lin · 2020
Closest in time.
Sketchy empirical natural gradient methods for deep learning
Minghan Yang, Dong Xu, Yongfeng Li, Zaiwen Wen, and Mengyun Chen · 2020
Closest in time.