Fetching the paper…
Reading the bibliography…
A stochastic second-order trust region method is proposed, which can be viewed as a second-order extension of the trust-region-ish (TRish) algorithm proposed by Curtis et al.
Fast convergence of natural gradient descent for overparameterized neural networks
Guodong Zhang, James Martens, and Roger Grosse · 1911
Earlier work this paper cites.
A Stochastic Approximation Method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
On a stochastic approximation method
K. L. Chung · 1954
Earlier work this paper cites.
A convergence theorem for nonnegative almost supermartingales and some applications
H. Robbins and D. Siegmund · 1971
Earlier work this paper cites.
A stochastic Newton-Raphson method
Dan Anbar · 1978
Earlier work this paper cites.
The conjugate gradient method and trust regions in large scale optimization
Trond Steihaug · 1983
Earlier work this paper cites.
Trust Region Methods
A. R. Conn, N. I. M. Gould, and Ph. L. Toint · 2000
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Yu. Nesterov · 2004
Earlier work this paper cites.
Numerical Optimization
J. Nocedal and S. J. Wright · 2006
Earlier work this paper cites.
A stochastic quasi-Newton method for online convex optimization
Nicol N Schraudolph, Jin Yu, and Simon Günter · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Robust Stochastic Approximation Approach to Stochastic Programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Sample Size Selection in Optimization Methods for Machine Learning
R. H. Byrd, G. M. Chin, J. Nocedal, and Y. Wu · 2012
Cited alongside, same era.
Hybrid deterministic-stochastic methods for data-fitting
M. P. Friedlander and M. Schmidt · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Cited alongside, same era.
Lecture 6.5. RMSPROP: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
A Kronecker-Factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Later among the works it cites.
Stochastic Quasi-Newton Methods for Nonconvex Stochastic Optimization
X. Wang, S. Ma, D. Goldfarb, and W. Liu · 2017
Later among the works it cites.
Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Later among the works it cites.
Exact and inexact subsampled Newton methods for optimization
Raghu Bollapragada, Richard H Byrd, and Jorge Nocedal · 2018
Later among the works it cites.
Optimization Methods for Large-Scale Machine Learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Later among the works it cites.
Global convergence rate analysis of unconstrained optimization methods based on probabilistic models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
A lower bound for the optimization of finite sums
A. Agarwal and L. Bottou · 2015
Cited alongside, same era.
Stochastic block mirror descent methods for nonsmooth and stochastic optimization
C. Dang and G. Lan · 2015
Cited alongside, same era.
Natural neural networks
Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
A stochastic quasi-Newton method for large-scale optimization
Richard H Byrd, Samantha L Hansen, Jorge Nocedal, and Yoram Singer · 2016
Cited alongside, same era.
Coralia Cartis and Katya Scheinberg · 2018
Later among the works it cites.
Stochastic optimization using a trust-region method and random models
Ruobing Chen, Matt Menickelly, and Katya Scheinberg · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel · 2018
Later among the works it cites.
On the convergence of adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Exploiting Negative Curvature in Deterministic and Stochastic Optimization
F. E. Curtis and D. P. Robinson · 2019
Closest in time.
A Stochastic Trust Region Algorithm Based on Careful Step Normalization
F. E. Curtis, K. Scheinberg, and R. Shi · 2019
Closest in time.
Probability: Theory and Examples
R. Durrett · 2019
Closest in time.