Fetching the paper…
Reading the bibliography…
We present an algorithm for minimizing a sum of functions that combines the computational efficiency of stochastic gradient descent (SGD) with the second order curvature information leveraged by quasi-Newton methods.
A stochastic approximation method
Robbins, Herbert and Monro, Sutton · 1951
Earlier work this paper cites.
The convergence of a class of double-rank minimization algorithms 2. The new algorithm
Broyden, CG · 1970
Earlier work this paper cites.
A new approach to variable metric algorithms
Fletcher, R · 1970
Earlier work this paper cites.
A family of variable-metric methods derived by variational means
Goldfarb, D · 1970
Earlier work this paper cites.
Conditioning of quasi-Newton methods for function minimization
Shanno, DF · 1970
Earlier work this paper cites.
Quasi-Newton methods, motivation and theory
Dennis Jr, John E and Moré, Jorge J · 1977
Earlier work this paper cites.
Robust statistics
Huber, PJ · 1981
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Liu, Dong C DC and Nocedal, Jorge · 1989
Earlier work this paper cites.
Stochastic gradient learning in neural networks
Bottou, Léon · 1991
Earlier work this paper cites.
An information-maximization approach to blind separation and blind deconvolution
Bell, AJ and Sejnowski, TJ · 1995
Earlier work this paper cites.
Natural Gradient Works Efficiently in Learning
Amari, Shun-Ichi · 1998
Earlier work this paper cites.
Local gain adaptation in stochastic gradient descent
Schraudolph, Nicol N · 1999
Earlier work this paper cites.
Convex optimization
Boyd, S P and Vandenberghe, L · 2004
Earlier work this paper cites.
A convergent incremental gradient method with a constant step size
Blatt, Doron, Hero, Alfred O, and Gauchman, Hillel · 2007
Cited alongside, same era.
A stochastic quasi-Newton method for online convex optimization
Schraudolph, Nicol, Yu, Jin, and Günter, Simon · 2007
Cited alongside, same era.
Trust region newton method for logistic regression
Lin, Chih-Jen, Weng, Ruby C, and Keerthi, S Sathiya · 2008
Cited alongside, same era.
SGD-QN: Careful quasi-Newton stochastic gradient descent
Bordes, Antoine, Bottou, Léon, and Gallinari, Patrick · 2009
Cited alongside, same era.
Historical Development of the BFGS Secant Method and Its Characterization Properties
Papakonstantinou, JM · 2009
Cited alongside, same era.
Contractive auto-encoders: Explicit invariance during feature extraction
Rifai, Salah, Vincent, Pascal, Muller, Xavier, Glorot, Xavier, and Bengio, Yoshua · 2011
Later among the works it cites.
Krylov subspace descent for deep learning
Vinyals, Oriol and Povey, Daniel · 2011
Later among the works it cites.
Efficient and optimal binary Hopfield associative memory storage using minimum probability flow
Hillar, Christopher, Sohl-Dickstein, Jascha, and Koepsell, Kilian · 2012
Later among the works it cites.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, Geoffrey E., Srivastava, Nitish, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan R · 2012
Later among the works it cites.
On the difficulty of training Recurrent Neural Networks
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2012
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sunehag, Peter, Trumpf, Jochen, Vishwanathan, S V N, and Schraudolph, Nicol · 2009
Cited alongside, same era.
Theano: a CPU and GPU math expression compiler
Bergstra, J and Breuleux, O · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2010
Cited alongside, same era.
Deep learning via Hessian-free optimization
Martens, James · 2010
Cited alongside, same era.
Non-Asymptotic Analysis of Stochastic Approximation Algorithms for Machine Learning
Bach, FR and Moulines, E · 2011
Cited alongside, same era.
On the use of stochastic hessian information in optimization methods for machine learning
Byrd, RH Richard H, Chin, GM Gillian M, Neveitt, Will, and Nocedal, Jorge · 2011
Cited alongside, same era.
On optimization methods for deep learning
Le, Quoc V., Ngiam, Jiquan, Coates, Adam, Lahiri, Abhik, Prochnow, Bobby, and Ng, Andrew Y · 2011
Cited alongside, same era.
Later among the works it cites.
A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets
Roux, N Le, Schmidt, M, and Bach, F · 2012
Later among the works it cites.
The Natural Gradient by Analogy to Signal Whitening, and Recipes and Tricks for its Use
Sohl-Dickstein, Jascha · 2012
Later among the works it cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O (1/n)
Bach, F and Moulines, E · 2013
Closest in time.
Fast probabilistic optimization from noisy gradients
Hennig, P · 2013
Closest in time.
Optimization with First-Order Surrogate Functions
Mairal, J · 2013
Closest in time.
A Stochastic Quasi-Newton Method for Large-Scale Optimization
Byrd, RH, Hansen, SL, Nocedal, J, and Singer, Y · 2014
Closest in time.
Incremental Majorization-Minimization Optimization with Application to Large-Scale Machine Learning
Mairal, Julien · 2014
Closest in time.