Fetching the paper…
Reading the bibliography…
We propose a novel method for speeding up stochastic optimization algorithms via sketching methods, which recently became a powerful tool for accelerating algorithms for numerical linear algebra.
Improving the convergence of back-propagation learning with second order methods
Sue Becker and Yann Le Cun · 1988
Earlier work this paper cites.
Backpropagation: Past and future
Paul J Werbos · 1988
Earlier work this paper cites.
Exact calculation of the product of the hessian matrix of feed-forward network error functions and a vector in 0 (n) time
Martin F Møller · 1993
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Partial bfgs update and efficient step-length calculation for three-layer neural networks
Kazumi Saito and Ryohei Nakano · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
The mnist database of handwritten digits, 1998
Yann LeCun and Corinna Cortes · 1998
Earlier work this paper cites.
Adaptive method of realizing natural gradient learning for multilayer perceptrons
Shun-Ichi Amari, Hyeyoung Park, and Kenji Fukumizu · 2000
Cited alongside, same era.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Cited alongside, same era.
Introductory lectures on convex optimization
Yurii Nesterov · 2004
Cited alongside, same era.
Improved approximation algorithms for large matrices via random projections
Tamas Sarlos · 2006
Cited alongside, same era.
A stochastic quasi-newton method for online convex optimization
Nicol Schraudolph, Jin Yu, and Simon Günter · 2007
Cited alongside, same era.
Topmoumoute online natural gradient algorithm
Nicolas L Roux, Pierre-Antoine Manzagol, and Yoshua Bengio · 2008
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Later among the works it cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Later among the works it cites.
Krylov subspace descent for deep learning
Oriol Vinyals and Daniel Povey · 2011
Later among the works it cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Later among the works it cites.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sgd-qn: Careful quasi-newton stochastic gradient descent
Antoine Bordes, Léon Bottou, and Patrick Gallinari · 2009
Cited alongside, same era.
Deep learning via hessian-free optimization
James Martens · 2010
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Later among the works it cites.
Sketching as a tool for numerical linear algebra
David P Woodruff · 2014
Later among the works it cites.