Fetching the paper…
Reading the bibliography…
To optimize a neural network one often thinks of optimizing its parameters, but it is ultimately a matter of optimizing the function that maps inputs to outputs.
Eigenvalues of covariance matrices: Application to neural-network learning
Yann LeCun, Ido Kanter, and Sara A Solla · 1991
Earlier work this paper cites.
A new learning algorithm for blind signal separation
Shun-ichi Amari, Andrzej Cichocki, and Howard Hua Yang · 1996
Earlier work this paper cites.
Tempering backpropagation networks: Not all weights are created equal
Nicol N Schraudolph and Terrence J Sejnowski · 1996
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
A bayesian approach to on-line learning
Manfred Opper and Ole Winther · 1998
Earlier work this paper cites.
Accelerated gradient descent by factor-centering decomposition
Nicol Schraudolph · 1998
Earlier work this paper cites.
Transformation invariance in pattern recognition—tangent distance and tangent propagation
Patrice Simard, Yann LeCun, John Denker, and Bernard Victorri · 1998
Earlier work this paper cites.
Adaptive method of realizing natural gradient learning for multilayer perceptrons
Shun-Ichi Amari, Hyeyoung Park, and Kenji Fukumizu · 2000
Earlier work this paper cites.
On-line variational bayesian learning
Antti Honkela and Harri Valpola · 2003
Earlier work this paper cites.
Motor adaptation to single force pulses: sensitive to direction but insensitive to within-movement pulse placement and magnitude
Michael S Fine and Kurt A Thoroughman · 2006
Earlier work this paper cites.
Adaptive regularization of weight vectors
Koby Crammer, Alex Kulesza, and Mark Dredze · 2009
Cited alongside, same era.
Deep learning made easier by linear transformations in perceptrons
Tapani Raiko, Harri Valpola, and Yann LeCun · 2012
Cited alongside, same era.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Variance reduction for stochastic gradient optimization
Chong Wang, Xi Chen, Alexander J Smola, and Eric P Xing · 2013
Cited alongside, same era.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size
Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer · 2016
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P Kingma · 2016
Later among the works it cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Later among the works it cites.
Gradient episodic memory for continual learning
David Lopez-Paz et al · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
James Martens · 2014
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Deep neural networks with random Gaussian weights: a universal classification strategy?
Raja Giryes, Guillermo Sapiro, and Alexander M Bronstein · 2016
Cited alongside, same era.
Gabriel Pereyra, George Tucker, Jan Chorowski, Łukasz Kaiser, and Geoffrey Hinton · 2017
Later among the works it cites.
icarl: Incremental classifier and representation learning
Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert · 2017
Later among the works it cites.
Continual learning through synaptic intelligence
Friedemann Zenke, Ben Poole, and Surya Ganguli · 2017
Later among the works it cites.
Online structured laplace approximations for overcoming catastrophic forgetting
Hippolyt Ritter, Aleksandar Botev, and David Barber · 2018
Closest in time.