Fetching the paper…
Reading the bibliography…
Recently, we proposed to transform the outputs of each hidden neuron in a multi-layer perceptron network to have zero output and zero slope on average, and use separate shortcut connections to model the linear dependencies instead.
Natural gradient works efficiently in learning
S. Amari · 1998
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 1998
Earlier work this paper cites.
Accelerated gradient descent by factor-centering decomposition
N. N. Schraudolph · 1998
Earlier work this paper cites.
Centering neural network gradient factors
N. N. Schraudolph · 1998
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Cited alongside, same era.
Topmoumoute online natural gradient algorithm
N. Le Roux, P. A. Manzagol, and Y. Bengio · 2008
Cited alongside, same era.
Deep big simple neural nets excel on handwritten digit recognition
D. C. Ciresan, U. Meier, L. M. Gambardella, and J. Schmidhuber · 2010
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Cited alongside, same era.
Deep learning via Hessian-free optimization
J. Martens · 2010
Later among the works it cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2012
Later among the works it cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Later among the works it cites.
Deep learning made easier by linear transformations in perceptrons
Tapani Raiko, Harri Valpola, and Yann LeCun · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…