Fetching the paper…
Reading the bibliography…
We present weight normalization: a reparameterization of the weight vectors in a neural network that decouples the length of those weight vectors from their direction.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Neural learning in structured parameter spaces - natural Riemannian gradient
S. Amari · 1997
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Rank, trace-norm and max-norm
N. Srebro and A. Shraibman · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
J. Martens · 2010
Earlier work this paper cites.
Deep learning made easier by linear transformations in perceptrons
T. Raiko, H. Valpola, and Y. LeCun · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Maxout networks
I. J. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio · 2013
Earlier work this paper cites.
Auto-Encoding Variational Bayes
D. P. Kingma and M. Welling · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Deeply-supervised nets
C.-Y. Lee, S. Xie, P. Gallagher, Z. Zhang, and Z. Tu · 2014
Cited alongside, same era.
Network in network
M. Lin, C. Qiang, and S. Yan · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Later among the works it cites.
Data-dependent initializations of convolutional neural networks
P. Krähenbühl, C. Doersch, J. Donahue, and T. Darrell · 2015
Later among the works it cites.
Optimizing neural networks with kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Later among the works it cites.
D. Mishkin and J. Matas · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Natural neural networks
G. Desjardins, K. Simonyan, R. Pascanu, et al · 2015
Cited alongside, same era.
Draw: A recurrent neural network for image generation
K. Gregor, I. Danihelka, A. Graves, and D. Wierstra · 2015
Cited alongside, same era.
Scaling up natural gradient by sparsely factorizing the inverse fisher matrix
R. Grosse and R. Salakhudinov · 2015
Cited alongside, same era.
Later among the works it cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Later among the works it cites.
Markov chain Monte Carlo and variational inference: Bridging the gap
T. Salimans, D. P. Kingma, and M. Welling · 2015
Later among the works it cites.
Striving for simplicity: The all convolutional net
J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller · 2015
Later among the works it cites.
Deep learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Closest in time.