Fetching the paper…
Reading the bibliography…
Standard gradient descent methods are susceptible to a range of issues that can impede training, such as high correlations and different scaling in parameter space.These difficulties can be addressed by second-order approaches that apply a pre-conditioning matrix to the gradient to improve convergence.
Possible principles underlying the transformations of sensory messages
H. Barlow · 1961
Earlier work this paper cites.
Gradient methods for the minimisation of functionals
B. Polyak · 1963
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. Polyak · 1964
Earlier work this paper cites.
Unsupervised learning
H. Barlow · 1989
Earlier work this paper cites.
What does the retina know about natural scenes?
J. J. Atick and A. N. Redlich · 1992
Earlier work this paper cites.
A theory of maximizing sensory information
J. H. Van Hateren · 1992
Earlier work this paper cites.
Towards a discrete newton method with memory for large-scale optimization
R. H. Byrd, J. Nocedal, and C. Zhu · 1996
Earlier work this paper cites.
Parameter adaptation in stochastic optimization
L. B. Almeida, T. Langlois, J. D. Amaral, and A. Plakhov · 1998
Earlier work this paper cites.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Y. LeCun and C. Cortes · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Striving for Simplicity: The All Convolutional Net
J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng · 2015
Cited alongside, same era.
Bipedalwalkerhardcore-v2
O. Klimov · 2016
Later among the works it cites.
Random synaptic feedback weights support error backpropagation for deep learning
T. P. Lillicrap, D. Cownden, D. B. Tweed, and C. J. Akerman · 2016
Later among the works it cites.
Second-Order Optimization for Neural Networks
J. Martens · 2016
Later among the works it cites.
Online learning rate adaptation with hypergradient descent
A. G. Baydin, R. Cornish, D. Martínez-Rubio, M. Schmidt, and F. D. Wood · 2017
Later among the works it cites.
Openai baselines
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y. Wu, and P. Zhokhov · 2017
Later among the works it cites.
Evolving stable strategies
D. Ha · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Desjardins, K. Simonyan, R. Pascanu, and K. Kavukcuoglu · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Preconditioned Stochastic Gradient Descent
X.-L. Li · 2015
Cited alongside, same era.
Gradient-based Hyperparameter Optimization through Reversible Learning
D. Maclaurin, D. Duvenaud, and R. P. Adams · 2015
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
J. Martens and R. B. Grosse · 2015
Cited alongside, same era.
Optimization Methods for Large-Scale Machine Learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
A Kronecker-factored approximate Fisher matrix for convolution layers
R. Grosse and J. Martens · 2016
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Reinforcement learning for improving agent design
D. Ha · 2018
Later among the works it cites.
Decorrelated Batch Normalization
L. Huang, D. Yang, B. Lang, and J. Deng · 2018
Later among the works it cites.
Recurrent deterministic policy gradient method for bipedal locomotion on rough terrain challenge
D. R. Song, C. Yang, C. McGreavy, and Z. Li · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
R. Wang, J. Lehman, J. Clune, and K. O. Stanley · 2019
Closest in time.