Fetching the paper…
Reading the bibliography…
We formulate the problem of neural network optimization as Bayesian filtering, where the observations are the backpropagated gradients.
Decoupled extended kalman filter training of feedforward layered networks
Puskorius, G. V. and Feldkamp, L. A · 1991
Earlier work this paper cites.
Optimal filtering algorithms for fast learning in feedforward neual networks
Sha, S., Palmieri, F., and Datum, M · 1992
Earlier work this paper cites.
Neurocontrol of nonlinear dynamical systems with kalman filter trained recurrent networks
Puskorius, G. V. and Feldkamp, L. A · 1994
Earlier work this paper cites.
How to train neural networks
Neuneier, R. and Zimmermann, H. G · 1998
Earlier work this paper cites.
Parameter-based kalman filter training: theory and implementation
Puskorius, G. V. and Feldkamp, L. A · 2001
Earlier work this paper cites.
Simple and conditioned adaptive behavior from kalman filter trained recurrent networks
Feldkamp, L. A., Prokhorov, D. V., and Feldkamp, T. M · 2003
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, A · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Cited alongside, same era.
A kronecker-factored approximate fisher matrix for convolution layers
Grosse, R. and Martens, J · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Cited alongside, same era.
Fixing weight decay regularization in adam
Loshchilov, I. and Hutter, F · 2017
Later among the works it cites.
Online natural gradient as a kalman filter
Ollivier, Y · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Later among the works it cites.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K · 2017
Later among the works it cites.
Noisy natural gradient as variational inference
Zhang, G., Sun, S., Duvenaud, D., and Grosse, R · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving generalization performance by switching from adam to sgd
Keskar, N. S. and Socher, R · 2017
Cited alongside, same era.
Khan, M. E. and Lin, W · 2017
Cited alongside, same era.
Vprop: Variational inference using rmsprop
Khan, M. E., Liu, Z., Tangkaratt, V., and Gal, Y · 2017
Cited alongside, same era.
Khan, M. E., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A · 2018
Closest in time.
On the convergence of adam and beyond
Reddi, S. J., Kale, S., and Kumar, S · 2018
Closest in time.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L., Xiong, Y., Liu, Y., and Sun, X · 2019
Closest in time.