Fetching the paper…
Reading the bibliography…
We introduce a novel framework for the estimation of the posterior distribution over the weights of a neural network, based on a new probabilistic interpretation of adaptive optimisation algorithms such as AdaGrad and Adam.
A Simple Baseline for Bayesian Uncertainty in Deep Learning
Maddox, W., Garipov, T., Izmailov, P., Vetrov, D., and Wilson, A. G · 1902
Earlier work this paper cites.
On The Likelihood That One Unknown Probability Exceeds Another in View of the Evidence of Two Samples
Thompson, W. R · 1933
Earlier work this paper cites.
A Practical Bayesian Framework for Backpropagation Networks
MacKay, D. J. C · 1992
Earlier work this paper cites.
Variational Inference in Probabilistic Models
Lawrence, N. D · 2000
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J. C., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Practical Variational Inference for Neural Networks
Graves, A · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
ADADELTA: An Adaptive Learning Rate Method
Zeiler, M. D · 2012
Cited alongside, same era.
Auto-Encoding Variational Bayes
Kingma, D. P. and Welling, M · 2014
Cited alongside, same era.
New insights and perspectives on the natural gradient method
Martens, J · 2014
Cited alongside, same era.
Weight Uncertainty in Neural Networks
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Lei, J. B · 2015
Cited alongside, same era.
Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning
Gal, Y. and Ghahramani, Z · 2016
Cited alongside, same era.
Implicit Reparameterization Gradients
Figurnov, M., Mohamed, S., and Mnih, A · 2018
Closest in time.
Fast and Scalable Bayesian Deep Learning by Weight-Perturbation in Adam
Khan, M. E., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A · 2018
Closest in time.
On the Convergence of Adam and Beyond
Reddi, S. J., Kale, S., and Kumar, S · 2018
Closest in time.
Deep Bayesian Bandits Showdown: An Empirical Comparison of Bayesian Deep Networks for Thompson Sampling
Riquelme, C., Tucker, G., and Snoek, J · 2018
Closest in time.
A Scalable Laplace Approximation For Neural Networks
Ritter, H., Botev, A., and Barber, D · 2018
Closest in time.
Noisy Natural Gradient as Variational Inference
Zhang, G., Sun, S., Duvenaud, D., and Grosse, R · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Conjugate-Computation Variational Inference : Converting Variational Inference in Non-Conjugate Models to Inferences in Conjugate Models
Khan, M. E. and Lin, W · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R · 2017
Cited alongside, same era.
’In-Between’ Uncertainty in Bayesian Neural Networks
Foong, A. Y. K., Li, Y., Miguel Hernández-Lobato, J., and Turner, R. E · 2019
Closest in time.