Fetching the paper…
Reading the bibliography…
We show that many machine-learning algorithms are specific instances of a single algorithm called the \emph{Bayesian learning rule}.
Stein’s lemma for the reparameterization trick with exponential family mixtures
W. Lin, M. E. Khan, and M. Schmidt · 1910
Earlier work this paper cites.
A note on the delta-method for finding variance formulae
R. A. Dorfman · 1938
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
A useful theorem for nonlinear devices having Gaussian inputs
R. Price · 1958
Earlier work this paper cites.
Statistical forecasting for inventory control
R. G. Brown · 1959
Earlier work this paper cites.
Planning Production, Inventories, and Work Force
C. C. Holt, F. Modigliani, J. F. Muth, and H. A. Simon · 1960
Earlier work this paper cites.
A new approach to linear filtering and prediction problems
R. E. Kalman · 1960
Earlier work this paper cites.
Transformations des signaux aléatoires a travers les systemes non linéaires sans mémoire
G. Bonnet · 1964
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Conditional markov processes
R. L. Stratonovich · 1965
Earlier work this paper cites.
A new family of life distributions
Z. W. Birnbaum and S. C. Saunders · 1969
Earlier work this paper cites.
Ridge regression: Biased estimation for nonorthogonal problems
A. E. Hoerl and R. W. Kennard · 1970
Earlier work this paper cites.
Characterization of exponentially modified Gaussian peaks in chromatography
E. Grushka · 1972
Earlier work this paper cites.
Scale mixtures of normal distributions
D. F. Andrews and C. L. Mallows · 1974
Earlier work this paper cites.
Evolutionsstrategie. optimierung technischer systeme nach prinzipien der biologischen evolution, 1976
A. Huning · 1976
Earlier work this paper cites.
Maximum likelihood from incomplete data via the em algorithm
A. P. Dempster, N. M. Laird, and D. B. Rubin · 1977
Earlier work this paper cites.
On Cesaro’s convergence of the gradient descent method for finding saddle points of convex-concave functions
A. Nemirovski and D. Yudin · 1978
Earlier work this paper cites.
Goal seeking components for adaptive intelligence: An initial assessment
A. G. Barto and R. S. Sutton · 1981
Earlier work this paper cites.
On the rationale of maximum-entropy methods
E. T. Jaynes · 1982
Earlier work this paper cites.
Recursive parameter estimation using incomplete data
D. M. Titterington · 1984
Earlier work this paper cites.
Exponential smoothing: The state of the art
E. S. Gardner Jr · 1985
Earlier work this paper cites.
Memoir on the probability of the causes of events
P. S. Laplace · 1986
Earlier work this paper cites.
Accurate approximations for posterior moments and marginal densities
L. Tierney and J. B. Kadane · 1986
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
S. Becker and Y. LeCun · 1988
Earlier work this paper cites.
Optimal information processing and Bayes’s theorem
A. Zellner · 1988
Earlier work this paper cites.
Aggregating strategies
V. G. Vovk · 1990
Earlier work this paper cites.
Bayesian Methods for Adaptive Models
D. Mackay · 1991
Earlier work this paper cites.
The weighted majority algorithm
N. Littlestone and M. K. Warmuth · 1994
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
S. Hochreiter and J. Schmidhuber · 1995
Earlier work this paper cites.
A variational approach to Bayesian logistic regression problems and their extensions
T. Jaakkola and M. Jordan · 1996
Earlier work this paper cites.
Normal inverse Gaussian distributions and stochastic volatility modelling
O. E. Barndorff-Nielsen · 1997
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
S. Amari · 1998
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, and G. B. Orrand K-R. Müller · 1998
Earlier work this paper cites.
A view of the EM algorithm that justifies incremental, sparse, and other variants
R. M. Neal and G. E. Hinton · 1998
Earlier work this paper cites.
Convergence of a stochastic approximation version of the EM algorithm
B. Delyon, M. Lavielle, and E. Moulines · 1999
Earlier work this paper cites.
A unifying review of linear Gaussian models
S. Roweis and Z. Ghahramani · 1999
Earlier work this paper cites.
Fast learning of on-line EM algorithm
M-A. Sato · 1999
Earlier work this paper cites.
Theoretical analysis of a class of randomized regularization methods
T. Zhang · 1999
Earlier work this paper cites.
Relative loss bounds for on-line density estimation with the exponential family of distributions
K. S. Azoury and M. K. Warmuth · 2001
Earlier work this paper cites.
Online model selection based on the variational Bayes
M-A. Sato · 2001
Earlier work this paper cites.
Variational extensions to EM and multinomial PCA
W. Buntine · 2002
Earlier work this paper cites.
On an equivalence between PLSI and LDA
Mark Girolami and Ata Kabán · 2003
Earlier work this paper cites.
The skew-normal distribution and related multivariate families
A. Azzalini · 2005
Earlier work this paper cites.
Clustering with Bregman divergences
A. Banerjee, S. Merugu, I. S. Dhillon, and J. Ghosh · 2005
Earlier work this paper cites.
Variational message passing
J. Winn and C. M. Bishop · 2005
Earlier work this paper cites.
Pattern Recognition and Machine Learning (Information Science and Statistics)
C. M. Bishop · 2006
Earlier work this paper cites.
Multivariate scale mixture of Gaussians modeling
T. Eltoft, T. Kim, and T-W. Lee · 2006
Cited alongside, same era.
Variational Bayesian multinomial probit regression with Gaussian process priors
M. Girolami and S. Rogers · 2006
Cited alongside, same era.
Gaussian Processes for Machine Learning
C. E. Rasmussen and C. K. I. Williams · 2006
Cited alongside, same era.
A correlated topic model of science
D. M. Blei and J. D. Lafferty · 2007
Cited alongside, same era.
PAC-Bayesian supervised classification: The thermodynamics of statistical learning. institute of mathematical statistics lecture notes—monograph series 56
O Catoni · 2007
Cited alongside, same era.
Approximate Bayesian inference for hierarchical Gaussian Markov random fields models
H. Rue and S. Martino · 2007
A theoretical analysis of optimization by Gaussian continuation
H. Mobahi and J. W. Fisher III · 2015
Later among the works it cites.
The information geometry of mirror descent
G. Raskutti and S. Mukherjee · 2015
Later among the works it cites.
Evolutionary multimodal optimization: A short survey
K-C. Wong · 2015
Later among the works it cites.
Information geometry and its applications
S. Amari · 2016
Later among the works it cites.
A general framework for updating belief distributions
P. G. Bissiri, C. C. Holmes, and S. G. Walker · 2016
Later among the works it cites.
Uncertainty in Deep Learning
Y. Gal · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Smoothing-based optimization
M. Leordeanu and M. Hebert · 2008
Cited alongside, same era.
Graphical models, exponential families, and variational inference
M. J. Wainwright and M. I. Jordan · 2008
Cited alongside, same era.
Deterministic latent variable models and their pitfalls
M. Welling, C. Chemudugunta, and N. Sutter · 2008
Cited alongside, same era.
On-line expectation–maximization algorithm for latent data models
O. Cappé and E. Moulines · 2009
Cited alongside, same era.
Probabilistic graphical models: principles and techniques
D. Koller and N. Friedman · 2009
Cited alongside, same era.
The variational Gaussian approximation revisited
M. Opper and C. Archambeau · 2009
Cited alongside, same era.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Later among the works it cites.
On graduated optimization for stochastic non-convex problems
E. Hazan, K. Y. Levy, and S. Shalev-Shwartz · 2016
Later among the works it cites.
Categorical reparameterization with gumbel-softmax
E. Jang, S Gu, and B. Poole · 2016
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Later among the works it cites.
Structured and efficient variational deep learning with matrix Gaussian posteriors
C. Louizos and M. Welling · 2016
Later among the works it cites.
The concrete distribution: A continuous relaxation of discrete random variables
C. J. Maddison, A. Mnih, and Y. W. Teh · 2016
Later among the works it cites.
Monte carlo structured SVI for two-level non-conjugate models
R. Sheth and R. Khardon · 2016
Later among the works it cites.
Variational inference: A review for statisticians
D. M. Blei, A. Kucukelbir, and J. D. McAuliffe · 2017
Later among the works it cites.
G. K. Dziugaite and D. M. Roy · 2017
Later among the works it cites.
Concrete dropout
Y. Gal, J. Hron, and A. Kendall · 2017
Later among the works it cites.
Conjugate-computation variational inference: converting variational inference in non-conjugate models to inferences in conjugate models
M. E. Khan and W. Lin · 2017
Later among the works it cites.
Stochastic gradient descent as approximate Bayesian inference
S. Mandt, M. D. Hoffman, and D. M. Blei · 2017
Later among the works it cites.
Information-geometric optimization algorithms: A unifying picture via invariance principles
Y. Ollivier, L. Arnold, A. Auger, and N. Hansen · 2017
Later among the works it cites.
Bayesian computing with INLA: A review
H. Rue, A. Riebler, S. H. Sørbye, J. B. Illian, D. P. Simpson, and F. K. Lindgren · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
P. Chaudhari and S. Soatto · 2018
Later among the works it cites.
Matrix variate distributions , volume 104
A. K. Gupta and D. K. Nagar · 2018
Later among the works it cites.
Fast yet simple natural-gradient descent for variational inference in complex models
M. E. Khan and D. Nielsen · 2018
Later among the works it cites.
Fast and scalable Bayesian deep learning by weight-perturbation in Adam
M. E. Khan, D. Nielsen, V. Tangkaratt, W. Lin, Y. Gal, and A. Srivastava · 2018
Later among the works it cites.
SLANG: Fast structured covariance approximations for Bayesian deep learning with natural gradient
A. Mishkin, F. Kunstner, D. Nielsen M. Schmidt, and M. E. Khan · 2018
Later among the works it cites.
Online natural gradient as a Kalman filter
Y. Ollivier · 2018
Later among the works it cites.
A scalable Laplace approximation for neural networks
H. Ritter, A. Botev, and D. Barber · 2018
Later among the works it cites.
Natural gradients in practice: Non-conjugate variational inference in Gaussian process models
H. Salimbeni, S. Eleftheriadis, and J. Hensman · 2018
Later among the works it cites.
An introduction to probabilistic programming
J-W. van de Meent, B. Paige, H. Yang, and F. Wood · 2018
Later among the works it cites.
The many faces of exponential weights in online learning
D. van der Hoeven, T. van Erven, and W. Kotłowski · 2018
Later among the works it cites.
An empirical study of binary neural networks’ optimisation
M. Alizadeh, J. Fernández-Marqués, N. D. Lane, and Y. Gal · 2019
Later among the works it cites.
Latent weights do not exist: Rethinking binarized neural network optimization
K. Helwegen, J. Widdicombe, L. Geiger, Z. Liu, K-T. Cheng, and R. Nusselder · 2019
Later among the works it cites.
Approximate inference turns deep networks into Gaussian processes
M. E. Khan, A. Immer, E. Abedi, and M. Korzepa · 2019
Later among the works it cites.
A simple baseline for Bayesian uncertainty in deep learning
W. J. Maddox, P. Izmailov, T. Garipov, D. P. Vetrov, and A. G. Wilson · 2019
Later among the works it cites.
Practical deep learning with Bayesian principles
K. Osawa, S. Swaroop, A. Jain, R. Eschenhagen, R. E. Turner, R. Yokota, and M. E. Khan · 2019
Later among the works it cites.
Algorithms and Theory for Variational Inference in Two-Level Non-conjugate Models
R. Sheth · 2019
Later among the works it cites.
Understanding straight-through estimator in training activation quantized neural nets
P. Yin, J. Lyu, S. Zhang, S. Osher, Y. Qi, and J. Xin · 2019
Later among the works it cites.
Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods
L. Aitchison · 2020
Later among the works it cites.
Divergence-based motivation for online EM and combining hidden variable models
E. Amid and M. K. Warmuth · 2020
Later among the works it cites.
Fast variational learning in state-space Gaussian process models
P. E. Chang, W. J. Wilkinson, M. E. Khan, and A. Solin · 2020
Later among the works it cites.
New insights and perspectives on the natural gradient method
J. Martens · 2020
Later among the works it cites.
Training binary neural networks using the Bayesian learning rule
X. Meng, R. Bachmann, and M. E. Khan · 2020
Later among the works it cites.
Homeomorphic-invariance of EM: Non-asymptotic convergence in KL divergence for exponential families via mirror descent
F. Kunstner, R. Kumar, and M. Schmidt · 2021
Closest in time.
Tractable structured natural-gradient descent using local parameterizations
W. Lin, F. Nielsen, M. E. Khan, and M. Schmidt · 2021
Closest in time.
Advances in variational inference
C. Zhang, J. Bütepage, H. Kjellström, and S. Mandt · 2026
Closest in time.