Fetching the paper…
Reading the bibliography…
Recent work has argued that stochastic gradient descent can approximate the Bayesian uncertainty in model parameters near local minima.
An invariant form for the prior probability in estimation problems
Harold Jeffreys · 1946
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Handbook of Stochastic Methods , volume 4
Crispin W Gardiner · 1985
Earlier work this paper cites.
Bias reduction of maximum likelihood estimates
David Firth · 1993
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Prior knowledge and preferential structures in gradient descent learning algorithms
Robert E Mahony and Robert C Williamson · 2001
Earlier work this paper cites.
Bayes, jeffreys, prior distributions and the philosophy of statistics
Andrew Gelman · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Riemann manifold langevin and hamiltonian monte carlo methods
Mark Girolami and Ben Calderhead · 2011
Earlier work this paper cites.
Information-geometric optimization algorithms: A unifying picture via invariance principles
Yann Ollivier, Ludovic Arnold, Anne Auger, and Nikolaus Hansen · 2011
Cited alongside, same era.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Cited alongside, same era.
Bayesian posterior sampling via stochastic gradient fisher scoring
Sungjin Ahn, Anoop Korattikara, and Max Welling · 2012
Cited alongside, same era.
Stochastic gradient riemannian langevin dynamics on the probability simplex
Sam Patterson and Yee Whye Teh · 2013
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
Accurate, large minibatch SGD: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Later among the works it cites.
A tutorial on fisher information
Alexander Ly, Maarten Marsman, Josine Verhagen, Raoul PPP Grasman, and Eric-Jan Wagenmakers · 2017
Later among the works it cites.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Later among the works it cites.
Natural langevin dynamics for neural networks
Gaétan Marceau-Caron and Yann Ollivier · 2017
Later among the works it cites.
Noisy natural gradient as variational inference
Guodong Zhang, Shengyang Sun, David Duvenaud, and Roger Grosse · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Coupling adaptive batch sizes with learning rates
Lukas Balles, Javier Romero, and Philipp Hennig · 2016
Cited alongside, same era.
Preconditioned stochastic gradient langevin dynamics for deep neural networks
Chunyuan Li, Changyou Chen, David E Carlson, and Lawrence Carin · 2016
Cited alongside, same era.
Stochastic gradient langevin dynamics that exploit neural network structure
Zachary Nado, Jasper Snoek, Roger Grosse, David Duvenaud, Bowen Xu, and James Martens · 2018
Closest in time.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L. Smith and Quoc V. Le · 2018
Closest in time.
Don’t decay the learning rate, increase the batch size
Samuel L. Smith, Pieter-Jan Kindermans, and Quoc V. Le · 2018
Closest in time.