Fetching the paper…
Reading the bibliography…
The generalized Gauss-Newton (GGN) approximation is often used to make practical Bayesian deep learning approaches scalable by replacing a second order derivative with a product of first order derivatives.
“Gauss-Newton approximation to Bayesian learning”
F Foresee and Martin Hagan · 1935
Earlier work this paper cites.
“Bayesian model comparison and backprop nets”
David MacKay · 1992
Earlier work this paper cites.
“The evidence framework applied to classification networks”
David MacKay · 1992
Earlier work this paper cites.
“Probable networks and plausible predictions—a review of practical Bayesian methods for supervised neural networks”
David MacKay · 1995
Earlier work this paper cites.
“Natural gradient works efficiently in learning”
Shun-Ichi Amari · 1998
Earlier work this paper cites.
“Matrix variate distributions”
Arjun Gupta and Daya Nagar · 1999
Earlier work this paper cites.
“Variational Inference in Probabilistic Models”
Neil. Lawrence · 2001
Earlier work this paper cites.
“Fast curvature matrix-vector products for second-order gradient descent”
Nicol Schraudolph · 2002
Earlier work this paper cites.
“Pattern recognition and machine learning”, Information Science and Statistics
Christopher Bishop · 2006
Earlier work this paper cites.
“Gaussian Processes for Machine Learning”
Carl Rasmussen and Christopher.. Williams · 2006
Earlier work this paper cites.
“Learning multiple layers of features from tiny images”, 2009
Alex Krizhevsky · 2009
Earlier work this paper cites.
“Variational Learning of Inducing Variables in Sparse Gaussian Processes”
Michalis Titsias · 2009
Earlier work this paper cites.
“MNIST handwritten digit database”, http://yann.lecun.com/exdb/mnist/, 2010
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
“MCMC Using Hamiltonian Dynamics”
Radford. Neal · 2010
Earlier work this paper cites.
“Learning Recurrent Neural Networks with Hessian-Free Optimization”
James Martens and Ilya Sutskever · 2011
Cited alongside, same era.
“Weight Uncertainty in Neural Networks”
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu and Daan Wierstra · 2015
Cited alongside, same era.
“Scalable Variational Gaussian Process Classification”
James Hensman, Alexander Matthews and Zoubin Ghahramani · 2015
Cited alongside, same era.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
“Variational dropout and the local reparameterization trick”
Durk Kingma, Tim Salimans and Max Welling · 2015
Cited alongside, same era.
“Optimizing neural networks with kronecker-factored approximate curvature”
James Martens and Roger Grosse · 2015
Cited alongside, same era.
“A scalable laplace approximation for neural networks”
Hippolyt Ritter, Aleksandar Botev and David Barber · 2018
Later among the works it cites.
“DeepOBS: A Deep Learning Optimizer Benchmark Suite”
Frank Schneider, Lukas Balles and Philipp Hennig · 2018
Later among the works it cites.
“Noisy Natural Gradient as Variational Inference”
Guodong Zhang, Shengyang Sun, David Duvenaud and Roger Grosse · 2018
Later among the works it cites.
“BackPACK: Packing more into Backprop”
Felix Dangel, Frederik Kunstner and Philipp Hennig · 2019
Later among the works it cites.
“’In-Between’Uncertainty in Bayesian Neural Networks”
Andrew Foong, Yingzhen Li, José Hernández-Lobato and Richard Turner · 2019
Later among the works it cites.
“Approximate Inference Turns Deep Networks into Gaussian Processes”
Mohammad Khan, Alexander Immer, Ehsan Abedi and Maciej Korzepa · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Obtaining Well Calibrated Probabilities Using Bayesian Binning”
Mahdi Naeini, Gregory. Cooper and Milos Hauskrecht · 2015
Cited alongside, same era.
“A kronecker-factored approximate fisher matrix for convolution layers”
Roger Grosse and James Martens · 2016
Cited alongside, same era.
“Practical Gauss-Newton Optimisation for Deep Learning”
Aleksandar Botev, Hippolyt Ritter and David Barber · 2017
Cited alongside, same era.
“Automatic differentiation in pytorch”, 2017
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga and Adam Lerer · 2017
Cited alongside, same era.
“Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms”, 2017
Han Xiao, Kashif Rasul and Roland Vollgraf · 2017
Cited alongside, same era.
“Optimization methods for large-scale machine learning”
Léon Bottou, Frank Curtis and Jorge Nocedal · 2018
Cited alongside, same era.
Later among the works it cites.
“Wide neural networks of any depth evolve as linear models under gradient descent”
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein and Jeffrey Pennington · 2019
Later among the works it cites.
“Do deep generative models know what they don’t know?”
Eric Nalisnick, Akihiro Matsukawa, Yee Teh, Dilan Gorur and Balaji Lakshminarayanan · 2019
Later among the works it cites.
“Practical deep learning with bayesian principles”
Kazuki Osawa, Siddharth Swaroop, Mohammad Khan, Anirudh Jain, Runa Eschenhagen, Richard Turner and Rio Yokota · 2019
Later among the works it cites.
“Uncertainty quantification using Bayesian neural networks in classification: Application to biomedical image segmentation”
Yongchan Kwon, Joong-Ho Won, Beom Kim and Myunghee Paik · 2020
Closest in time.
“New Insights and Perspectives on the Natural Gradient Method”
James Martens · 2020
Closest in time.
“Continual Deep Learning by Functional Regularisation of Memorable Past”
Pingbo Pan, Siddharth Swaroop, Alexander Immer, Runa Eschenhagen, Richard Turner and Mohammad Khan · 2020
Closest in time.
“How good is the bayes posterior in deep neural networks really?”
Florian Wenzel, Kevin Roth, Bastiaan Veeling, Jakub Światkowski, Linh Tran, Stephan Mandt, Jasper Snoek, Tim Salimans, Rodolphe Jenatton and Sebastian Nowozin · 2020
Closest in time.