Continual learning with hypernetworks
Original
Johannes von Oswald, Christian Henning, João Sacramento, and Benjamin F. Grewe · 1906
Earlier work this paper cites.
On estimation of a probability density function and mode
Emanuel Parzen · 1962
Earlier work this paper cites.
Improvement on some known nonparametric uniformly consistent estimators of derivatives of a density
Radhey S Singh · 1977
Earlier work this paper cites.
Neural network ensembles
L. K. Hansen and P. Salamon · 1990
Earlier work this paper cites.
On the algebraic structure of feedforward network weight spaces
Robert Hecht-Nielsen · 1990
Earlier work this paper cites.
A statistical approach to learning and generalization in layered neural networks
E. Levin, N. Tishby, and S. A. Solla · 1990
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
David JC MacKay · 1992
Earlier work this paper cites.
On the geometry of feedforward neural network error surfaces
An Mei Chen, Haw-minn Lu, and Robert Hecht-Nielsen · 1993
Earlier work this paper cites.
Bayesian learning for neural networks
Radford M Neal · 1995
Earlier work this paper cites.
Bagging predictors
Leo Breiman · 1996
Earlier work this paper cites.
The variational formulation of the fokker–planck equation
Richard Jordan, David Kinderlehrer, and Felix Otto · 1998
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Computation with infinite neural networks
Christopher KI Williams · 1998
Earlier work this paper cites.
On the surprising behavior of distance metrics in high dimensional space
Charu C Aggarwal, Alexander Hinneburg, and Daniel A Keim · 2001
Earlier work this paper cites.
The geometry of dissipative evolution equations: the porous medium equation
Felix Otto · 2001
Earlier work this paper cites.
How good is the bayes posterior in deep neural networks really?
Original
Florian Wenzel, Kevin Roth, Bastiaan S Veeling, Jakub Światkowski, Linh Tran, Stephan Mandt, Jasper Snoek, Tim Salimans, Rodolphe Jenatton, and Sebastian Nowozin · 2002
Earlier work this paper cites.
Gradient flows: in metric spaces and in the space of probability measures
Luigi Ambrosio, Nicola Gigli, and Giuseppe Savaré · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky et al · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Bayesian learning for neural networks , volume 118
Radford M Neal · 2012
Earlier work this paper cites.
Continuity equations and ode flows with non-smooth velocity
Luigi Ambrosio and Gianluca Crippa · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Symmetry-invariant optimization in deep networks
Original
Vijay Badrinarayanan, Bamdev Mishra, and Roberto Cipolla · 2015
Earlier work this paper cites.
Weight uncertainty in neural networks
Original
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.