Fetching the paper…
Reading the bibliography…
We introduce the Kronecker factored online Laplace approximation for overcoming catastrophic forgetting in neural networks.
Some Methods of Speeding up the Convergence of Iteration Methods
B. T. Polyak · 1964
Earlier work this paper cites.
Stochastic Models, Estimation and Control
P. Maybeck · 1982
Earlier work this paper cites.
A Method of Solving a Convex Programming Problem with Convergence Rate O (1/k2)
Y. Nesterov · 1983
Earlier work this paper cites.
Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem
M. McCloskey and N. J. Cohen · 1989
Earlier work this paper cites.
Connectionist Models of Recognition Memory: Constraints Imposed by Learning and Forgetting Functions
R. Ratcliff · 1990
Earlier work this paper cites.
A Practical Bayesian Framework for Backpropagation Networks
D. J. C. MacKay · 1992
Earlier work this paper cites.
Gradient-based Learning Applied to Document Recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
A Bayesian Approach to On-line Learning
M. Opper and O. Winther · 1998
Earlier work this paper cites.
Catastrophic Forgetting in Connectionist Networks
R. M. French · 1999
Earlier work this paper cites.
Matrix Variate Distributions
A. K. Gupta and D. K. Nagar · 1999
Earlier work this paper cites.
Online Variational Bayesian Learning
Z. Ghahramani · 2000
Earlier work this paper cites.
Expectation Propagation for Approximate Bayesian Inference
T. P. Minka · 2001
Earlier work this paper cites.
On-line Variational Bayesian Learning
A. Honkela and H. Valpola · 2003
Earlier work this paper cites.
Laplace Propagation
E. Eskin, A. J. Smola, and S. Vishwanathan · 2004
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
A. Krizhevsky and G. Hinton · 2009
Cited alongside, same era.
Reading Digits in Natural Images with Unsupervised Feature Learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng · 2011
Cited alongside, same era.
An Empirical Investigation of Catastrophic Forgetting in Gradient-based Neural Networks
I. J. Goodfellow, M. Mirza, D. Xiao, A. Courville, and Y. Bengio · 2013
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
New Insights and Perspectives on the Natural Gradient Method
J. Martens · 2014
Cited alongside, same era.
Overcoming Catastrophic Forgetting by Incremental Moment Matching
S.-W. Lee, J.-H. Kim, J. Jun, J.-W. Ha, and B.-T. Zhang · 2017
Later among the works it cites.
Gradient Episodic Memory for Continual Learning
D. Lopez-Paz and M. Ranzato · 2017
Later among the works it cites.
Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
L. Sagun, U. Evci, V. U. Guney, Y. Dauphin, and L. Bottou · 2017
Later among the works it cites.
Continual Learning with Deep Generative Replay
H. Shin, J. K. Lee, J. Kim, and J. Kim · 2017
Later among the works it cites.
Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
H. Xiao, K. Rasul, and R. Vollgraf · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weight Uncertainty in Neural Networks
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra · 2015
Cited alongside, same era.
Lasagne: First release., August 2015
S. Dieleman, J. Schlüter, C. Raffel, E. Olson, S. K. Sønderby, D. Nouri, et al · 2015
Cited alongside, same era.
Optimizing Neural Networks with Kronecker-factored Approximate Curvature
J. Martens and R. Grosse · 2015
Cited alongside, same era.
A Kronecker-factored Approximate Fisher Matrix for Convolution Layers
R. Grosse and J. Martens · 2016
Cited alongside, same era.
A. A. Rusu, N. C. Rabinowitz, G. Desjardins, H. Soyer, J. Kirkpatrick, K. Kavukcuoglu, R. Pascanu, and R. Hadsell · 2016
Cited alongside, same era.
Theano: A Python Framework for Fast Computation of Mathematical Expressions
Theano Development Team · 2016
Cited alongside, same era.
Practical Gauss-Newton Optimisation for Deep Learning
A. Botev, H. Ritter, and D. Barber · 2017
Cited alongside, same era.
Continual Learning through Synaptic Intelligence
F. Zenke, B. Poole, and S. Ganguli · 2017
Later among the works it cites.
Recasting Gradient-Based Meta-Learning as Hierarchical Bayes
E. Grant, C. Finn, S. Levine, T. Darrell, and T. Griffiths · 2018
Closest in time.
Overcoming Catastrophic Interference using Conceptor-Aided Backpropagation
X. He and H. Jaeger · 2018
Closest in time.
Note on the Quadratic Penalties in Elastic Weight Consolidation
F. Huszár · 2018
Closest in time.
Reply to Huszár: The Elastic Weight Consolidation Penalty is Empirically Valid
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, D. Hassabis, C. Clopath, D. Kumaran, and R. Hadsell · 2018
Closest in time.
Variational Continual Learning
C. V. Nguyen, Y. Li, T. D. Bui, and R. E. Turner · 2018
Closest in time.
A Scalable Laplace Approximation for Neural Networks
H. Ritter, A. Botev, and D. Barber · 2018
Closest in time.
Overcoming Catastrophic Forgetting with Hard Attention to the Task
J. Serrà, D. Surís, M. Miron, and A. Karatzoglou · 2018
Closest in time.