Fetching the paper…
Reading the bibliography…
Second-order optimizers are thought to hold the potential to speed up neural network training, but due to the enormous size of the curvature matrix, they typically require approximations to be computationally tractable.
Limitations of the empirical fisher approximation for natural gradient descent
Kunstner, F., Balles, L., and Hennig, P. (2019) · 1905
Earlier work this paper cites.
Fast convergence of natural gradient descent for overparameterized neural networks
Zhang, G., Martens, J., and Grosse, R. (2019) · 1905
Earlier work this paper cites.
Limitations of the empirical fisher approximation for natural gradient descent
Kunstner, F., Balles, L., and Hennig, P. (2019) · 1905
Earlier work this paper cites.
Fast convergence of natural gradient descent for overparameterized neural networks
Zhang, G., Martens, J., and Grosse, R. (2019) · 1905
Earlier work this paper cites.
Efficient subsampled gauss-newton and natural gradient methods for training neural networks
Ren, Y. and Goldfarb, D. (2019) · 1906
Earlier work this paper cites.
Efficient subsampled gauss-newton and natural gradient methods for training neural networks
Ren, Y. and Goldfarb, D. (2019) · 1906
Earlier work this paper cites.
A theoretical framework for back-propagation
LeCun, Y. (1988) · 1988
Earlier work this paper cites.
A theoretical framework for back-propagation
LeCun, Y. (1988) · 1988
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A. (1990) · 1990
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A. (1990) · 1990
Earlier work this paper cites.
Constrained realizations of gaussian fields-a simple algorithm
Hoffman, Y. and Ribak, E. (1991) · 1991
Earlier work this paper cites.
Constrained realizations of gaussian fields-a simple algorithm
Hoffman, Y. and Ribak, E. (1991) · 1991
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
MacKay, D. J. (1992) · 1992
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
MacKay, D. J. (1992) · 1992
Earlier work this paper cites.
Gradient-based learning algorithms for recurrent
Williams, R. J. and Zipser, D. (1995) · 1995
Earlier work this paper cites.
Gradient-based learning algorithms for recurrent
Williams, R. J. and Zipser, D. (1995) · 1995
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I. (1998) · 1998
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I. (1998) · 1998
Earlier work this paper cites.
Methods of information geometry
Amari, S.-i. and Nagaoka, H. (2000) · 2000
Earlier work this paper cites.
Methods of information geometry
Amari, S.-i. and Nagaoka, H. (2000) · 2000
Earlier work this paper cites.
Unifying regularisation methods for continual learning
Benzing, F. (2020) · 2006
Earlier work this paper cites.
Practical quasi-newton methods for training deep neural networks
Goldfarb, D., Ren, Y., and Bahamou, A. (2020) · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G. E. and Salakhutdinov, R. R. (2006) · 2006
Earlier work this paper cites.
Unifying regularisation methods for continual learning
Benzing, F. (2020) · 2006
Earlier work this paper cites.
Practical quasi-newton methods for training deep neural networks
Goldfarb, D., Ren, Y., and Bahamou, A. (2020) · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G. E. and Salakhutdinov, R. R. (2006) · 2006
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
Roux, N., Manzagol, P.-a., and Bengio, Y. (2007) · 2007
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
Roux, N., Manzagol, P.-a., and Bengio, Y. (2007) · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. (2009) · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. (2009) · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
Martens, J. et al. (2010) · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
Martens, J. et al. (2010) · 2010
Earlier work this paper cites.
Practical variational inference for neural networks
Graves, A. (2011) · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
Practical variational inference for neural networks
Graves, A. (2011) · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y. (2013) · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013) · 2013
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y. (2013) · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013) · 2013
Earlier work this paper cites.
How auto-encoders could provide credit assignment in deep networks via target propagation
Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Distributed optimization of deeply nested systems
Carreira-Perpinan, M. and Wang, W. (2014) · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
Martens, J. (2014) · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Cited alongside, same era.
How auto-encoders could provide credit assignment in deep networks via target propagation
Bengio, Y. (2014) · 2014
Cited alongside, same era.
Distributed optimization of deeply nested systems
Carreira-Perpinan, M. and Wang, W. (2014) · 2014
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014) · 2014
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R. (2017) · 2017
Later among the works it cites.
Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods
Aitchison, L. (2018) · 2018
Later among the works it cites.
Fast approximate natural gradient descent in a kronecker-factored eigenbasis
George, T., Laurent, C., Bouthillier, X., Ballas, N., and Vincent, P. (2018) · 2018
Later among the works it cites.
Decoupling backpropagation using constrained optimization methods
Gotmare, A., Thomas, V., Brea, J., and Jaggi, M. (2018) · 2018
Later among the works it cites.
Recasting gradient-based meta-learning as hierarchical bayes
Grant, E., Finn, C., Levine, S., Darrell, T., and Griffiths, T. (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Cited alongside, same era.
New insights and perspectives on the natural gradient method
Martens, J. (2014) · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Cited alongside, same era.
Weight uncertainty in neural network
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
Desjardins, G., Simonyan, K., Pascanu, R., and Kavukcuoglu, K. (2015) · 2015
Cited alongside, same era.
Scaling up natural gradient by sparsely factorizing the inverse fisher matrix
Grosse, R. and Salakhudinov, R. (2015) · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Cited alongside, same era.
Later among the works it cites.
Fast and scalable bayesian deep learning by weight-perturbation in adam
Khan, M., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A. (2018) · 2018
Later among the works it cites.
Kronecker-factored curvature approximations for recurrent neural networks
Martens, J., Ba, J., and Johnson, M. (2018) · 2018
Later among the works it cites.
Approximating real-time recurrent learning with random kronecker factors
Mujika, A., Meier, F., and Steger, A. (2018) · 2018
Later among the works it cites.
Online natural gradient as a kalman filter
Ollivier, Y. (2018) · 2018
Later among the works it cites.
A scalable laplace approximation for neural networks
Ritter, H., Botev, A., and Barber, D. (2018b) · 2018
Later among the works it cites.
Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods
Aitchison, L. (2018) · 2018
Later among the works it cites.
Fast approximate natural gradient descent in a kronecker-factored eigenbasis
George, T., Laurent, C., Bouthillier, X., Ballas, N., and Vincent, P. (2018) · 2018
Later among the works it cites.
Decoupling backpropagation using constrained optimization methods
Gotmare, A., Thomas, V., Brea, J., and Jaggi, M. (2018) · 2018
Later among the works it cites.
Recasting gradient-based meta-learning as hierarchical bayes
Grant, E., Finn, C., Levine, S., Darrell, T., and Griffiths, T. (2018) · 2018
Later among the works it cites.
Fast and scalable bayesian deep learning by weight-perturbation in adam
Khan, M., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A. (2018) · 2018
Later among the works it cites.
Kronecker-factored curvature approximations for recurrent neural networks
Martens, J., Ba, J., and Johnson, M. (2018) · 2018
Later among the works it cites.
Approximating real-time recurrent learning with random kronecker factors
Mujika, A., Meier, F., and Steger, A. (2018) · 2018
Later among the works it cites.
Online natural gradient as a kalman filter
Ollivier, Y. (2018) · 2018
Later among the works it cites.
A scalable laplace approximation for neural networks
Ritter, H., Botev, A., and Barber, D. (2018b) · 2018
Later among the works it cites.
Efficient full-matrix adaptive regularization
Agarwal, N., Bullins, B., Chen, X., Hazan, E., Singh, K., Zhang, C., and Zhang, Y. (2019) · 2019
Later among the works it cites.
Optimal kronecker-sum approximation of real time recurrent learning
Benzing, F., Gauy, M. M., Mujika, A., Martinsson, A., and Steger, A. (2019) · 2019
Later among the works it cites.
Exact natural gradient in deep linear networks and application to the nonlinear case
Bernacchia, A., Lengyel, M., and Hennequin, G. (2019) · 2019
Later among the works it cites.
Large-scale distributed second-order optimization using kronecker-factored approximate curvature for deep convolutional neural networks
Osawa, K., Tsuji, Y., Ueno, Y., Naruse, A., Yokota, R., and Matsuoka, S. (2019) · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019) · 2019
Later among the works it cites.
Efficient full-matrix adaptive regularization
Agarwal, N., Bullins, B., Chen, X., Hazan, E., Singh, K., Zhang, C., and Zhang, Y. (2019) · 2019
Later among the works it cites.
Optimal kronecker-sum approximation of real time recurrent learning
Benzing, F., Gauy, M. M., Mujika, A., Martinsson, A., and Steger, A. (2019) · 2019
Later among the works it cites.
Exact natural gradient in deep linear networks and application to the nonlinear case
Bernacchia, A., Lengyel, M., and Hennequin, G. (2019) · 2019
Later among the works it cites.
Large-scale distributed second-order optimization using kronecker-factored approximate curvature for deep convolutional neural networks
Osawa, K., Tsuji, Y., Ueno, Y., Naruse, A., Yokota, R., and Matsuoka, S. (2019) · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019) · 2019
Later among the works it cites.
Modular block-diagonal curvature approximations for feedforward architectures
Dangel, F., Harmeling, S., and Hennig, P. (2020) · 2020
Later among the works it cites.
A theoretical framework for target propagation
Meulemans, A., Carzaniga, F., Suykens, J., Sacramento, J., and Grewe, B. F. (2020) · 2020
Later among the works it cites.
Modular block-diagonal curvature approximations for feedforward architectures
Dangel, F., Harmeling, S., and Hennig, P. (2020) · 2020
Later among the works it cites.
A theoretical framework for target propagation
Meulemans, A., Carzaniga, F., Suykens, J., Sacramento, J., and Grewe, B. F. (2020) · 2020
Later among the works it cites.
Locoprop: Enhancing backprop via local loss optimization
Amid, E., Anil, R., and Warmuth, M. K. (2021) · 2021
Later among the works it cites.
Vivit: Curvature access through the generalized gauss-newton’s low-rank structure
Dangel, F., Tatzel, L., and Hennig, P. (2021) · 2021
Later among the works it cites.
Laplace redux–effortless bayesian deep learning
Daxberger, E., Kristiadi, A., Immer, A., Eschenhagen, R., Bauer, M., and Hennig, P. (2021) · 2021
Later among the works it cites.
A Note on Efficient Conditional Simulation of Gaussian Distributions
Doucet, A. (2010) · 2021
Later among the works it cites.
Scalable marginal likelihood estimation for model selection in deep learning
Immer, A., Bauer, M., Fortuin, V., Rätsch, G., and Khan, M. E. (2021) · 2021
Later among the works it cites.
Pathological spectra of the fisher information metric and its variants in deep neural networks
Karakida, R., Akaho, S., and Amari, S.-i. (2021) · 2021
Later among the works it cites.
Global inducing point variational posteriors for bayesian neural networks and deep gaussian processes
Ober, S. W. and Aitchison, L. (2021) · 2021
Later among the works it cites.
Locoprop: Enhancing backprop via local loss optimization
Amid, E., Anil, R., and Warmuth, M. K. (2021) · 2021
Later among the works it cites.
Vivit: Curvature access through the generalized gauss-newton’s low-rank structure
Dangel, F., Tatzel, L., and Hennig, P. (2021) · 2021
Later among the works it cites.
Laplace redux–effortless bayesian deep learning
Daxberger, E., Kristiadi, A., Immer, A., Eschenhagen, R., Bauer, M., and Hennig, P. (2021) · 2021
Later among the works it cites.
A Note on Efficient Conditional Simulation of Gaussian Distributions
Doucet, A. (2010) · 2021
Later among the works it cites.
Scalable marginal likelihood estimation for model selection in deep learning
Immer, A., Bauer, M., Fortuin, V., Rätsch, G., and Khan, M. E. (2021) · 2021
Later among the works it cites.
Pathological spectra of the fisher information metric and its variants in deep neural networks
Karakida, R., Akaho, S., and Amari, S.-i. (2021) · 2021
Later among the works it cites.
Global inducing point variational posteriors for bayesian neural networks and deep gaussian processes
Ober, S. W. and Aitchison, L. (2021) · 2021
Later among the works it cites.