Fetching the paper…
Reading the bibliography…
In this paper, we develop an efficient sketchy empirical natural gradient method (SENG) for large-scale deep learning problems.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Neural learning in structured parameter spaces-natural riemannian gradient
Amari, S.-i · 1997
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Hazan, E., Agarwal, A., and Kale, S · 2007
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
Roux, N. L., Manzagol, P.-A., and Bengio, Y · 2008
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
A stochastic quasi-Newton method for large-scale optimization
Byrd, R. H., Hansen, S. L., Nocedal, J., and Singer, Y · 2016
Earlier work this paper cites.
Spsd matrix approximation vis column selection: Theories, algorithms, and extensions
Wang, S., Luo, L., and Zhang, Z · 2016
Earlier work this paper cites.
Practical Gauss-Newton optimisation for deep learning
Botev, A., Ritter, H., and Barber, D · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Cited alongside, same era.
Exact natural gradient in deep linear networks and its application to the nonlinear case
Bernacchia, A., Lengyel, M., and Hennequin, G · 2018
Cited alongside, same era.
A gram-gauss-newton method learning overparameterized deep neural networks for regression problems
Cai, T., Gao, R., Hou, J., Chen, S., Wang, D., He, D., Zhang, Z., and Wang, L · 2019
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2019
Later among the works it cites.
Optimization for deep learning: theory and algorithms
Sun, R · 2019
Later among the works it cites.
A stochastic extra-step quasi-newton method for nonsmooth nonconvex optimization
Yang, M., Milzarek, A., Wen, Z., and Zhang, T · 2019
Later among the works it cites.
Fast convergence of natural gradient descent for over-parameterized neural networks
Zhang, G., Martens, J., and Grosse, R. B · 2019
Later among the works it cites.
Understanding approximate fisher information for fast convergence of natural gradient descent in wide neural networks
Karakida, R. and Osawa, K · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mattson, P., Cheng, C., Coleman, C., Diamos, G., Micikevicius, P., Patterson, D., Tang, H., Wei, G.-Y., Bailis, P., Bittorf, V., et al · 2019
Cited alongside, same era.
Efficient subsampled gauss-newton and natural gradient methods for training neural networks
Ren, Y. and Goldfarb, D · 2019
Cited alongside, same era.
Sketched ridge regression: Optimization perspective, statistical perspective, and model averaging
Wang, S., Gittens, A., and Mahoney, M. W
Cited in the paper.
Stochastic Quasi-Newton Methods for Nonconvex Stochastic Optimization
Wang, X., Ma, S., Goldfarb, D., and Liu, W
Cited in the paper.
New insights and perspectives on the natural gradient method
Martens, J · 2020
Closest in time.
Scalable and practical natural gradient for large-scale deep learning
Osawa, K., Tsuji, Y., Ueno, Y., Naruse, A., Foo, C. S., and Yokota, R · 2020
Closest in time.