Fetching the paper…
Reading the bibliography…
Leveraging second-order information about the loss at the scale of deep networks is one of the main lines of approach for improving the performance of current optimizers for deep learning.
Natural gradient works efficiently in learning
Amari, S.-i · 1998
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J. C., Hazan, E., and Singer, Y · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
Martens, J · 2010
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
A Kronecker-factored approximate fisher matrix for convolution layers
Grosse, R. and Martens, J · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Second-order stochastic optimization for machine learning in linear time
Agarwal, N., Bullins, B., and Hazan, E · 2017
Earlier work this paper cites.
Distributed second-order optimization using Kronecker-factored approximations
Ba, J., Grosse, R., and Martens, J · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Alistarh, D., Hoefler, T., Johansson, M., Konstantinov, N., Khirirat, S., and Renggli, C · 2018
Cited alongside, same era.
Pytorch-sparse library
Fey, M. et al · 2018
Cited alongside, same era.
Shampoo: Preconditioned stochastic tensor optimization
Gupta, V., Koren, T., and Singer, Y · 2018
Cited alongside, same era.
Neumann optimizer: A practical optimization algorithm for deep neural networks
Krishnan, S., Xiao, Y., and Saurous, R. A · 2018
Cited alongside, same era.
An evaluation of Fisher approximations beyond Kronecker factorization
Laurent, C., George, T., Bouthillier, X., Ballas, N., and Vincent, P · 2018
Cited alongside, same era.
PowerSGD: Practical Low-Rank Gradient Compression for Distributed Optimization
Vogels, T., Karimireddy, S. P., and Jaggi, M · 2019
Later among the works it cites.
MLPrune: Multi-layer pruning for automated neural network compression, 2019
Zeng, W. and Urtasun, R · 2019
Later among the works it cites.
Three mechanisms of weight decay regularization
Zhang, G., Wang, C., Xu, B., and Grosse, R · 2019
Later among the works it cites.
WoodFisher: Efficient second-order approximation for neural network compression
Singh, S. P. and Alistarh, D · 2020
Later among the works it cites.
The error-feedback framework: Better rates for sgd with delayed gradients and compressed communication
Stich, S. U. and Karimireddy, S. P · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kronecker-factored curvature approximations for recurrent neural networks
Martens, J., Ba, J., and Johnson, M · 2018
Cited alongside, same era.
Sparsified sgd with memory
Stich, S. U., Cordonnier, J.-B., and Jaggi, M · 2018
Cited alongside, same era.
Efficient full-matrix adaptive regularization
Agarwal, N., Bullins, B., Chen, X., Hazan, E., Singh, K., Zhang, C., and Zhang, Y · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Error feedback fixes SignSGD and other gradient compression schemes
Karimireddy, S. P., Rebjock, Q., Stich, S. U., and Jaggi, M · 2019
Cited alongside, same era.
Large-scale distributed second-order optimization using Kronecker-factored approximate curvature for deep convolutional neural networks
Osawa, K., Tsuji, Y., Ueno, Y., Naruse, A., Yokota, R., and Matsuoka, S · 2019
Cited alongside, same era.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Later among the works it cites.
M-FAC: Efficient matrix-free approximations of second-order information
Frantar, E., Kurtic, E., and Alistarh, D · 2021
Later among the works it cites.
Elastic consistency: A general consistency model for distributed stochastic gradient descent
Nadiradze, G., Markov, I., Chatterjee, B., Kungurtsev, V., and Alistarh, D · 2021
Later among the works it cites.
ADAHESSIAN: An adaptive second order optimizer for machine learning
Yao, Z., Gholami, A., Shen, S., Keutzer, K., and Mahoney, M. W · 2021
Later among the works it cites.
FFCV: Accelerating training by removing data bottlenecks
Leclerc, G., Ilyas, A., Engstrom, L., Park, S. M., Salman, H., and Madry, A · 2022
Later among the works it cites.
On distributed adaptive optimization with gradient compression
Li, X., Karimi, B., and Li, P · 2022
Later among the works it cites.
ASDL: A unified interface for gradient preconditioning in pytorch, 2023
Osawa, K., Ishikawa, S., Yokota, R., Li, S., and Hoefler, T · 2023
Closest in time.