Fetching the paper…
Reading the bibliography…
Second-order methods such as KFAC can be useful for neural net training.
A Stochastic Approximation Method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Fisher’s method of scoring
Osborne, M. R · 1992
Earlier work this paper cites.
Partitioned algorithms for maximum likelihood and other non-linear estimation
Smyth, G. K · 1996
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I · 1998
Earlier work this paper cites.
On “natural” learning and pruning in multilayered perceptrons
Heskes, T · 2000
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Schraudolph, N. N · 2002
Earlier work this paper cites.
Fisher scoring: An interpolation family and its Monte Carlo implementations
Wang, Y · 2010
Earlier work this paper cites.
Practical variational inference for neural networks
Graves, A · 2011
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
Martens, J · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
Optimization and nonlinear equations
Smyth, G. K · 2015
Earlier work this paper cites.
A kronecker-factored approximate fisher matrix for convolution layers
Grosse, R. and Martens, J · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Kipf, T. N. and Welling, M · 2016
Earlier work this paper cites.
Practical Gauss-Newton optimisation for deep learning
Botev, A., Ritter, H., and Barber, D · 2017
Earlier work this paper cites.
Conjugate-computation variational inference: Converting variational inference in non-conjugate models to inferences in conjugate models
Khan, M. and Lin, W · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
George, T., Laurent, C., Bouthillier, X., Ballas, N., and Vincent, P · 2018
Cited alongside, same era.
Fast yet Simple Natural-Gradient Descent for Variational Inference in Complex Models
Khan, M. E. and Nielsen, D · 2018
Cited alongside, same era.
Fast and scalable Bayesian deep learning by weight-perturbation in Adam
Khan, M. E., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A · 2018
Cited alongside, same era.
Preconditioner on matrix lie group for sgd
Li, X.-L · 2018
Cited alongside, same era.
Kronecker-factored curvature approximations for recurrent neural networks
Martens, J., Ba, J., and Johnson, M · 2018
Cited alongside, same era.
Mixed precision training
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., and Wu, H · 2018
Escaping the big data paradigm with compact transformers
Hassani, A., Walton, S., Shah, N., Abuduweili, A., Li, J., and Shi, H · 2021
Later among the works it cites.
Scalable marginal likelihood estimation for model selection in deep learning
Immer, A., Bauer, M., Fortuin, V., Rätsch, G., and Emtiyaz, K. M · 2021
Later among the works it cites.
Khan, M. E. and Rue, H · 2021
Later among the works it cites.
Tractable structured natural-gradient descent using local parameterizations
Lin, W., Nielsen, F., Emtiyaz, K. M., and Schmidt, M · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Noisy natural gradient as variational inference
Zhang, G., Sun, S., Duvenaud, D., and Grosse, R · 2018
Cited alongside, same era.
Limitations of the empirical Fisher approximation for natural gradient descent
Kunstner, F., Balles, L., and Hennig, P · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Practical deep learning with Bayesian principles
Osawa, K., Swaroop, S., Khan, M. E. E., Jain, A., Eschenhagen, R., Turner, R. E., and Yokota, R · 2019
Cited alongside, same era.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
If influence functions are the answer, then what is the question?
Bae, J., Ng, N., Lo, A., Ghassemi, M., and Grosse, R. B · 2022
Later among the works it cites.
Black box lie group preconditioners for sgd
Li, X · 2022
Later among the works it cites.
Bridging the gap between vision transformers and convolutional neural networks on small datasets
Lu, Z., Xie, H., Liu, C., and Zhang, Y · 2022
Later among the works it cites.
Analytic natural gradient updates for cholesky factor in gaussian variational approximation
Tan, L. S · 2022
Later among the works it cites.
Convolutions through the lens of tensor networks
Dangel, F · 2023
Closest in time.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., et al · 2023
Closest in time.
Kronecker-Factored Approximate Curvature for modern neural network architectures
Eschenhagen, R., Immer, A., Turner, R. E., Schneider, F., and Hennig, P · 2023
Closest in time.
Studying large language model generalization with influence functions
Grosse, R., Bae, J., Anil, C., Elhage, N., Tamkin, A., Tajdini, A., Steiner, B., Li, D., Durmus, E., Perez, E., et al · 2023
Closest in time.
Global context vision transformers
Hatamizadeh, A., Yin, H., Heinrich, G., Kautz, J., and Molchanov, P · 2023
Closest in time.
Simplifying momentum-based positive-definite submanifold optimization with applications to deep learning
Lin, W., Duruisseaux, V., Leok, M., Nielsen, F., Khan, M. E., and Schmidt, M · 2023
Closest in time.
PipeFisher: Efficient training of large language models using pipelining and Fisher information matrices
Osawa, K., Li, S., and Hoefler, T · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.
Patches are all you need?
Trockman, A. and Kolter, J. Z · 2023
Closest in time.
Repvit: Revisiting mobile cnn from vit perspective
Wang, A., Chen, H., Lin, Z., Pu, H., and Ding, G · 2023
Closest in time.