Fetching the paper…
Reading the bibliography…
Given the massive cost of language model pre-training, a non-trivial improvement of the optimization algorithm would lead to a material reduction on the time and cost of training.
Approximate confidence intervals
Bartlett, M · 1953
Earlier work this paper cites.
The convergence of a class of double-rank minimization algorithms 1. general considerations
Broyden, C. G · 1970
Earlier work this paper cites.
Improving the convergence of back-propagation learning with
Becker, S. and Le Cun, Y · 1988
Earlier work this paper cites.
A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines
Hutchinson, M. F · 1989
Earlier work this paper cites.
Rprop: a fast adaptive learning algorithm
Braun, H. and Riedmiller, M · 1992
Earlier work this paper cites.
Numerical methods for unconstrained optimization and nonlinear equations
Dennis Jr, J. E. and Schnabel, R. B · 1996
Earlier work this paper cites.
Trust-region methods, siam
Conn, A. R., Gould, N., and Toint, P. L · 2000
Earlier work this paper cites.
Iterative solution of nonlinear equations in several variables
Ortega, J. M. and Rheinboldt, W. C · 2000
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Schraudolph, N. N · 2002
Earlier work this paper cites.
Convex optimization
Boyd, S. P. and Vandenberghe, L · 2004
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Nesterov, Y. and Polyak, B. T · 2006
Earlier work this paper cites.
Deep learning via hessian-free optimization
Martens, J. et al · 2010
Earlier work this paper cites.
Improved preconditioner for hessian free optimization
Chapelle, O., Erhan, D., et al · 2011
Earlier work this paper cites.
Hessian matrix vs. gauss–newton hessian matrix
Chen, P · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y · 2013
Earlier work this paper cites.
No more pesky learning rates
Schaul, T., Zhang, S., and LeCun, Y · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
Improved bounds on sample size for implicit matrix trace estimators
Roosta-Khorasani, F. and Ascher, U · 2015
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Dozat, T · 2016
Earlier work this paper cites.
A kronecker-factored approximate fisher matrix for convolution layers
Grosse, R. and Martens, J · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Sagun, L., Bottou, L., and LeCun, Y · 2016
Earlier work this paper cites.
Distributed second-order optimization using kronecker-factored approximations
Ba, J., Grosse, R., and Martens, J · 2017
Cited alongside, same era.
Practical gauss-newton optimisation for deep learning
Botev, A., Ritter, H., and Barber, D · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Regularizing and optimizing lstm language models
Merity, S., Keskar, N. S., and Socher, R · 2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
On the promise of the stochastic generalized gauss-newton method for training dnns
Gargiani, M., Zanelli, A., Diehl, M., and Hutter, F · 2020
Later among the works it cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Later among the works it cites.
Understanding the difficulty of training transformers
Liu, L., Liu, X., Gao, J., Chen, W., and Han, J · 2020
Later among the works it cites.
New insights and perspectives on the natural gradient method
Martens, J · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Balles, L. and Hennig, P · 2018
Cited alongside, same era.
signsgd: Compressed optimisation for non-convex problems
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
George, T., Laurent, C., Bouthillier, X., Ballas, N., and Vincent, P · 2018
Cited alongside, same era.
Shampoo: Preconditioned stochastic tensor optimization
Gupta, V., Koren, T., and Singer, Y · 2018
Cited alongside, same era.
Kronecker-factored curvature approximations for recurrent neural networks
Martens, J., Ba, J., and Johnson, M · 2018
Cited alongside, same era.
Later among the works it cites.
The implicit and explicit regularization effects of dropout
Wei, C., Kakade, S., and Ma, T · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Later among the works it cites.
Pyhessian: Neural networks through the lens of the hessian
Yao, Z., Gholami, A., Keutzer, K., and Mahoney, M. W · 2020
Later among the works it cites.
Why are adaptive methods good for attention models?
Zhang, J., Karimireddy, S. P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S · 2020
Later among the works it cites.
Adabelief optimizer: Adapting stepsizes by the belief in observed gradients
Zhuang, J., Tang, T., Ding, Y., Tatikonda, S. C., Dvornek, N., Papademetris, X., and Duncan, J · 2020
Later among the works it cites.
How to train bert with an academic budget
Izsak, P., Berchansky, M., and Levy, O · 2021
Later among the works it cites.
Doubly adaptive scaled algorithm for machine learning using second-order information
Jahani, M., Rusakov, S., Shi, Z., Richtárik, P., Mahoney, M. W., and Takáč, M · 2021
Later among the works it cites.
Mistral – a journey towards reproducible language model training
Karamcheti, S., Orr, L., Bolton, J., Zhang, T., Goel, K., Narayan, A., Bommasani, R., Narayanan, D., Hashimoto, T., Jurafsky, D., Manning, C. D., Potts, C., Ré, C., and Liang, P · 2021
Later among the works it cites.
Stability and convergence of stochastic gradient clipping: Beyond lipschitz continuity and smoothness
Mai, V. V. and Johansson, M · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al · 2021
Later among the works it cites.
A deeper look at the hessian eigenspectrum of deep neural networks and its applications to regularization
Sankar, A. R., Khasbage, Y., Vigneswaran, R., and Balasubramanian, V. N · 2021
Later among the works it cites.
Adahessian: An adaptive second order optimizer for machine learning
Yao, Z., Gholami, A., Shen, S., Mustafa, M., Keutzer, K., and Mahoney, M · 2021
Later among the works it cites.
Gpt-neox-20b: An open-source autoregressive language model
Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., et al · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Later among the works it cites.
Robustness to unbounded smoothness of generalized signsgd
Crawshaw, M., Liu, M., Orabona, F., Zhang, W., and Zhuang, Z · 2022
Later among the works it cites.
Neural Network Training Dynamics
Grosse, R · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms
Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Liu, Y., Pham, H., Dong, X., Luong, T., Hsieh, C.-J., et al · 2023
Closest in time.
Kunstner, F., Chen, J., Lavington, J. W., and Schmidt, M · 2023
Closest in time.
Gpt-4 technical report
OpenAI · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.