Fetching the paper…
Reading the bibliography…
Large language models (LLMs) demonstrate remarkable performance, and improving their pre-training process appears to be key to enhancing their capabilities further.
The MNIST database of handwritten digits
LeCun, Y., Cortes, C., and Burges, C. J · 1998
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
On learning rates and schrödinger operators, 2020
Shi, B., Su, W. J., and Jordan, M. I · 2004
Earlier work this paper cites.
Bootstrap your own latent: A new approach to self-supervised learning, 2020
Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. A., Guo, Z. D., Azar, M. G., Piot, B., Kavukcuoglu, K., Munos, R., and Valko, M · 2006
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Sagun, L., Bottou, L., and LeCun, Y · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts, 2017
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
signsgd: Compressed optimisation for non-convex problems
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A · 2018
Earlier work this paper cites.
Shampoo: Preconditioned stochastic tensor optimization
Gupta, V., Koren, T., and Singer, Y · 2018
Earlier work this paper cites.
Gradient descent happens in a tiny subspace
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Earlier work this paper cites.
An alternative view: When does sgd escape local minima?
Kleinberg, B., Li, Y., and Yuan, Y · 2018
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2019
Earlier work this paper cites.
Openwebtext corpus, 2019
Gokaslan, A. and Cohen, V · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
How noise affects the hessian spectrum in overparameterized neural networks
Wei, M. and Schwab, D. J · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., et al · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., et al · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., et al · 2021
Cited alongside, same era.
Muon: An optimizer for hidden layers in neural networks, 2024
Jordan, K., Jin, Y., Boza, V., Jiacheng, Y., Cecista, F., Newhouse, L., and Bernstein, J · 2024
Later among the works it cites.
Scalable optimization in the modular norm
Large, T., Liu, Y., Huh, M., Bahng, H., Isola, P., and Bernstein, J · 2024
Later among the works it cites.
Cautious optimizers: Improving training with one line of code
Liang, K., Chen, L., Liu, B., and Liu, Q · 2024
Later among the works it cites.
Complex fractal trainability boundary can arise from trivial non-convexity, 2024
Liu, Y · 2024
Later among the works it cites.
Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits
Morales Brotons, D., Vogels, T., and Hendrikx, H · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Github copilot: Your ai pair programmer, 2022
GitHub · 2022
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Rae, J. W., Lai, A., Wang, J., Millican, K., Young, S., Tieleman, O., et al · 2022
Cited alongside, same era.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Minervini, P., Andreas, J., et al · 2022
Cited alongside, same era.
Galactica: A large language model for science
Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., et al · 2022
Cited alongside, same era.
Rethinking learning rate tuning in the era of large language models, 2023
Jin, H., Wei, W., Wang, X., Zhang, W., and Wu, Y · 2023
Cited alongside, same era.
On Quantum Speedups for Nonconvex Optimization via Quantum Tunneling Walks
Liu, Y., Su, W. J., and Li, T · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
The ademamix optimizer: Better, faster, older
Pagliardini, M., Ablin, P., and Grangier, D · 2024
Later among the works it cites.
The boundary of neural network trainability is fractal, 2024
Sohl-Dickstein, J · 2024
Later among the works it cites.
Hop, skip, jump to convergence: Dynamics of learning rate transitions for improved training of large language models, 2024
Subramanian, S., Ganapathiraman, V., and Barrett, C · 2024
Later among the works it cites.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T · 2024
Later among the works it cites.
Soap: Improving and stabilizing shampoo using adam
Vyas, N., Morwani, D., Zhao, R., Shapira, I., Brandfonbrener, D., Janson, L., and Kakade, S · 2024
Later among the works it cites.
How to set adamw’s weight decay as you scale model and dataset size, 2024
Wang, X. and Aitchison, L · 2024
Later among the works it cites.
Understanding warmup-stable-decay learning rates: A river valley loss landscape perspective
Wen, K., Li, Z., Wang, J., Hall, D., Liang, P., and Ma, T · 2024
Later among the works it cites.
Mars: Unleashing the power of variance reduction for training large models
Yuan, H., Liu, Y., Wu, S., Zhou, X., and Gu, Q · 2024
Later among the works it cites.
Adam-mini: Use fewer learning rates to gain more, 2024
Zhang, Y., Chen, C., Li, Z., Ding, T., Wu, C., Kingma, D. P., Ye, Y., Luo, Z.-Q., and Sun, R · 2024
Later among the works it cites.
Deconstructing what makes a good optimizer for language models
Zhao, R., Morwani, D., Brandfonbrener, D., Vyas, N., and Kakade, S · 2024
Later among the works it cites.