Fetching the paper…
Reading the bibliography…
Robust and effective scaling of models from small to large width typically requires the precise adjustment of many algorithmic and architectural details, such as parameterization and optimizer choices.
Priors for infinite networks
Neal, R. M · 1996
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Google vizier: A service for black-box optimization
Golovin, D., Solnik, B., Moitra, S., Kochanski, G., Karro, J., and Sculley, D · 2017
Earlier work this paper cites.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S. S., Pennington, J., and Sohl-Dickstein, J · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L. and Bach, F · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2018
Earlier work this paper cites.
Gaussian process behaviour in wide deep neural networks
Matthews, A. G. d. G., Rowland, M., Hron, J., Turner, R. E., and Ghahramani, Z · 2018
Earlier work this paper cites.
An empirical model of large-batch training
McCandlish, S., Kaplan, J., Amodei, D., and Team, O. D · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P.-M · 2018
Earlier work this paper cites.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
Rotskoff, G. and Vanden-Eijnden, E · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Earlier work this paper cites.
A mean-field limit for certain deep neural networks
Araújo, D., Oliveira, R. I., and Yukimura, D · 2019
Earlier work this paper cites.
On empirical comparisons of optimizers for deep learning
Choi, D., Shallue, C. J., Nado, Z., Lee, J., Maddison, C. J., and Dahl, G. E · 2019
Cited alongside, same era.
Notes on contemporary machine learning for physicists, 2019
Kaplan, J · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2019
Cited alongside, same era.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Zhang, G., Li, L., Nado, Z., Martens, J., Sachdeva, S., Dahl, G., Shallue, C., and Grosse, R. B · 2019
Cited alongside, same era.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Later among the works it cites.
Meta-principled family of hyperparameter scaling strategies
Yaida, S · 2022
Later among the works it cites.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J · 2022
Later among the works it cites.
St-moe: Designing stable and transferable sparse expert models
Zoph, B., Bello, I., Kumar, S., Du, N., Huang, Y., Dean, J., Shazeer, N., and Fedus, W · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The DeepMind JAX Ecosystem, 2020
Babuschkin, I., Baumli, K., Bell, A., Bhupatiraju, S., Bruce, J., Buchlovsky, P., Budden, D., Cai, T., Clark, A., Danihelka, I., Dedieu, A., Fantacci, C., Godwin, J., Jones, C., Hemsley, R., Hennigan, T., Hessel, M., Hou, S., Kapturowski, S., Keck, T., Kemaev, I., King, M., Kunesch, M., Martens, L., Merzic, H., Mikulik, V., Norman, T., Papamakarios, G., Quan, J., Ring, R., Ruiz, F., Sanchez, A., Sartran, L., Schneider, R., Sezener, E., Spencer, S., Srinivasan, S., Stanojević, M., Stokowiec, W., Wang, L., Zhou, G., and Viola, F · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Disentangling feature and lazy training in deep neural networks
Geiger, M., Spigler, S., Jacot, A., and Wyart, M · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Mesh TensorFlow - Model Parallelism Made Easier; comments on unit scaling convention, 12 2020
Shazeer, N · 2020
Cited alongside, same era.
Blake, C., Orr, D., and Luschi, C · 2023
Later among the works it cites.
Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit
Bordelon, B., Noci, L., Li, M. B., Hanin, B., and Pehlevan, C · 2023
Later among the works it cites.
Steering deep feature learning with backward aligned feature updates
Chizat, L. and Netrapalli, P · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., et al · 2023
Later among the works it cites.
Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster
Dey, N., Gosal, G., Khachane, H., Marshall, W., Pathria, R., Tom, M., Hestness, J., et al · 2023
Later among the works it cites.
Effective theory of transformers at initialization
Dinan, E., Yaida, S., and Zhang, S · 2023
Later among the works it cites.
Flax: A neural network library and ecosystem for JAX, 2023
Heek, J., Levskaya, A., Oliver, A., Ritter, M., Rondepierre, B., Steiner, A., and van Zee, M · 2023
Later among the works it cites.
On the parameterization of second-order optimization effective towards the infinite width
Ishikawa, S. and Karakida, R · 2023
Later among the works it cites.
Small-scale proxies for large-scale transformer training instabilities
Wortsman, M., Liu, P. J., Xiao, L., Everett, K., Alemi, A., Adlam, B., Co-Reyes, J. D., Gur, I., Kumar, A., Novak, R., et al · 2023
Later among the works it cites.
Tensor programs ivb: Adaptive optimization in the infinite-width limit
Yang, G. and Littwin, E · 2023
Later among the works it cites.
Stabilizing transformer training by preventing attention entropy collapse
Zhai, S., Likhomanenko, T., Littwin, E., Busbridge, D., Ramapuram, J., Zhang, Y., Gu, J., and Susskind, J. M · 2023
Later among the works it cites.
Deepseek llm: Scaling open-source language models with longtermism
Bi, X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., et al · 2024
Closest in time.
Scalable optimization in the modular norm
Large, T., Liu, Y., Huh, M., Bahng, H., Isola, P., and Bernstein, J · 2024
Closest in time.
A large-scale exploration of μ \mu -transfer
Lingle, L · 2024
Closest in time.
Nanodo: A minimal transformer decoder-only language model implementation in JAX., 2024
Liu, P. J., Novak, R., Lee, J., Wortsman, M., Xiao, L., Everett, K., Alemi, A. A., Kurzeja, M., Marcenac, P., Gur, I., Kornblith, S., Xu, K., Elsayed, G., Fischer, I., Pennington, J., Adlam, B., and Dickstein, J.-S · 2024
Closest in time.