Fetching the paper…
Reading the bibliography…
By classifying infinite-width neural networks and identifying the *optimal* limit, Tensor Programs IV and V demonstrated a universal way, called $\mu$P, for *widthwise hyperparameter transfer*, i.e., predicting optimal hyperparameters of wide neural networks from narrow ones.
Highway networks, 2015
R. K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Earlier work this paper cites.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Attention is all you need, 2017
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport, 2018
L. Chizat and F. Bach · 2018
Earlier work this paper cites.
How to start training: The effect of initialization and architecture, 2018
B. Hanin and D. Rolnick · 2018
Earlier work this paper cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks, 2018
D. Zou, Y. Cao, D. Zhou, and Q. Gu · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization, 2019
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Earlier work this paper cites.
Fixup initialization: Residual learning without normalization, 2019
H. Zhang, Y. N. Dauphin, and T. Ma · 2019
Cited alongside, same era.
On lazy training in differentiable programming, 2020
L. Chizat, E. Oyallon, and F. Bach · 2020
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks, 2020
A. Jacot, F. Gabriel, and C. Hongler · 2020
Cited alongside, same era.
Understanding the difficulty of training transformers
L. Liu, X. Liu, J. Gao, W. Chen, and J. Han · 2020
Cited alongside, same era.
Stable resnet
S. Hayou, E. Clerico, B. He, G. Deligiannidis, A. Doucet, and J. Rousseau · 2021
Cited alongside, same era.
Tensor programs i: Wide feedforward or recurrent neural networks of any architecture are gaussian processes, 2021
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
G. Yang, E. J. Hu, I. Babuschkin, S. Sidor, X. Liu, D. Farhi, N. Ryder, J. Pachocki, W. Chen, and J. Gao · 2022
Later among the works it cites.
On the infinite-depth limit of finite-width neural networks
S. Hayou · 2023
Closest in time.
Width and depth limits commute in residual networks
S. Hayou and G. Yang · 2023
Closest in time.
Depth dependence of μ \mu p learning rates in relu mlps, 2023
S. Jelassi, B. Hanin, Z. Ji, S. J. Reddi, S. Bhojanapalli, and S. Kumar · 2023
Closest in time.
The shaped transformer: Attention models in the infinite depth-and-width limit, 2023
L. Noci, C. Li, M. B. Li, B. He, T. Hofmann, C. Maddison, and D. M. Roy · 2023
Closest in time.
Gpt-4 technical report, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Yang · 2021
Cited alongside, same era.
Tensor programs iv: Feature learning in infinite-width neural networks
G. Yang and E. J. Hu · 2021
Cited alongside, same era.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse, 2022
L. Noci, S. Anagnostidis, L. Biggio, A. Orvieto, S. P. Singh, and A. Lucchi · 2022
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun
Cited in the paper.
Scaling limits of wide neural networks with weight sharing: Gaussian process behavior, gradient independence, and neural tangent kernel derivation, 2020a
G. Yang
Cited in the paper.
Tensor programs ii: Neural tangent kernel for any architecture, 2020b
G. Yang
Cited in the paper.
OpenAI · 2023
Closest in time.
Tensor programs ivb: Adaptive optimization in the infinite-width limit, 2023
G. Yang and E. Littwin · 2023
Closest in time.
Stabilize deep resnet with a sharp scaling factor τ \tau , 2023
H. Zhang, D. Yu, M. Yi, W. Chen, and T.-Y. Liu · 2023
Closest in time.