Fetching the paper…
Reading the bibliography…
We perform an effective-theory analysis of forward-backward signal propagation in wide and deep Transformers, i.e., residual neural networks with multi-head self-attention blocks and multilayer perceptron blocks.
1907
Earlier work this paper cites.
B. T. Polyak and A. B. Juditsky, “Acceleration of Stochastic Approximation by Averaging,” SIAM Journal on Control and Optimization
1992
Earlier work this paper cites.
Springer, 1996
R. M. Neal, “Priors for Infinite Networks,” in Bayesian Learning for Neural Networks · 1996
Earlier work this paper cites.
N. Shazeer, “GLU Variants Improve Transformer,” arXiv:2002.05202 [cs.CL]
2002
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al
2015
Cited alongside, same era.
Ravenio Books, 2016
G. K. Zipf, Human Behavior and the Principle of Least Effort: An Introduction to Human Ecology · 2016
Cited alongside, same era.
http://commoncrawl.org/2016/10/news-dataset-available
S. Nagel, “CC-News,” 2016 · 2016
Cited alongside, same era.
R. Wightman et al
2019
Cited alongside, same era.
Cited in the paper.
Cited in the paper.
Cited in the paper.
Cited in the paper.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer Normalization,” arXiv:1607.06450 [stat.ML]
Cited in the paper.
Cited in the paper.
S. Yaida, “Meta-Principled Family of Hyperparameter Scaling Strategies,” arXiv:2210.04909 [cs.LG]
Cited in the paper.
Cited in the paper.
http://Skylion007.github.io/OpenWebTextCorpus
A. Gokaslan and V. Cohen, “Openwebtext corpus,” 2019 · 2019
Later among the works it cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al
2019
Later among the works it cites.
Cambridge University Press, 2022
D. A. Roberts, S. Yaida, and B. Hanin, The Principles of Deep Learning Theory · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…