Fetching the paper…
Reading the bibliography…
A simple design recipe for deep Transformers is to compose identical building blocks.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Earlier work this paper cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
David Balduzzi, Marcus Frean, Lennox Leary, JP Lewis, Kurt Wan-Duo Ma, and Brian McWilliams · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Deep information propagation
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick · 2018
Earlier work this paper cites.
Deep Neural Networks as Gaussian Processes
Jaehoon Lee, Jascha Sohl-dickstein, Jeffrey Pennington, Roman Novak, Sam Schoenholz, and Yasaman Bahri · 2018
Earlier work this paper cites.
Gaussian Process Behaviour in Wide Deep Neural Networks
Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani · 2018
Earlier work this paper cites.
Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz, and Jeffrey Pennington · 2018
Earlier work this paper cites.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N Dauphin, and Tengyu Ma · 2018
Earlier work this paper cites.
How to initialize your network? robust initialization for weightnorm & resnets
Devansh Arpit, Víctor Campos, and Yoshua Bengio · 2019
Earlier work this paper cites.
On the impact of the activation function on deep neural networks training
Soufiane Hayou, Arnaud Doucet, and Judith Rousseau · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Augmenting self-attention with persistent memory
Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin · 2019
Earlier work this paper cites.
Dynamical isometry is achieved in residual networks in a universal way for any activation function
Wojciech Tarnowski, Piotr Warchoł, Stanisław Jastrzebski, Jacek Tabor, and Maciej Nowak · 2019
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Wide feedforward or recurrent neural networks of any architecture are gaussian processes
Greg Yang · 2019
Earlier work this paper cites.
A mean field theory of batch normalization
Greg Yang, Jeffrey Pennington, Vinay Rao, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2019
Cited alongside, same era.
Root mean square layer normalization
Biao Zhang and Rico Sennrich · 2019
Cited alongside, same era.
Improving deep transformer with depth-scaled initialization and merged attention
Biao Zhang, Ivan Titov, and Rico Sennrich · 2019
Cited alongside, same era.
Muse: Parallel multi-scale attention for sequence to sequence learning
Guangxiang Zhao, Xu Sun, Jingjing Xu, Zhiyuan Zhang, and Liangchen Luo · 2019
Cited alongside, same era.
Batch normalization biases residual blocks towards the identity function in deep networks
Soham De and Sam Smith · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Going deeper with image transformers
Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou · 2021
Later among the works it cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Later among the works it cites.
The neural covariance sde: Shaped infinite depth-and-width networks at initialization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Cited alongside, same era.
Provable benefit of orthogonal initialization in optimizing deep linear networks
Wei Hu, Lechao Xiao, and Jeffrey Pennington · 2020
Cited alongside, same era.
Improving transformer optimization through better initialization
Xiao Shi Huang, Felipe Perez, Jimmy Ba, and Maksims Volkovs · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
Understanding the difficulty of training transformers
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han · 2020
Cited alongside, same era.
Disentangling trainability and generalization in deep neural networks
Lechao Xiao, Jeffrey Pennington, and Samuel Schoenholz · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Cited alongside, same era.
Mufan Bill Li, Mihai Nica, and Daniel M Roy · 2022
Later among the works it cites.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto, Sidak Pal Singh, and Aurelien Lucchi · 2022
Later among the works it cites.
Introducing nvFuser, a deep learning compiler for PyTorch
Christian Sarofeen, Piotr Bialecki, Jie Jiang, Kevin Stephano, Masaki Kozuki, Neal Vaidya, and Stas. Bekman · 2022
Later among the works it cites.
Trimbert: Tailoring bert for trade-offs
Sharath Nittur Sridhar, Anthony Sarah, and Sairam Sundaresan · 2022
Later among the works it cites.
Deep learning without shortcuts: Shaping the kernel with tailored rectifiers
Guodong Zhang, Aleksandar Botev, and James Martens · 2022
Later among the works it cites.
Centered self-attention layers
Ameen Ali, Tomer Galanti, and Lior Wolf · 2023
Closest in time.
Scaling vision transformers to 22 billion parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Peter Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al · 2023
Closest in time.
Re-parameterizing your optimizers rather than architectures
Xiaohan Ding, Honghao Chen, Xiangyu Zhang, Kaiqi Huang, Jungong Han, and Guiguang Ding · 2023
Closest in time.
Cramming: Training a language model on a single GPU in one day, 2023
Jonas Geiping and Tom Goldstein · 2023
Closest in time.
Width and depth limits commute in residual networks
Soufiane Hayou and Greg Yang · 2023
Closest in time.
Deep transformers without shortcuts: Modifying self-attention for faithful signal propagation
Bobby He, James Martens, Guodong Zhang, Aleksandar Botev, Andrew Brock, Samuel L Smith, and Yee Whye Teh · 2023
Closest in time.
Towards training without depth limits: Batch normalization without gradient explosion
Alexandru Meterez, Amir Joudaki, Francesco Orabona, Alexander Immer, Gunnar Rätsch, and Hadi Daneshmand · 2023
Closest in time.
The shaped transformer: Attention models in the infinite depth-and-width limit
Lorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He, Thomas Hofmann, Chris J. Maddison, and Daniel M. Roy · 2023
Closest in time.
One wide feedforward is all you need
Telmo Pessoa Pires, António V Lopes, Yannick Assogba, and Hendra Setiawan · 2023
Closest in time.
The languini kitchen: Enabling language modelling research at different scales of compute
Aleksandar Stanić, Dylan Ashley, Oleg Serikov, Louis Kirsch, Francesco Faccio, Jürgen Schmidhuber, Thomas Hofmann, and Imanol Schlag · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Mimetic initialization of self-attention layers
Asher Trockman and J Zico Kolter · 2023
Closest in time.
Pre-training via denoising for molecular property prediction
Sheheryar Zaidi, Michael Schaarschmidt, James Martens, Hyunjik Kim, Yee Whye Teh, Alvaro Sanchez-Gonzalez, Peter Battaglia, Razvan Pascanu, and Jonathan Godwin · 2023
Closest in time.