Fetching the paper…
Reading the bibliography…
In deep learning theory, the covariance matrix of the representations serves as a proxy to examine the network's trainability.
“Language models are few-shot learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“Finite depth and width corrections to the neural tangent kernel”, 2019
Boris Hanin and Mihai Nica · 1909
Earlier work this paper cites.
“Bayesian learning for neural networks”, 1995
Radford Neal · 1995
Earlier work this paper cites.
“Scaling laws for neural language models”, 2020
Jared Kaplan et al · 2001
Earlier work this paper cites.
“Feature learning in infinite-width neural networks”, 2020
Greg Yang and Edward Hu · 2011
Earlier work this paper cites.
“Brownian motion and stochastic calculus”
Ioannis Karatzas and Steven Shreve · 2012
Earlier work this paper cites.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
“Lyapunov exponents for products of rectangular real, complex and quaternionic Ginibre matrices”
JR Ipsen · 2015
Earlier work this paper cites.
“Delving deep into rectifiers: Surpassing human-level performance on imagenet classification”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2015
Earlier work this paper cites.
“Stochastic Calculus (Lecture Notes)”
Jason Miller · 2015
Earlier work this paper cites.
“Exponential expressivity in deep neural networks through transient chaos”
Ben Poole et al · 2016
Earlier work this paper cites.
Jimmy Ba, Jamie Kiros and Geoffrey Hinton · 2016
Earlier work this paper cites.
“Residual networks behave like ensembles of relatively shallow networks”
Andreas Veit, Michael Wilber and Serge Belongie · 2016
Earlier work this paper cites.
“Attention is all you need”, 2017
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Deep learning scaling is predictable, empirically”, 2017
Joel Hestness et al · 2017
Earlier work this paper cites.
“Deep Information Propagation”
Samuel. Schoenholz, Justin Gilmer, Surya Ganguli and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
“Mean field residual networks: on the edge of chaos”
Greg Yang and Samuel Schoenholz · 2017
Earlier work this paper cites.
“Spectral radii of large non-Hermitian random matrices”
Tiefeng Jiang and Yongcheng Qi · 2017
Earlier work this paper cites.
“Deep Neural Networks as Gaussian Processes”
Jaehoon Lee et al · 2018
Earlier work this paper cites.
“Gaussian process behaviour in wide deep neural networks”, 2018
Alexander Matthews et al · 2018
Earlier work this paper cites.
“Neural tangent kernel: Convergence and generalization in neural networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Earlier work this paper cites.
“A mean field view of the landscape of two-layer neural networks”
Song Mei, Andrea Montanari and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
“On the global convergence of gradient descent for over-parameterized models using optimal transport”
Lenaic Chizat and Francis Bach · 2018
Earlier work this paper cites.
“Trainability and Accuracy of Neural Networks: An Interacting Particle System Approach”, 2018
Grant. Rotskoff and Eric Vanden-Eijnden · 2018
Earlier work this paper cites.
“GLUE: A multi-task benchmark and analysis platform for natural language understanding”
Alex Wang et al · 2018
Cited alongside, same era.
“On the impact of the activation function on deep neural networks training”
Soufiane Hayou, Arnaud Doucet and Judith Rousseau · 2019
Cited alongside, same era.
“Asymptotics of wide networks from feynman diagrams”
Ethan Dyer and Guy Gur-Ari · 2019
Cited alongside, same era.
“Products of many large random matrices and gradients in deep neural networks”
Boris Hanin and Mihai Nica · 2019
Cited alongside, same era.
“From integrable to chaotic systems: Universal local statistics of Lyapunov exponents”
Gernot Akemann, Zdzislaw Burda and Mario Kieburg · 2019
Cited alongside, same era.
“Foundations of Modern Probability”, Probability theory and stochastic modelling
O. Kallenberg · 2021
Later among the works it cites.
“Training compute-optimal large language models”, 2022
Jordan Hoffmann et al · 2022
Later among the works it cites.
“Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?”, 2022
Yi Tay et al · 2022
Later among the works it cites.
“Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank Collapse”, 2022
Lorenzo Noci et al · 2022
Later among the works it cites.
“Activation function design for deep networks: linearity and effective initialisation”
Michael Murray, Vinayak Abrol and Jared Tanner · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Dynamical isometry is achieved in residual networks in a universal way for any activation function”
Wojciech Tarnowski et al · 2019
Cited alongside, same era.
“Disentangling trainability and generalization in deep neural networks”
Lechao Xiao, Jeffrey Pennington and Samuel Schoenholz · 2020
Cited alongside, same era.
“On layer normalization in the transformer architecture”
Ruibin Xiong et al · 2020
Cited alongside, same era.
“Infinite attention: NNGP and NTK for deep attention networks”
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein and Roman Novak · 2020
Cited alongside, same era.
“Mean field analysis of neural networks: A law of large numbers”
Justin Sirignano and Konstantinos Spiliopoulos · 2020
Cited alongside, same era.
“Non-Gaussian processes and neural networks at finite widths”
Sho Yaida · 2020
Cited alongside, same era.
“Tensor programs ii: Neural tangent kernel for any architecture”
Greg Yang · 2020
Cited alongside, same era.
“Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers”, 2022
Guodong Zhang, Aleksandar Botev and James Martens · 2022
Later among the works it cites.
“The neural covariance SDE: Shaped infinite depth-and-width networks at initialization”
Mufan Li, Mihai Nica and Dan Roy · 2022
Later among the works it cites.
“Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer”, 2022
Greg Yang et al · 2022
Later among the works it cites.
“The principles of deep learning theory”
Daniel Roberts, Sho Yaida and Boris Hanin · 2022
Later among the works it cites.
“Correlation Functions in Random Fully Connected Neural Networks at Finite Width”, 2022
Boris Hanin · 2022
Later among the works it cites.
“On the infinite-depth limit of finite-width neural networks”, 2022
Soufiane Hayou · 2022
Later among the works it cites.
“The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks”
Emmanuel Abbe, Enric Adsera and Theodor Misiakiewicz · 2022
Later among the works it cites.
“High-dimensional asymptotics of feature learning: How one gradient step improves the representation”
Jimmy Ba et al · 2022
Later among the works it cites.
“Neural networks can learn representations with gradient descent”
Alexandru Damian, Jason Lee and Mahdi Soltanolkotabi · 2022
Later among the works it cites.
“Neural Networks Efficiently Learn Low-Dimensional Representations with SGD”, 2022
Alireza Mousavi-Hosseini et al · 2022
Later among the works it cites.
“Scaling Laws for Generative Mixed-Modal Language Models”, 2023
Armen Aghajanyan et al · 2023
Closest in time.
“Effective Theory of Transformers at Initialization”, 2023
Emily Dinan, Sho Yaida and Susan Zhang · 2023
Closest in time.
Bobby He et al · 2023
Closest in time.
“Scaling vision transformers to 22 billion parameters”, 2023
Mostafa Dehghani et al · 2023
Closest in time.
“Width and Depth Limits Commute in Residual Networks”, 2023
Soufiane Hayou and Greg Yang · 2023
Closest in time.
“Stabilizing Transformer Training by Preventing Attention Entropy Collapse”, 2023
Shuangfei Zhai et al · 2023
Closest in time.
“SGD learning on neural networks: leap complexity and saddle-to-saddle dynamics”, 2023
Emmanuel Abbe, Enric Boix-Adsera and Theodor Misiakiewicz · 2023
Closest in time.
“Learning time-scales in two-layers neural networks”, 2023
Raphaël Berthier, Andrea Montanari and Kangjie Zhou · 2023
Closest in time.