Fetching the paper…
Reading the bibliography…
Our understanding of learning dynamics of deep neural networks (DNNs) remains incomplete.
Generalisation dynamics of online learning in over-parameterised neural networks
Sebastian Goldt, Madhu S. Advani, Andrew M. Saxe, Florent Krzakala, and Lenka Zdeborová · 1901
Earlier work this paper cites.
Learning representations by back–propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Inequalities for the singular values of Hadamard products
Xingzhi Zhan · 1997
Earlier work this paper cites.
Using algebraic geometry , volume 185
David A. Cox, John Little, and Donal O’shea · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey E. Hinton, et al · 2009
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
The wikitext long term dependency language modeling dataset
Stephen Merity · 2016
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
The full spectrum of deepnet Hessians at scale: Dynamics with SGD training and sample size
Vardan Papyan · 2018
Cited alongside, same era.
Linearized sigmoidal activation: A novel activation function with tractable non-linear characteristics to boost representation capability
Vivek Singh Bawa and Vinay Kumar · 2019
Cited alongside, same era.
Singular values for ReLU layers
Sören Dittmer, Emily J. King, and Peter Maass · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau · 2019
Prevalence of neural collapse during the terminal phase of deep learning training
Vardan Papyan, X. Y. Han, and David L. Donoho · 2020
Later among the works it cites.
Peering beyond the gradient veil with distributed auto differentiation
Bradley T. Baker, Aashis Khanal, Vince D. Calhoun, Barak A. Pearlmutter, and Sergey M. Plis · 2021
Later among the works it cites.
On the role of neural collapse in transfer learning
Tomer Galanti, András György, and Marcus Hutter · 2021
Later among the works it cites.
Neural collapse under MSE loss: Proximity to and dynamics on the central path
X. Y. Han, Vardan Papyan, and David L. Donoho · 2021
Later among the works it cites.
A geometric analysis of neural collapse with unconstrained features
Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A mathematical theory of semantic development in deep neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2019
Cited alongside, same era.
PowerSGD: Practical low-rank gradient compression for distributed optimization
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi · 2019
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S. Advani, Andrew M. Saxe, and Haim Sompolinsky · 2020
Cited alongside, same era.
Batch normalization provably avoids ranks collapse for randomly initialised deep networks
Hadi Daneshmand, Jonas Kohler, Francis Bach, Thomas Hofmann, and Aurelien Lucchi · 2020
Cited alongside, same era.
Neural collapse with cross-entropy loss
Jianfeng Lu and Stefan Steinerberger · 2020
Cited alongside, same era.
Neural collapse with unconstrained features
Dustin G. Mixon, Hans Parshall, and Jianzong Pi · 2020
Cited alongside, same era.
Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
Sebastian Goldt, Madhu Advani, Andrew M. Saxe, Florent Krzakala, and Lenka Zdeborová
Cited in the paper.
Later among the works it cites.
The neural race reduction: dynamics of abstraction in gated networks
Andrew Saxe, Shagun Sodhani, and Sam Jay Lewallen · 2022
Later among the works it cites.
Imbalance trouble: Revisiting neural-collapse geometry
Christos Thrampoulidis, Ganesh Ramachandra Kini, Vala Vakilian, and Tina Behnia · 2022
Later among the works it cites.
Extended unconstrained features model for exploring deep neural collapse
Tom Tirer and Joan Bruna · 2022
Later among the works it cites.
Neural collapse with normalized features: A geometric analysis over the riemannian manifold
Can Yaras, Peng Wang, Zhihui Zhu, Laura Balzano, and Qing Qu · 2022
Later among the works it cites.
Neural collapse: A review on modelling principles and generalization
Vignesh Kothapalli · 2023
Later among the works it cites.