Fetching the paper…
Reading the bibliography…
Epoch-wise double descent is the phenomenon where generalisation performance improves beyond the point of overfitting, resulting in a generalisation curve exhibiting two descents under the course of learning.
Temporal Evolution of Generalization during Learning in Linear Networks
P. Baldi and Y. Chauvin · 1991
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2014
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P.-M. Nguyen · 2018
Earlier work this paper cites.
Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
S. Arora, S. S. Du, W. Hu, Z. Li, and R. Wang · 2019
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias-variance tradeoff
M. Belkin, D. Hsu, S. Ma, and S. Mandal · 2019
Earlier work this paper cites.
Implicit Regularization of Discrete Gradient Dynamics in Linear Neural Networks
G. Gidel, F. Bach, and S. Lacoste-Julien · 2019
Earlier work this paper cites.
Generalisation dynamics of online learning in over-parameterised neural networks
S. Goldt, M. S. Advani, A. M. Saxe, F. Krzakala, and L. Zdeborová · 2019
Earlier work this paper cites.
An analytic theory of generalization dynamics and transfer learning in deep linear networks
A. K. Lampinen and S. Ganguli · 2019
Earlier work this paper cites.
A mathematical theory of semantic development in deep neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2019
Earlier work this paper cites.
Understanding Double Descent Requires a Fine-Grained Bias-Variance Decomposition
B. Adlam and J. Pennington · 2020
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks
M. S. Advani, A. M. Saxe, and H. Sompolinsky · 2020
Earlier work this paper cites.
Triple Descent and the Two Kinds of Overfitting: Where & Why do they Appear?
S. D’Ascoli, L. Sagun, and G. Biroli · 2020
Earlier work this paper cites.
The implicit bias of depth: How incremental learning drives generalization
D. Gissin, S. Shalev-Shwartz, and A. Daniely · 2020
Earlier work this paper cites.
Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
S. Goldt, M. S. Advani, A. M. Saxe, F. Krzakala, and L. Zdeborová · 2020
Cited alongside, same era.
Model, sample, and epoch-wise descents: exact solution of gradient flow in the random feature model
A. Bodin and N. Macris · 2021
Cited alongside, same era.
Early stopping in deep networks: double descent and how to eliminate it
R. Heckel and F. F. Yilmaz · 2021
Cited alongside, same era.
Implicit Sparse Regularization: The Impact of Depth and Early Stopping
J. Li, T. V. Nguyen, C. Hegde, and R. K. Wong · 2021
Cited alongside, same era.
On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear Networks
H. Min, S. Tarmoun, R. Vidal, and E. Mallada · 2021
Cited alongside, same era.
Exact learning dynamics of deep linear networks with prior knowledge
L. Braun, C. C. J. Dominé, J. E. Fitzgerald, and A. M. Saxe · 2022
Later among the works it cites.
Surprises in High-Dimensional Ridgeless Least Squares Interpolation
T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani · 2022
Later among the works it cites.
Multi-scale Feature Learning Dynamics: Insights for Double Descent
M. Pezeshki, A. Mitra, Y. Bengio, and G. Lajoie · 2022
Later among the works it cites.
On the generalization of models trained With SGD: Information-Theoretic Bounds and Implications
Z. Wang and Y. Mao · 2022
Later among the works it cites.
Incremental Learning in Diagonal Linear Networks
R. Berthier · 2023
Later among the works it cites.
Deep Linear Networks can Benignly Overfit when Shallow Ones Do
N. S. Chatterji and P. M. Long · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever · 2021
Cited alongside, same era.
Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of Stochasticity
S. Pesme, L. Pillaud-Vivien, and N. Flammarion · 2021
Cited alongside, same era.
When and how epochwise double descent happens
C. Stephenson and T. Lee · 2021
Cited alongside, same era.
Understanding the Dynamics of Gradient Flow in Overparameterized Linear Models
S. Tarmoun, G. França, B. Haeffele, and R. Vidal · 2021
Cited alongside, same era.
A unifying view on implicit bias in training linear neural networks
C. Yun, S. Krishnan, and H. Mobahi · 2021
Cited alongside, same era.
Neural networks as kernel learners: the silent alignment effect
A. Atanasov, B. Bordelon, and C. Pehlevan · 2022
Cited alongside, same era.
A. Bodin and N. Macris · 2022
Cited alongside, same era.
Later among the works it cites.
A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning
A. Curth, A. Jeffares, and M. van der Schaar · 2023
Later among the works it cites.
(S) GD over Diagonal Linear Networks: Implicit Bias, Large Stepsizes and Edge of Stability
M. Even, S. Pesme, S. Gunasekar, and N. Flammarion · 2023
Later among the works it cites.
Implicit regularization for group sparsity
J. Li, T. V. Nguyen, C. Hegde, and R. K. W. Wong · 2023
Later among the works it cites.
On the Convergence of Gradient Flow on Multi-layer Linear Models
H. Min, R. Vidal, and E. Mallada · 2023
Later among the works it cites.
Saddle-to-Saddle Dynamics in Diagonal Linear Networks
S. Pesme and N. Flammarion · 2023
Later among the works it cites.
Linear Convergence of Gradient Descent for Finite Width Over-parametrized Linear Networks with General Initialization
Z. Xu, H. Min, S. Tarmoun, E. Mallada, and R. Vidal · 2023
Later among the works it cites.
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
D. Kunin, A. Raventós, C. Dominé, F. Chen, D. Klindt, A. Saxe, and S. Ganguli · 2024
Closest in time.