Fetching the paper…
Reading the bibliography…
A wide range of empirical and theoretical works have shown that overparameterisation can amplify the performance of neural networks.
The transfer of a discrimination along a continuum
Lawrence, D. H · 1952
Earlier work this paper cites.
Discrimination transfer along a pitch continuum
Baker, R. A. and Osgood, S. W · 1954
Earlier work this paper cites.
Science and human behavior
Skinner, B. F · 1965
Earlier work this paper cites.
Task difficulty and the specificity of perceptual learning
Ahissar, M. and Hochstein, S · 1997
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism
Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G · 2003
Earlier work this paper cites.
The reverse hierarchy theory of visual perceptual learning
Ahissar, M. and Hochstein, S · 2004
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
When does fading enhance perceptual category learning?
Pashler, H. and Mozer, M. C · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L. and Bach, F · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P.-M · 2018
Cited alongside, same era.
Curriculum learning by transfer learning: Theory and experiments with deep networks
Weinshall, D., Cohen, G., and Amir, D · 2018
Cited alongside, same era.
Gradient descent with random initialization: Fast global convergence for nonconvex phase retrieval
Chen, Y., Chi, Y., Fan, J., and Ma, C · 2019
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
Kornblith, S., Norouzi, M., Lee, H., and Hinton, G · 2019
Cited alongside, same era.
Theory of curriculum learning, with convex loss functions
Weinshall, D. and Amir, D · 2020
Curriculum learning by optimizing learning dynamics
Zhou, T., Wang, S., and Bilmes, J · 2021
Later among the works it cites.
Trainability and accuracy of artificial neural networks: An interacting particle system approach
Rotskoff, G. and Vanden-Eijnden, E · 2022
Later among the works it cites.
An analytical theory of curriculum learning in teacher-student networks
Saglietti, L., Mannelli, S., and Saxe, A · 2022
Later among the works it cites.
Bias-inducing geometries: exactly solvable data model with fairness implications
Sarao Mannelli, S., Gerace, F., Rostamzadeh, N., and Saglietti, L · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Sorscher, B., Geirhos, R., Shekhar, S., Ganguli, S., and Morcos, A · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The staircase property: How hierarchical structure can guide deep learning
Abbe, E., Boix-Adsera, E., Brennan, M. S., Bresler, G., and Nagaraj, D · 2021
Cited alongside, same era.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Loureiro, B., Gerbelot, C., Cui, H., Goldt, S., Krzakala, F., Mezard, M., and Zdeborová, L · 2021
Cited alongside, same era.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Refinetti, M., Goldt, S., Krzakala, F., and Zdeborová, L · 2021
Cited alongside, same era.
When do curricula work?
Wu, X., Dyer, E., and Neyshabur, B · 2021
Cited alongside, same era.
Procrustean training for imbalanced deep learning
Ye, H.-J., Zhan, D.-C., and Chao, W.-L · 2021
Cited alongside, same era.
Complex dynamics in simple neural networks: Understanding gradient flow in phase retrieval
Sarao Mannelli, S., Biroli, G., Cammarota, C., Krzakala, F., Urbani, P., and Zdeborová, L
Cited in the paper.
Soviany, P., Ionescu, R. T., Rota, P., and Sebe, N · 2022
Later among the works it cites.
High-dimensional limit theorems for sgd: Effective dynamics and critical scaling
Ben Arous, G., Gheissari, R., and Jagannath, A · 2023
Later among the works it cites.
A mathematical model for curriculum learning for parities
Cornacchia, E. and Mossel, E · 2023
Later among the works it cites.
On the impact of machine learning randomness on group fairness
Ganesh, P., Chang, H., Strobel, M., and Shokri, R · 2023
Later among the works it cites.
Adaptive algorithms for shaping behavior
Tong, W. L., Iyer, A., Murthy, V. N., and Reddy, G · 2023
Later among the works it cites.
Provable advantage of curriculum learning on parity targets with mixed inputs
Abbe, E., Cornacchia, E., and Lotfi, A · 2024
Closest in time.