Fetching the paper…
Reading the bibliography…
This work reports deep-learning-unique first-order and second-order phase transitions, whose phenomenology closely follows that in statistical physics.
Diagnosing and enhancing vae models
Dai, B. and Wipf, D. (2019) · 1903
Earlier work this paper cites.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J. (2019) · 1903
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
Hopfield, J. J. (1982) · 1982
Earlier work this paper cites.
Learnability with respect to fixed distributions
Benedek, G. M. and Itai, A. (1991) · 1991
Earlier work this paper cites.
The statistical mechanics of learning a rule
Watkin, T. L., Rau, A., and Biehl, M. (1993) · 1993
Earlier work this paper cites.
Rigorous learning curve bounds from statistical mechanics
Haussler, D., Kearns, M., Seung, H. S., and Tishby, N. (1996) · 1996
Earlier work this paper cites.
Estimation of dependences based on empirical data
Vapnik, V. (2006) · 2006
Earlier work this paper cites.
The free-energy principle: a rough guide to the brain?
Friston, K. (2009) · 2009
Earlier work this paper cites.
Identity matters in deep learning
Hardt, M. and Ma, T. (2016) · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K. (2016) · 2016
Cited alongside, same era.
Martin, C. H. and Mahoney, M. W. (2017) · 2017
Cited alongside, same era.
Fixing a broken elbo
Alemi, A., Poole, B., Fischer, I., Dillon, J., Saurous, R. A., and Murphy, K. (2018) · 2018
Cited alongside, same era.
Deep linear networks with arbitrary loss: All local minima are global
Laurent, T. and Brecht, J. (2018) · 2018
Cited alongside, same era.
Don’t blame the elbo! a linear vae perspective on posterior collapse
Transport analysis of infinitely deep neural network
Sonoda, S. and Murata, N. (2019) · 2019
Later among the works it cites.
Statistical mechanics of deep learning
Bahri, Y., Kadmon, J., Pennington, J., Schoenholz, S. S., Sohl-Dickstein, J., and Ganguli, S. (2020) · 2020
Later among the works it cites.
Two models of double descent for weak features
Belkin, M., Hsu, D., and Xu, J. (2020) · 2020
Later among the works it cites.
A random matrix analysis of random fourier features: beyond the gaussian kernel, a precise phase transition, and the corresponding double descent
Liao, Z., Couillet, R., and Mahoney, M. W. (2020) · 2020
Later among the works it cites.
Statistical mechanics of deep linear neural networks: The backpropagating kernel renormalization
Li, Q. and Sompolinsky, H. (2021) · 2021
Later among the works it cites.
Noether’s learning dynamics: Role of symmetry breaking in neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lucas, J., Tucker, G., Grosse, R., and Norouzi, M. (2019) · 2019
Cited alongside, same era.
On dropout and nuclear norm regularization
Mianjy, P. and Arora, R. (2019) · 2019
Cited alongside, same era.
Generalization in a linear perceptron in the presence of noise
Krogh, A. and Hertz, J. A. (1992a)
Cited in the paper.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J. A. (1992b)
Cited in the paper.
Tanaka, H. and Kunin, D. (2021) · 2021
Later among the works it cites.
Exact solutions of a deep linear network
Ziyin, L., Li, B., and Meng, X. (2022) · 2022
Closest in time.