Fetching the paper…
Reading the bibliography…
The main goal of this thesis is to point out that the bias-variance tradeoff is not always true (e.g.
Scaling description of generalization with number of parameters in deep learning
Geiger, M., Jacot, A., Spigler, S., Gabriel, F., Sagun, L., d’Ascoli, S., Biroli, G., Hongler, C., and Wyart, M · 1901
Earlier work this paper cites.
On empirical spectral analysis of stochastic processes
Grenander, U · 1952
Earlier work this paper cites.
A completely automatic french curve : fitting spline functions by cross validation
Wahba, G. and Wold, S · 1975
Earlier work this paper cites.
Bootstrap methods : Another look at the jackknife
Efron, B · 1979
Earlier work this paper cites.
Réseaux de neurones pour la reconnaissance des formes : architectures et apprentissage
Guyon, I · 1988
Earlier work this paper cites.
What size net gives valid generalization ?
Baum, E. B. and Haussler, D · 1989
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
Generalized additive models
Hastie, T. and Tibshirani, R · 1990
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen. Diploma thesis, Institut für Informatik, Lehrstuhl Prof. Brauer, Technische Universität München, 1991
Hochreiter, S · 1991
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Hornik, K · 1991
Earlier work this paper cites.
Eigenvalues of covariance matrices : Application to neural-network learning
LeCun, Y., Kanter, I., and Solla, S · 1991
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
Geman, S., Bienenstock, E., and Doursat, R · 1992
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Leshno, M. and Schocken, S · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Barron, A. R · 1994
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P · 1994
Earlier work this paper cites.
Polynomial bounds for vc dimension of sigmoidal neural networks
Karpinski, M. and Macintyre, A · 1995
Earlier work this paper cites.
Statistical mechanics of learning : Generalization
Opper, M · 1995
Earlier work this paper cites.
Bias plus variance decomposition for zero-one loss functions
Kohavi, R. and Wolpert, D · 1996
Earlier work this paper cites.
Dynamics of training
Bös, S. and Opper, M · 1997
Earlier work this paper cites.
Almost linear vc dimension bounds for piecewise polynomial networks
Bartlett, P. L., Maiorov, V., and Meir, R · 1998
Earlier work this paper cites.
Efficient backprop
LeCun, Y., Bottou, L., Orr, G. B., and Müller, K.-R · 1998
Earlier work this paper cites.
Statistical learning theory
Vapnik, V. N · 1998
Earlier work this paper cites.
Improved boosting algorithms using confidence-rated predictions
Schapire, R. E. and Singer, Y · 1999
Earlier work this paper cites.
An overview of statistical learning theory
Vapnik, V. N · 1999
Earlier work this paper cites.
A unified bias-variance decomposition and its applications
Domingos, P · 2000
Earlier work this paper cites.
Pattern Classification
Duda, R. O., Hart, P. E., and Stork, D. G · 2001
Earlier work this paper cites.
The Elements of Statistical Learning
Hastie, T., Tibshirani, R., and Friedman, J · 2001
Earlier work this paper cites.
The Concentration of Measure Phenomenon
Ledoux, M · 2001
Earlier work this paper cites.
Learning to generalize
Opper, M · 2001
Cited alongside, same era.
Stability and generalization
Bousquet, O. and Elisseeff, A · 2002
Cited alongside, same era.
Contributions to decision tree induction : bias/variance tradeoff and time series classification
Geurts, P · 2002
Cited alongside, same era.
Rademacher and gaussian complexities : Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2003
Cited alongside, same era.
Boosting with the l 2 loss : regression and classification
Bühlmann, P. and Yu, B · 2003
Cited alongside, same era.
Variance and bias for general loss functions
James, G. M · 2003
Cited alongside, same era.
Three factors influencing minima in SGD
Jastrzkebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A. J · 2017
Later among the works it cites.
On large-batch training for deep learning : Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Later among the works it cites.
Fisher-rao metric, geometry, and complexity of neural networks
Liang, T., Poggio, T. A., Rakhlin, A., and Stokes, J · 2017
Later among the works it cites.
Implicit regularization in deep learning
Neyshabur, B · 2017
Later among the works it cites.
Resurrecting the sigmoid in deep learning through dynamical isometry : theory and practice
Pennington, J., Schoenholz, S., and Ganguli, S · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bishop, C. M · 2006
Cited alongside, same era.
The tradeoffs of large scale learning
Bottou, L. and Bousquet, O · 2008
Cited alongside, same era.
The elements of statistical learning : data mining, inference and prediction
Hastie, T., Tibshirani, R., and Friedman, J · 2009
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Cited alongside, same era.
Learning From Data
Abu-Mostafa, Y. S., Magdon-Ismail, M., and Lin, H.-T · 2012
Cited alongside, same era.
Understanding the bias-variance tradeoff, June 2012
Fortmann-Roe, S · 2012
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2017
Later among the works it cites.
Deep information propagation
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J · 2017
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Later among the works it cites.
Theory of deep learning iib : Optimization properties of SGD
Zhang, C., Liao, Q., Rakhlin, A., Miranda, B., Golowich, N., and Poggio, T. A · 2017
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
Du, S. and Lee, J · 2018
Later among the works it cites.
Wtf is the bias-variance tradeoff ? (infographic), May 2018
EliteDataScience · 2018
Later among the works it cites.
Characterizing implicit bias in terms of optimization geometry
Gunasekar, S., Lee, J., Soudry, D., and Srebro, N · 2018
Later among the works it cites.
Gradient descent happens in a tiny subspace
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Later among the works it cites.
Deep neural networks as gaussian processes
Lee, J., Sohl-dickstein, J., Pennington, J., Novak, R., Schoenholz, S., and Bahri, Y · 2018
Later among the works it cites.
Measuring the intrinsic dimension of objective landscapes
Li, C., Farkhoor, H., Liu, R., and Yosinski, J · 2018
Later among the works it cites.
A modern take on the bias-variance tradeoff in neural networks, 2018
Neal, B., Mittal, S., Baratin, A., Tantia, V., Scicluna, M., Lacoste-Julien, S., and Mitliagkas, I · 2018
Later among the works it cites.
Sensitivity and generalization in neural networks : an empirical study
Novak, R., Bahri, Y., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J · 2018
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P.-J., and Le, Q. V · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., and Srebro, N · 2018
Later among the works it cites.
A jamming transition from under- to over-parametrization affects loss landscape and generalization
Spigler, S., Geiger, M., d’Ascoli, S., Sagun, L., Biroli, G., and Wyart, M · 2018
Later among the works it cites.
Dynamical isometry and a mean field theory of CNNs : How to train 10,000-layer vanilla convolutional neural networks
Xiao, L., Bahri, Y., Sohl-Dickstein, J., Schoenholz, S., and Pennington, J · 2018
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S., Zhai, X., Poczos, B., and Singh, A · 2019
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation, 2019
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J · 2019
Closest in time.
Deep double descent : Where bigger models and more data hurt, 2019
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2019
Closest in time.
In support of over-parametrization in deep reinforcement learning : an empirical study
Neal, B. and Mitliagkas, I · 2019
Closest in time.
The role of over-parametrization in generalization of neural networks
Neyshabur, B., Li, Z., Bhojanapalli, S., LeCun, Y., and Srebro, N · 2019
Closest in time.
The effect of network width on stochastic gradient descent and generalization : an empirical study
Park, D., Sohl-Dickstein, J., Le, Q., and Smith, S · 2019
Closest in time.