Fetching the paper…
Reading the bibliography…
The bias-variance tradeoff tells us that as model complexity increases, bias falls and variances increases, leading to a U-shaped test error curve.
Bootstrap methods: Another look at the jackknife
Efron, B · 1979
Earlier work this paper cites.
Generalized additive models
Hastie, T. and Tibshirani, R · 1990
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen. Diploma thesis, Institut für Informatik, Lehrstuhl Prof. Brauer, Technische Universität München, 1991
Hochreiter, S · 1991
Earlier work this paper cites.
Eigenvalues of covariance matrices: Application to neural-network learning
LeCun, Y., Kanter, I., and Solla, S · 1991
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
Geman, S., Bienenstock, E., and Doursat, R · 1992
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P · 1994
Earlier work this paper cites.
Bias plus variance decomposition for zero-one loss functions
Kohavi, R. and Wolpert, D · 1996
Earlier work this paper cites.
Efficient backprop
LeCun, Y., Bottou, L., Orr, G. B., and Müller, K.-R · 1998
Earlier work this paper cites.
Statistical learning theory
Vapnik, V. N · 1998
Earlier work this paper cites.
Improved boosting algorithms using confidence-rated predictions
Schapire, R. E. and Singer, Y · 1999
Earlier work this paper cites.
An overview of statistical learning theory
Vapnik, V. N · 1999
Earlier work this paper cites.
A unified bias-variance decomposition and its applications
Domingos, P · 2000
Earlier work this paper cites.
Pattern Classification
Duda, R. O., Hart, P. E., and Stork, D. G · 2001
Earlier work this paper cites.
The Elements of Statistical Learning
Hastie, T., Tibshirani, R., and Friedman, J · 2001
Earlier work this paper cites.
The Concentration of Measure Phenomenon
Ledoux, M · 2001
Earlier work this paper cites.
Stability and generalization
Bousquet, O. and Elisseeff, A · 2002
Earlier work this paper cites.
Boosting with the l 2 loss: regression and classification
Bühlmann, P. and Yu, B · 2003
Earlier work this paper cites.
Variance and bias for general loss functions
James, G. M · 2003
Earlier work this paper cites.
Pattern Recognition and Machine Learning (Information Science and Statistics)
Bishop, C. M · 2006
Earlier work this paper cites.
The elements of statistical learning: data mining, inference and prediction
Hastie, T., Tibshirani, R., and Friedman, J · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Learning From Data
Abu-Mostafa, Y. S., Magdon-Ismail, M., and Lin, H.-T · 2012
Cited alongside, same era.
Understanding the bias-variance tradeoff, June 2012
Fortmann-Roe, S · 2012
Cited alongside, same era.
An Introduction to Statistical Learning: With Applications in R
James, G., Witten, D., Hastie, T., and Tibshirani, R · 2014
Cited alongside, same era.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., and Shamir, O · 2014
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural network
Saxe, A. M., Mcclelland, J. L., and Ganguli, S · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Deep information propagation
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J · 2017
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Later among the works it cites.
Theory of deep learning iib: Optimization properties of SGD
Zhang, C., Liao, Q., Rakhlin, A., Miranda, B., Golowich, N., and Poggio, T. A · 2017
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
Du, S. and Lee, J · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N · 2015
Cited alongside, same era.
An analysis of deep neural network models for practical applications
Canziani, A., Paszke, A., and Culurciello, E · 2016
Cited alongside, same era.
Linear regression and the bias variance tradeoff, 2016
Gonzalez, J. E · 2016
Cited alongside, same era.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Wtf is the bias-variance tradeoff? (infographic), May 2018
EliteDataScience · 2018
Closest in time.
Characterizing implicit bias in terms of optimization geometry
Gunasekar, S., Lee, J., Soudry, D., and Srebro, N · 2018
Closest in time.
Gradient descent happens in a tiny subspace
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Closest in time.
Deep neural networks as gaussian processes
Lee, J., Sohl-dickstein, J., Pennington, J., Novak, R., Schoenholz, S., and Bahri, Y · 2018
Closest in time.
Measuring the intrinsic dimension of objective landscapes
Li, C., Farkhoor, H., Liu, R., and Yosinski, J · 2018
Closest in time.
Sensitivity and generalization in neural networks: an empirical study
Novak, R., Bahri, Y., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J · 2018
Closest in time.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P.-J., and Le, Q. V · 2018
Closest in time.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., and Srebro, N · 2018
Closest in time.
A jamming transition from under- to over-parametrization affects loss landscape and generalization
Spigler, S., Geiger, M., d’Ascoli, S., Sagun, L., Biroli, G., and Wyart, M · 2018
Closest in time.
Dynamical isometry and a mean field theory of CNNs: How to train 10,000-layer vanilla convolutional neural networks
Xiao, L., Bahri, Y., Sohl-Dickstein, J., Schoenholz, S., and Pennington, J · 2018
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
Du, S., Zhai, X., Poczos, B., and Singh, A · 2019
Closest in time.
Scaling description of generalization with number of parameters in deep learning
Geiger, M., Jacot, A., Spigler, S., Gabriel, F., Sagun, L., d’Ascoli, S., Biroli, G., Hongler, C., and Wyart, M · 2019
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation, 2019
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J · 2019
Closest in time.
The role of over-parametrization in generalization of neural networks
Neyshabur, B., Li, Z., Bhojanapalli, S., LeCun, Y., and Srebro, N · 2019
Closest in time.
The effect of network width on stochastic gradient descent and generalization: an empirical study
Park, D., Sohl-Dickstein, J., Le, Q., and Smith, S · 2019
Closest in time.