Fetching the paper…
Reading the bibliography…
To assess generalization, machine learning scientists typically either (i) bound the generalization gap and then (after training) plug in the empirical risk to obtain a bound on the true risk; or (ii) validate empirically on holdout data.
Deterministic pac-bayesian generalization bounds for deep networks via generalizing noise-resilience
Nagarajan, V. and Kolter, J. Z · 1905
Earlier work this paper cites.
Adjustment of an inverse matrix corresponding to a change in one element of a given matrix
Sherman, J. and Morrison, W. J · 1950
Earlier work this paper cites.
Learnability and the vapnik-chervonenkis dimension
Blumer, A., Ehrenfeucht, A., Haussler, D., and Warmuth, M. K · 1989
Earlier work this paper cites.
On the method of bounded differences , pp. 148–188
McDiarmid, C · 1989
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Hoeffding, W · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
On nonparametric estimation of density level sets
Tsybakov, A. B. et al · 1997
Earlier work this paper cites.
Gradient-Based Learning Applied to Document Recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Algorithmic stability and sanity-check bounds for leave-one-out cross-validation
Kearns, M. and Ron, D · 1999
Earlier work this paper cites.
An overview of statistical learning theory
Vapnik, V. N · 1999
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2002
Earlier work this paper cites.
Stability and generalization
Bousquet, O. and Elisseeff, A · 2002
Earlier work this paper cites.
Leave-one-out error and stability of learning algorithms with applications
Elisseeff, A., Pontil, M., et al · 2003
Earlier work this paper cites.
Gradient directed regularization for linear regression and classification
Friedman, J. and Popescu, B. E · 2003
Earlier work this paper cites.
Learning theory: stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization
Mukherjee, S., Niyogi, P., Poggio, T., and Rifkin, R · 2006
Earlier work this paper cites.
On early stopping in gradient descent learning
Yao, Y., Rosasco, L., and Caponnetto, A · 2007
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Learnability, stability and uniform convergence
Shalev-Shwartz, S., Shamir, O., Srebro, N., and Sridharan, K · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Machine Learning: A Probabilistic Perspective
Murphy, K. P · 2012
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
Concentration inequalities for sampling without replacement
Bardenet, R., Maillard, O.-A., et al · 2015
Earlier work this paper cites.
The ladder: A reliable leaderboard for machine learning competitions
Blum, A. and Hardt, M · 2015
Earlier work this paper cites.
Preserving statistical validity in adaptive data analysis
Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., and Roth, A. L · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Neyshabur, B., Tomioka, R., and Srebro, N · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Cited alongside, same era.
Non-vacuous generalization bounds at the imagenet scale: a pac-bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., and Orbanz, P · 2018
Later among the works it cites.
An exponential efron-stein inequality for lq stable learning rules
Abou-Moustafa, K. and Szepesvári, C · 2019
Later among the works it cites.
Arora, S., Du, S. S., Hu, W., Li, Z., and Wang, R · 2019
Later among the works it cites.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Achille, A., Rovere, M., and Soatto, S · 2017
Cited alongside, same era.
A closer look at memorization in deep networks
Arpit, D., Jastrzebski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A., Bengio, Y., et al · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M. J · 2017
Cited alongside, same era.
Dziugaite, G. K. and Roy, D. M · 2017
Cited alongside, same era.
Deep learning is robust to massive label noise
Rolnick, D., Veit, A., Belongie, S., and Shavit, N · 2017
Cited alongside, same era.
A continuous-time view of early stopping for least squares
Ali, A., Kolter, J. Z., and Tibshirani, R. J · 2018
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Arora, S., Ge, R., Neyshabur, B., and Zhang, Y · 2018
Cited alongside, same era.
Hu, W., Li, Z., and Yu, D · 2019
Later among the works it cites.
The implicit bias of gradient descent on nonseparable data
Ji, Z. and Telgarsky, M · 2019
Later among the works it cites.
Li, M., Soltanolkotabi, M., and Oymak, S · 2019
Later among the works it cites.
Sgd on neural networks learns functions of increasing complexity
Nakkiran, P., Kaplun, G., Kalimeris, D., Yang, T., Edelman, B. L., Zhang, F., and Barak, B · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
Do imagenet classifiers generalize to imagenet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2019
Later among the works it cites.
The implicit regularization of stochastic gradient flow for least squares
Ali, A., Dobriban, E., and Tibshirani, R. J · 2020
Later among the works it cites.
For self-supervised learning, rationality implies generalization, provably
Bansal, Y., Kaplun, G., and Barak, B · 2020
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L. and Bach, F · 2020
Later among the works it cites.
The early phase of neural network training
Frankle, J., Schwab, D. J., and Morcos, A. S · 2020
Later among the works it cites.
The surprising simplicity of the early-time learning dynamics of neural networks
Hu, W., Xiao, L., Adlam, B., and Pennington, J · 2020
Later among the works it cites.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Hui, L. and Belkin, M · 2020
Later among the works it cites.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
Li, M., Soltanolkotabi, M., and Oymak, S · 2020
Later among the works it cites.
Early-learning regularization prevents memorization of noisy labels
Liu, S., Niles-Weed, J., Razavian, N., and Fernandez-Granda, C · 2020
Later among the works it cites.
Classification vs regression in overparameterized regimes: Does the loss function matter?
Muthukumar, V., Narang, A., Subramanian, V., Belkin, M., Hsu, D., and Sahai, A · 2020
Later among the works it cites.
In defense of uniform convergence: Generalization via derandomization with an application to interpolating predictors
Negrea, J., Dziugaite, G. K., and Roy, D · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Later among the works it cites.
On uniform convergence and low-norm interpolation learning
Zhou, L., Sutherland, D. J., and Srebro, N · 2020
Later among the works it cites.