Fetching the paper…
Reading the bibliography…
This paper introduces a distribution-dependent PAC-Chernoff bound that exhibits perfect tightness for interpolators, even within over-parameterized model classes.
Sur un nouveau théoreme-limite de la théorie des probabilités
Cramér, H. (1938) · 1938
Earlier work this paper cites.
Structural risk minimization over data-dependent hierarchies
Shawe-Taylor, J., Bartlett, P. L., Williamson, R. C., & Anthony, M. (1998) · 1940
Earlier work this paper cites.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Chernoff, H. (1952) · 1952
Earlier work this paper cites.
Convex Analysis
Rockafellar, R. T. (1970) · 1970
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., & Haffner, P. (1998) · 1998
Earlier work this paper cites.
PAC-Bayesian model averaging
McAllester, D. A. (1999) · 1999
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y. (2003) · 2003
Earlier work this paper cites.
Entropies, convexity, and functional inequalities, on Φ \Phi -entropies and Φ \Phi -sobolev inequalities
Chafaï, D. (2004) · 2004
Earlier work this paper cites.
On the benefits of invariance in neural networks
Lyle, C., van der Wilk, M., Kwiatkowska, M., Gal, Y., & Bloem-Reddy, B. (2020) · 2005
Earlier work this paper cites.
Information-theoretic upper and lower bounds for statistical estimation
Zhang, T. (2006) · 2006
Earlier work this paper cites.
PAC-Bayesian supervised classification: The thermodynamics of statistical learning
Catoni, O. (2007) · 2007
Earlier work this paper cites.
Measuring invariances in deep networks
Goodfellow, I., Lee, H., Le, Q., Saxe, A., & Ng, A. (2009) · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al. (2009) · 2009
Earlier work this paper cites.
The theorem of bahadur and rao and large portfolio losses
Herdegen, M. (2008) · 2011
Earlier work this paper cites.
Entropy, large deviations, and statistical mechanics
Ellis, R. S. (2012) · 2012
Earlier work this paper cites.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., & Shamir, O. (2014) · 2014
Earlier work this paper cites.
Understanding Machine Learning: From Theory to Algorithms
Shalev-Shwartz, S. & Ben-David, S. (2014) · 2014
Earlier work this paper cites.
Path-sgd: Path-normalized optimization in deep neural networks
Neyshabur, B., Salakhutdinov, R. R., & Srebro, N. (2015) · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Tishby, N. & Zaslavsky, N. (2015) · 2015
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
Vapnik, V. N. & Chervonenkis, A. Y. (2015) · 2015
Earlier work this paper cites.
Why are deep nets reversible: A simple theory, with implications for training
Arora, S., Liang, Y., & Ma, T. (2016) · 2016
Earlier work this paper cites.
Deep learning
Goodfellow, I., Bengio, Y., Courville, A., & Bengio, Y. (2016) · 2016
Cited alongside, same era.
A probabilistic framework for deep learning
Patel, A. B., Nguyen, M. T., & Baraniuk, R. (2016) · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Poole, B., Lahiri, S., Raghu, M., Sohl-Dickstein, J., & Ganguli, S. (2016) · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., & Wojna, Z. (2016) · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., & Telgarsky, M. J. (2017) · 2017
Cited alongside, same era.
Towards understanding the invertibility of convolutional neural networks
Gilbert, A. C., Zhang, Y., Lee, K., Zhang, Y., & Lee, H. (2017) · 2017
Cited alongside, same era.
The role of over-parametrization in generalization of neural networks
Neyshabur, B., Li, Z., Bhojanapalli, S., LeCun, Y., & Srebro, N. (2019) · 2019
Later among the works it cites.
A survey on image data augmentation for deep learning
Shorten, C. & Khoshgoftaar, T. M. (2019) · 2019
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a PAC-Bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., & Orbanz, P. (2019) · 2019
Later among the works it cites.
Benign overfitting in linear regression
Bartlett, P. L., Long, P. M., Lugosi, G., & Tsigler, A. (2020) · 2020
Later among the works it cites.
A group-theoretic framework for data augmentation
Chen, S., Dobriban, E., & Lee, J. H. (2020) · 2020
Later among the works it cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Identity matters in deep learning
Hardt, M. & Ma, T. (2017) · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Nocedal, J., Tang, P., Mudigere, D., & Smelyanskiy, M. (2017b) · 2017
Cited alongside, same era.
Generalization in deep networks: The role of distance from initialization
Nagarajan, V. & Kolter, J. Z. (2017) · 2017
Cited alongside, same era.
Deep information propagation
Schoenholz, S. S., Gilmer, J., Ganguli, S., & Sohl-Dickstein, J. (2017) · 2017
Cited alongside, same era.
Generalization error of invariant classifiers
Sokolic, J., Giryes, R., Sapiro, G., & Rodrigues, M. (2017) · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., & Vinyals, O. (2017) · 2017
Cited alongside, same era.
Feldman, V. & Zhang, C. (2020) · 2020
Later among the works it cites.
Simple and effective regularization methods for training on noisily labeled data with generalization guarantee
Hu, W., Li, Z., & Yu, D. (2020) · 2020
Later among the works it cites.
Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., & Bengio, S. (2020) · 2020
Later among the works it cites.
In defense of uniform convergence: Generalization via derandomization with an application to interpolating predictors
Negrea, J., Dziugaite, G. K., & Roy, D. (2020) · 2020
Later among the works it cites.
Mathematics of deep learning
Vidal, R., Bruna, J., Giryes, R., & Soatto, S. (2020) · 2020
Later among the works it cites.
Geometric deep learning: Grids, groups, graphs, geodesics, and gauges
Bronstein, M. M., Bruna, J., Cohen, T., & Veličković, P. (2021) · 2021
Later among the works it cites.
A law of robustness for two-layers neural networks
Bubeck, S., Li, Y., & Nagaraj, D. M. (2021) · 2021
Later among the works it cites.
Regularisation of neural networks by enforcing lipschitz continuity
Gouk, H., Frank, E., Pfahringer, B., & Cree, M. J. (2021) · 2021
Later among the works it cites.
Explaining generalization in deep learning: progress and fundamental limits
Nagarajan, V. (2021) · 2021
Later among the works it cites.
A PAC-Bayesian generalization bound for equivariant networks
Behboodi, A., Cesa, G., & Cohen, T. S. (2022) · 2022
Later among the works it cites.
Group symmetry in PAC learning
Elesedy, B. (2022) · 2022
Later among the works it cites.
Generalization in deep learning
Kawaguchi, K., Kaelbling, L. P., & Bengio, Y. (2022) · 2022
Later among the works it cites.
A universal law of robustness via isoperimetry
Bubeck, S. & Sellke, M. (2023) · 2023
Later among the works it cites.
PAC-Bayes-Chernoff bounds for unbounded losses
Casado, I., Ortega, L. A., Masegosa, A. R., & Pérez, A. (2024) · 2024
Later among the works it cites.
Fantastic generalization measures are nowhere to be found
Gastpar, M., Nachum, I., Shafer, J., & Weinberger, T. (2024) · 2024
Later among the works it cites.
Near-interpolators: Rapid norm growth and the trade-off between interpolation and generalization
Wang, Y., Sonthalia, R., & Hu, W. (2024) · 2024
Later among the works it cites.