Fetching the paper…
Reading the bibliography…
Empirically, the PAC-Bayesian analysis is known to produce tight risk bounds for practical machine learning algorithms.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
The sizes of compact subsets of hilbert space and continuity of gaussian processes
Dudley, R. M. (1967) · 1967
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time, i
Donsker, M. D. and Varadhan, S. S. (1975) · 1975
Earlier work this paper cites.
Some pac-bayesian theorems
McAllester, D. A. (1999) · 1999
Earlier work this paper cites.
Majorizing measures without measures
Talagrand, M. (2001) · 2001
Earlier work this paper cites.
Combining pac-bayesian and generic chaining bounds
Audibert, J.-Y. and Bousquet, O. (2007) · 2007
Earlier work this paper cites.
Pac-bayesian supervised classification: the thermodynamics of statistical learning
Catoni, O. (2007) · 2007
Earlier work this paper cites.
Optimal transport: old and new
Villani, C. (2008) · 2008
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W. (2011) · 2011
Cited alongside, same era.
Pac-bayesian inequalities for martingales
Seldin, Y., Laviolette, F., Cesa-Bianchi, N., Shawe-Taylor, J., and Auer, P. (2012) · 2012
Cited alongside, same era.
Concentration inequalities: A nonasymptotic theory of independence
Boucheron, S., Lugosi, G., and Massart, P. (2013) · 2013
Cited alongside, same era.
Pac-bayes-empirical-bernstein inequality
Tolstikhin, I. O. and Seldin, Y. (2013) · 2013
Cited alongside, same era.
Adding gradient noise improves learning for very deep networks
Neelakantan, A., Vilnis, L., Le, Q. V., Sutskever, I., Kaiser, L., Kurach, K., and Martens, J. (2015) · 2015
Cited alongside, same era.
How much does your data exploration overfit? controlling bias via information usage
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Dziugaite, G. K. and Roy, D. M. (2017) · 2017
Later among the works it cites.
Nearly-tight vc-dimension bounds for piecewise linear neural networks
Harvey, N., Liaw, C., and Mehrabian, A. (2017) · 2017
Later among the works it cites.
Entropy-sgd optimizes the prior of a pac-bayes bound: Generalization properties of entropy-sgd and data-dependent priors
Dziugaite, G. K. and Roy, D. M. (2018) · 2018
Later among the works it cites.
Generalization bounds of SGLD for non-convex learning: Two theoretical viewpoints
Mou, W., Wang, L., Zhai, X., and Zheng, K. (2018) · 2018
Later among the works it cites.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
Neyshabur, B., Bhojanapalli, S., and Srebro, N. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Russo, D. and Zou, J. (2015) · 2015
Cited alongside, same era.
An introduction to matrix concentration inequalities
Tropp, J. A. et al. (2015) · 2015
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M. J. (2017) · 2017
Cited alongside, same era.
Deterministic PAC-bayesian generalization bounds for deep networks via generalizing noise-resilience
Nagarajan, V. and Kolter, Z. (2019) · 2019
Closest in time.
Non-vacuous generalization bounds at the imagenet scale: a PAC-bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., and Orbanz, P. (2019) · 2019
Closest in time.