Fetching the paper…
Reading the bibliography…
We show that Entropy-SGD (Chaudhari et al., 2017), when viewed as a learning algorithm, optimizes a PAC-Bayes bound on the risk of a Gibbs (posterior) classifier, i.e., a randomized classifier obtained by a risk-sensitive perturbation of the weights of a learned classifier.
“Keeping the Neural Networks Simple by Minimizing the Description Length of the Weights”
Geoffrey. Hinton and Drew van Camp · 1993
Earlier work this paper cites.
“Flat Minima”
Sepp Hochreiter and J\"urgen Schmidhuber · 1997
Earlier work this paper cites.
“PAC-Bayesian Model Averaging”
David. McAllester · 1999
Earlier work this paper cites.
“Quantitatively tight sample complexity bounds”, 2002
John Langford · 2002
Earlier work this paper cites.
“(Not) Bounding the True Error”
John Langford and Rich Caruana · 2002
Earlier work this paper cites.
“Differential Privacy”
Cynthia Dwork · 2006
Earlier work this paper cites.
“From ε \varepsilon -entropy to KL-entropy: Analysis of minimum information complexity density estimation”
Tong Zhang · 2006
Earlier work this paper cites.
“Information-theoretic upper and lower bounds for statistical estimation”
Tong Zhang · 2006
Earlier work this paper cites.
“PAC-Bayesian Supervised Classification: The Thermodynamics of Statistical Learning”, Lecture Notes-Monograph Series
Olivier Catoni · 2007
Earlier work this paper cites.
“Mechanism Design via Differential Privacy”
Frank McSherry and Kunal Talwar · 2007
Earlier work this paper cites.
“Learning Multiple Layers of Features from Tiny Images” https://www.cs.toronto.edu/˜kriz/learning-features-2009-TR.pdf , 2009
Alex Krizhevsky · 2009
Earlier work this paper cites.
“MNIST handwritten digit database”, http://yann.lecun.com/exdb/mnist/, 2010
Yann LeCun, Corinna Cortes and Christopher J.. Burges · 2010
Earlier work this paper cites.
“Differentially private empirical risk minimization”
Kamalika Chaudhuri, Claire Monteleoni and Anand Sarwate · 2011
Earlier work this paper cites.
“Bayesian learning via stochastic gradient Langevin dynamics”
Max Welling and Yee Teh · 2011
Earlier work this paper cites.
“The Safe Bayesian-Learning the Learning Rate via the Mixability Gap.”
Peter Gr\"unwald · 2012
Cited alongside, same era.
“Private convex empirical risk minimization and high-dimensional regression”
Daniel Kifer, Adam Smith and Abhradeep Thakurta · 2012
Cited alongside, same era.
“A PAC-Bayesian Tutorial with a Dropout Bound”, 2013
David. McAllester · 2013
Cited alongside, same era.
“Differential privacy: an exploration of the privacy-utility landscape”, 2013
Darakhshan Mir · 2013
Cited alongside, same era.
Raef Bassily, Adam Smith and Abhradeep Thakurta · 2014
Cited alongside, same era.
“Algorithmic stability for adaptive data analysis”
Raef Bassily et al · 2016
Later among the works it cites.
“PAC-Bayesian Theory Meets Bayesian Inference”
Pascal Germain, Francis Bach, Alexandre Lacoste and Simon Lacoste-Julien · 2016
Later among the works it cites.
“Fast Rates with Unbounded Losses”, 2016
Peter. Gr\"unwald and Nishant. Mehta · 2016
Later among the works it cites.
“Differential Privacy without Sensitivity”
Kentaro Minami, Hitomi Arai, Issei Sato and Hiroshi Nakagawa · 2016
Later among the works it cites.
“On the Emergence of Invariance and Disentangling in Deep Representations”, 2017
Alessandro Achille and Stefano Soatto · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christos Dimitrakakis, Blaine Nelson, Aikaterini Mitrokotsa and Benjamin Rubinstein · 2014
Cited alongside, same era.
Behnam Neyshabur, Ryota Tomioka and Nathan Srebro · 2014
Cited alongside, same era.
“Subdominant Dense Clusters Allow for Simple Learning and High Computational Performance in Neural Networks with Discrete Synapses”
Carlo Baldassi et al · 2015
Cited alongside, same era.
“Generalization in adaptive data analysis and holdout reuse”
Cynthia Dwork et al · 2015
Cited alongside, same era.
“Preserving statistical validity in adaptive data analysis”
Cynthia Dwork et al · 2015
Cited alongside, same era.
“Variational Dropout and the Local Reparameterization Trick”
Diederik Kingma, Tim Salimans and Max Welling · 2015
Cited alongside, same era.
“Privacy for Free: Posterior Sampling and Stochastic Gradient Monte Carlo”
Yu-Xiang Wang, Stephen. Fienberg and Alexander. Smola · 2015
Cited alongside, same era.
Pratik Chaudhari et al · 2017
Closest in time.
Gintare Dziugaite and Daniel. Roy · 2017
Closest in time.
“A PAC-Bayesian Analysis of Randomized Learning with Applications to Stochastic Gradient Descent”
Ben London · 2017
Closest in time.
“Differential privacy and generalization: Sharper bounds with applications”
Luca Oneto, Sandro Ridella and Davide Anguita · 2017
Closest in time.
“Non-convex learning via Stochastic Gradient Langevin Dynamics: a nonasymptotic analysis”
Maxim Raginsky, Alexander Rakhlin and Matus Telgarsky · 2017
Closest in time.
“Understanding deep learning requires rethinking generalization”
Chiyuan Zhang et al · 2017
Closest in time.
“Information dropout: Learning optimal representations through noisy computation”
Alessandro Achille and Stefano Soatto · 2018
Closest in time.
“Data-dependent PAC-Bayes priors via differential privacy”, 2018
Gintare Dziugaite and Daniel. Roy · 2018
Closest in time.