Fetching the paper…
Reading the bibliography…
The properties of flat minima in the empirical risk landscape of neural networks have been debated for some time.
S. Lim, I. Kim, T. Kim, C. Kim, and S. Kim · 1905
Earlier work this paper cites.
Spin glass theory and beyond: An Introduction to the Replica Method and Its Applications , volume 9
M. Mézard, G. Parisi, and M. Virasoro · 1987
Earlier work this paper cites.
Non-vacuous generalization bounds at the imagenet scale: A pac-bayesian compression approach, 2018
W. Zhou, V. Veitch, M. Austern, R. P. Adams, and P. Orbanz · 1987
Earlier work this paper cites.
Bayesian back-propagation
W. L. Buntine and A. S. Weigend · 1991
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
G. E. Hinton and D. van Camp · 1993
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Information, physics, and computation
M. Mezard and A. Montanari · 2009
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
M. Welling and Y. W. Teh · 2011
Earlier work this paper cites.
Deep learning with elastic averaging sgd, 2014
S. Zhang, A. Choromanska, and Y. LeCun · 2014
Earlier work this paper cites.
Subdominant dense clusters allow for simple learning and high computational performance in neural networks with discrete synapses
C. Baldassi, A. Ingrosso, C. Lucibello, L. Saglietti, and R. Zecchina · 2015
Earlier work this paper cites.
Local entropy as a measure for sampling solutions in constraint satisfaction problems
C. Baldassi, A. Ingrosso, C. Lucibello, L. Saglietti, and R. Zecchina · 2016
Cited alongside, same era.
Deep pyramidal residual networks
D. Han, J. Kim, and J. Kim · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. P. Kingma · 2016
Cited alongside, same era.
Autoaugment: Learning augmentation policies from data
E. D. Cubuk, B. Zoph, D. Mané, V. Vasudevan, and Q. V. Le · 2018
Later among the works it cites.
Entropy-SGD optimizes the prior of a PAC-bayes bound: Data-dependent PAC-bayes priors via differential privacy, 2018
G. K. Dziugaite and D. M. Roy · 2018
Later among the works it cites.
Y. Yamada, M. Iwamura, and K. Kise · 2018
Later among the works it cites.
Properties of the geometry of solutions and capacity of multilayer neural networks with rectified linear unit activations
C. Baldassi, E. M. Malatesta, and R. Zecchina · 2019
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Chaudhari, C. Baldassi, R. Zecchina, S. Soatto, and A. Talwalkar · 2017
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
T. Devries and G. W. Taylor · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data, 2017
G. K. Dziugaite and D. M. Roy · 2017
Cited alongside, same era.
Deep relaxation: partial differential equations for optimizing deep neural networks
P. Chaudhari, A. Oberman, S. Osher, S. Soatto, and G. Carlier · 2018
Cited alongside, same era.
Unreasonable effectiveness of learning neural networks: From accessible states and robust ensembles to basic algorithmic schemes
C. Baldassi, C. Borgs, J. T. Chayes, A. Ingrosso, C. Lucibello, L. Saglietti, and R. Zecchina
Cited in the paper.
Later among the works it cites.
Fantastic generalization measures and where to find them, 2019
Y. Jiang, B. Neyshabur, H. Mobahi, D. Krishnan, and S. Bengio · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks, 2019
M. Tan and Q. V. Le · 2019
Later among the works it cites.
Lookahead optimizer: k steps forward, 1 step back
M. Zhang, J. Lucas, J. Ba, and G. E. Hinton · 2019
Later among the works it cites.
Shaping the learning landscape in neural networks around wide flat minima
C. Baldassi, F. Pittorino, and R. Zecchina · 2020
Closest in time.