Fetching the paper…
Reading the bibliography…
Recent advances in deep learning theory have evoked the study of generalizability across different local minima of deep neural networks (DNNs).
Assessing the accuracy of the maximum likelihood estimator: Observed versus expected fisher information
Efron, B. and Hinkley, D. V · 1978
Earlier work this paper cites.
Modeling by shortest data description
Rissanen, J · 1978
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
MacKay, D. J · 1992
Earlier work this paper cites.
Fisher information and stochastic complexity
Rissanen, J. J · 1996
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2002
Earlier work this paper cites.
(not) bounding the true error
Langford, J. and Caruana, R · 2002
Earlier work this paper cites.
Simplified pac-bayesian margin bounds
McAllester, D · 2003
Earlier work this paper cites.
The minimum description length principle
Grünwald, P. D. and Grunwald, A · 2007
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Robustness and generalization
Xu, H. and Mannor, S · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K · 2016
Cited alongside, same era.
Generalizing pooling functions in convolutional neural networks: Mixed, gated, and tree
Lee, C.-Y., Gallagher, P. W., and Tu, Z · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Cited alongside, same era.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M. J · 2017
Cited alongside, same era.
Robust large margin deep neural networks
Sokolić, J., Giryes, R., Sapiro, G., and Rodrigues, M. R · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Later among the works it cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Wu, L., Zhu, Z., et al · 2017
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Arora, S., Ge, R., Neyshabur, B., and Zhang, Y · 2018
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Du, S. S., Lee, J. D., Li, H., Wang, L., and Zhai, X · 2018
Later among the works it cites.
Large margin deep networks for classification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Cited alongside, same era.
Dziugaite, G. K. and Roy, D. M · 2017
Cited alongside, same era.
Nearly-tight vc-dimension bounds for piecewise linear neural networks
Harvey, N., Liaw, C., and Mehrabian, A · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Cited alongside, same era.
Elsayed, G., Krishnan, D., Mobahi, H., Regan, K., and Bengio, S · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G · 2018
Later among the works it cites.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
Neyshabur, B., Bhojanapalli, S., and Srebro, N · 2018
Later among the works it cites.
Optimization landscape and expressivity of deep cnns
Nguyen, Q. and Hein, M · 2018
Later among the works it cites.
Skip connections eliminate singularities
Orhan, E. and Pitkow, X · 2018
Later among the works it cites.
The spectrum of the fisher information matrix of a single-hidden-layer neural network
Pennington, J. and Worah, P · 2018
Later among the works it cites.
Empirical analysis of the hessian of over-parametrized neural networks, 2018
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2018
Later among the works it cites.
Predicting the generalization gap in deep networks with margin distributions
Jiang, Y., Krishnan, D., Mobahi, H., and Bengio, S · 2019
Closest in time.
Universal statistics of fisher information in deep neural networks: Mean field approach
Karakida, R., Akaho, S., and Amari, S.-i · 2019
Closest in time.
Lightlike neuromanifolds, occam’s razor and deep learning
Sun, K. and Nielsen, F · 2019
Closest in time.
Non-vacuous generalization bounds at the imagenet scale: a PAC-bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., and Orbanz, P · 2019
Closest in time.