Fetching the paper…
Reading the bibliography…
We investigate how the final parameters found by stochastic gradient descent are influenced by over-parameterization.
Effiicient backprop
LeCun, Y., Bottou, L., Orr, G. B., and Müller, K.-R · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
MNIST handwritten digit database
LeCun, Y., Cortes, C., and Burges, C. J · 2010
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M. J · 2017
Earlier work this paper cites.
Chaudhari, P. and Soatto, S · 2017
Earlier work this paper cites.
Dziugaite, G. K. and Roy, D. M · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R. B., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Nearly-tight vc-dimension bounds for piecewise linear neural networks
Harvey, N., Liaw, C., and Mehrabian, A · 2017
Cited alongside, same era.
Three factors influencing minima in SGD
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A. J · 2017
Cited alongside, same era.
Universal statistics of fisher information in deep neural networks: Mean field approach
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Later among the works it cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2018
Later among the works it cites.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S., Pennington, J., and Sohl-dickstein, J · 2018
Later among the works it cites.
An empirical model of large-batch training
McCandlish, S., Kaplan, J., Amodei, D., and Team, O. D · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karakida, R., Akaho, S., and ichi Amari, S · 2017
Cited alongside, same era.
Progressive growing of gans for improved quality, stability, and variation
Karras, T., Aila, T., Laine, S., and Lehtinen, J · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate bayesian inference
Mandt, S., Hoffman, M. D., and Blei, D. M · 2017
Cited alongside, same era.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Neyshabur, B., Bhojanapalli, S., and Srebro, N · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2017
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Smith, S. L. and Le, Q. V · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P., and Le, Q. V · 2017
Cited alongside, same era.
L2 regularization versus batch and weight normalization
van Laarhoven, T · 2017
Cited alongside, same era.
Later among the works it cites.
Deterministic pac-bayesian generalization bounds for deep networks via generalizing noise-resilience
Nagarajan, V. and Kolter, Z · 2018
Later among the works it cites.
Towards understanding the role of over-parametrization in generalization of neural networks
Neyshabur, B., Li, Z., Bhojanapalli, S., LeCun, Y., and Srebro, N · 2018
Later among the works it cites.
Sensitivity and generalization in neural networks: an empirical study
Novak, R., Bahri, Y., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J · 2018
Later among the works it cites.
Winner’s curse? on pace, progress, and empirical rigor
Sculley, D., Snoek, J., Wiltschko, A. B., and Rahimi, A · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J. M., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Later among the works it cites.
Fluctuation-dissipation relations for stochastic gradient descent
Yaida, S · 2018
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a pac-bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., and Orbanz, P · 2018
Later among the works it cites.