Fetching the paper…
Reading the bibliography…
Despite the non-convex nature of their loss functions, deep neural networks are known to generalize well when optimized with stochastic gradient descent (SGD).
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Hochreiter, S. and Schmidhuber, J · 1995
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
Cesa-bianchi, N., Conconi, A., and Gentile, C · 2002
Earlier work this paper cites.
Stochastic convex optimization
Shalev-Shwartz, S., Shamir, O., Sridharan, K., and Srebro, N · 2009
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K · 2012
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J. and Vinyals, O · 2014
Earlier work this paper cites.
Open problem: The landscape of the loss surfaces of multilayer networks
Choromanska, A., LeCun, Y., and Arous, G. B · 2015
Earlier work this paper cites.
Escaping from saddle points - online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Densely Connected Convolutional Networks
Huang, G., Liu, Z., Weinberger, K. Q., and van der Maaten, L · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M · 2017
Earlier work this paper cites.
Sharp Minima Can Generalize For Deep Nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Earlier work this paper cites.
Learning One-hidden-layer Neural Networks with Landscape Design
Ge, R., Lee, J. D., and Ma, T · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R. B., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Cited alongside, same era.
Snapshot ensembles: Train 1, get m for free
Huang, G., Li, Y., Pleiss, G., Liu, Z., Hopcroft, J. E., and Weinberger, K. Q · 2017
Cited alongside, same era.
Three factors influencing minima in SGD
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A. J · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I · 2017
Stronger generalization bounds for deep nets via a compression approach
Arora, S., Ge, R., Neyshabur, B., and Zhang, Y · 2018
Later among the works it cites.
The loss landscape of overparameterized neural networks
Cooper, Y · 2018
Later among the works it cites.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F · 2018
Later among the works it cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Later among the works it cites.
Characterizing implicit bias in terms of optimization geometry
Gunasekar, S., Lee, J., Soudry, D., and Srebro, N · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generalization in Deep Learning
Kawaguchi, K., Pack Kaelbling, L., and Bengio, Y · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
Pennington, J. and Bahri, Y · 2017
Cited alongside, same era.
Spurious Local Minima are Common in Two-Layer ReLU Neural Networks
Safran, I. and Shamir, O · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Güney, V. U., Dauphin, Y., and Bottou, L · 2017
Cited alongside, same era.
A Bayesian Perspective on Generalization and Stochastic Gradient Descent
Smith, S. L. and Le, Q. V · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P., and Le, Q. V · 2017
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G · 2018
Later among the works it cites.
Risk and parameter convergence of logistic regression
Ji, Z. and Telgarsky, M · 2018
Later among the works it cites.
Accelerated gradient descent escapes saddle points faster than gradient descent
Jin, C., Netrapalli, P., and Jordan, M. I · 2018
Later among the works it cites.
An Alternative View: When Does SGD Escape Local Minima?
Kleinberg, R., Li, Y., and Yuan, Y · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Later among the works it cites.
Revisiting small batch training for deep neural networks
Masters, D. and Luschi, C · 2018
Later among the works it cites.
Towards understanding the role of over-parametrization in generalization of neural networks
Neyshabur, B., Li, Z., Bhojanapalli, S., LeCun, Y., and Srebro, N · 2018
Later among the works it cites.
First-order stochastic algorithms for escaping from saddle points in almost linear time
Xu, Y., Rong, J., and Yang, T · 2018
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a PAC-bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., and Orbanz, P · 2019
Closest in time.