Fetching the paper…
Reading the bibliography…
The speed at which one can minimize an expected loss using stochastic methods depends on two properties: the curvature of the loss and the variance of the gradients.
A scale invariant flatness measure for deep network minima
Rangamani, A., Nguyen, N. H., Kumar, A., Phan, D., Chin, S. H., and Tran, T. D. (2019) · 1902
Earlier work this paper cites.
Limitations of the empirical fisher approximation
Kunstner, F., Balles, L., and Hennig, P. (2019) · 1905
Earlier work this paper cites.
A new look at the statistical model identification
Akaike, H. (1974) · 1974
Earlier work this paper cites.
The distribution of information statistics and the criterion of goodness of fit of models
Takeuchi, K. (1976) · 1976
Earlier work this paper cites.
Network information criterion-determining the number of hidden units for an artificial neural network model
Murata, N., Yoshizawa, S., and Amari, S.-i. (1994) · 1994
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Stability and generalization
Bousquet, O. and Elisseeff, A. (2002) · 2002
Earlier work this paper cites.
The tradeoffs of large scale learning
Bottou, L. and Bousquet, O. (2008) · 2008
Earlier work this paper cites.
Improving first and second-order methods by modeling uncertainty
Le Roux, N., Bengio, Y., and Fitzgibbon, A. (2011) · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate o (1/n)
Bach, F. and Moulines, E. (2013) · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Cited alongside, same era.
New insights and perspectives on the natural gradient method
Martens, J. (2014) · 2014
Cited alongside, same era.
Convergence rate of stochastic gradient with constant step size
Schmidt, M. (2014) · 2014
Cited alongside, same era.
From averaging to acceleration, there is only a step-size
Flammarion, N. and Bach, F. (2015) · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Cited alongside, same era.
Three factors influencing minima in sgd
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A. (2017) · 2017
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. (2017) · 2017
Later among the works it cites.
Fisher-rao metric, geometry, and complexity of neural networks
Liang, T., Poggio, T., Rakhlin, A., and Stokes, J. (2017) · 2017
Later among the works it cites.
Stochastic gradient descent as approximate bayesian inference
Mandt, S., Hoffman, M. D., and Blei, D. M. (2017) · 2017
Later among the works it cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nonparametric stochastic approximation with large step-sizes
Dieuleveut, A., Bach, F., et al. (2016) · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Sagun, L., Bottou, L., and LeCun, Y. (2016) · 2016
Cited alongside, same era.
On optimal generalizability in parametric learning
Beirami, A., Razaviyayn, M., Shahrampour, S., and Tarokh, V. (2017) · 2017
Cited alongside, same era.
Chaudhari, P. and Soatto, S. (2017) · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y. (2017) · 2017
Cited alongside, same era.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
George, T., Laurent, C., Bouthillier, X., Ballas, N., and Vincent, P. (2018) · 2018
Later among the works it cites.
Sensitivity and generalization in neural networks: an empirical study
Novak, R., Bahri, Y., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J. (2018) · 2018
Later among the works it cites.
Approximate leave-one-out for fast parameter tuning in high dimensions
Wang, S., Zhou, W., Lu, H., Maleki, A., and Mirrokni, V. (2018) · 2018
Later among the works it cites.
Fluctuation-dissipation relations for stochastic gradient descent
Yaida, S. (2018) · 2018
Later among the works it cites.
Zhu, Z., Wu, J., Yu, B., Wu, L., and Ma, J. (2018) · 2018
Later among the works it cites.