Fetching the paper…
Reading the bibliography…
We study the implicit regularization of mini-batch stochastic gradient descent, when applied to the fundamental problem of least squares regression.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
On asymptotic normality in stochastic approximation
Fabian, V · 1968
Earlier work this paper cites.
Theory and methods related to the singular-function expansion and Landweber’s iteration for integral equations of the first kind
Strand, O. N · 1974
Earlier work this paper cites.
Diffusions for global optimization
Geman, S. and Hwang, C.-R · 1986
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
Ruppert, D · 1988
Earlier work this paper cites.
Generalization and parameter estimation in feedforward nets: Some experiments
Morgan, N. and Bourlard, H · 1989
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Statistical mechanics of learning from examples
Seung, H. S., Sompolinsky, H., and Tishby, N · 1992
Earlier work this paper cites.
Worst-case quadratic loss bounds for prediction using linear functions and gradient descent
Cesa-Bianchi, N., Long, P. M., and Warmuth, M. K · 1996
Earlier work this paper cites.
Online learning and stochastic approximations
Bottou, L · 1998
Earlier work this paper cites.
Stochastic learning
Bottou, L · 2003
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
Kushner, H. and Yin, G. G · 2003
Earlier work this paper cites.
Stochastic differential equations
Øksendal, B · 2003
Earlier work this paper cites.
Gradient directed regularization
Friedman, J. and Popescu, B · 2004
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Zhang, T · 2004
Earlier work this paper cites.
Parameter flows
Ramsay, J · 2005
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N. and Lugosi, G · 2006
Earlier work this paper cites.
On early stopping in gradient descent learning
Yao, Y., Rosasco, L., and Caponnetto, A · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
Bousquet, O. and Bottou, L · 2008
Earlier work this paper cites.
Online gradient descent learning algorithms
Ying, Y. and Pontil, M · 2008
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L · 2010
Earlier work this paper cites.
Implementing regularization implicitly via approximate eigenvector computation
Mahoney, M. W. and Orecchia, L · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E. and Bach, F. R · 2011
Earlier work this paper cites.
Mcmc using hamiltonian dynamics
Neal, R. M. et al · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, B., Re, C., Wright, S., and Niu, F · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
Approximate computation and implicit regularization for very large-scale data analysis
Mahoney, M. W · 2012
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate o (1/n)
Bach, F. and Moulines, E · 2013
Earlier work this paper cites.
Gradient methods for convex minimization: better rates under weaker conditions
Zhang, H. and Yin, W · 2013
Cited alongside, same era.
Constant step size least-mean-square: Bias-variance trade-offs and optimal sampling distributions
Défossez, A. and Bach, F · 2014
Cited alongside, same era.
Anti-differentiating approximation algorithms: A case study with min-cuts, spectral, and flow
Gleich, D. and Mahoney, M · 2014
Cited alongside, same era.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A · 2014
Cited alongside, same era.
Approximation analysis of stochastic gradient langevin dynamics by using fokker-planck equation and ito process
Sato, I. and Nakagawa, H · 2014
Scaling sgd batch size to 32k for imagenet training
You, Y., Gitman, I., and Ginsburg, B · 2017
Later among the works it cites.
Determinantal point processes for mini-batch diversification
Zhang, C., Kjellstrom, H., and Mandt, S · 2017
Later among the works it cites.
A continuous-time view of early stopping for least squares regression
Ali, A., Kolter, J. Z., and Tibshirani, R. J · 2018
Later among the works it cites.
Constant step size stochastic gradient descent for probabilistic modeling
Babichev, D. and Bach, F · 2018
Later among the works it cites.
On exponential convergence of sgd in non-convex over-parametrized learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Asynchronous stochastic convex optimization
Duchi, J. C., Chaturapruek, S., and Ré, C · 2015
Cited alongside, same era.
Continuous-time limit of stochastic gradient descent revisited
Mandt, S., Hoffman, M. D., and Blei, D. M · 2015
Cited alongside, same era.
Learning with incremental iterative regularization
Rosasco, L. and Villa, S · 2015
Cited alongside, same era.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2016
Cited alongside, same era.
Statistical inference for model parameters in stochastic gradient descent
Chen, X., Lee, J. D., Tong, X. T., and Zhang, Y · 2016
Cited alongside, same era.
Early stopping as nonparametric variational inference
Duvenaud, D., Maclaurin, D., and Adams, R · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Cited alongside, same era.
Bassily, R., Belkin, M., and Ma, S · 2018
Later among the works it cites.
Statistical sparse online regression: A diffusion approximation perspective
Fan, J., Gong, W., Li, C. J., and Sun, Q · 2018
Later among the works it cites.
Parallelizing stochastic gradient descent for least squares regression: mini-batching, averaging, and model misspecification
Jain, P., Kakade, S. M., Kidambi, R., Netrapalli, P., and Sidford, A · 2018
Later among the works it cites.
An alternative view: When does sgd escape local minima?
Kleinberg, R., Li, Y., and Yuan, Y · 2018
Later among the works it cites.
Martin, C. H. and Mahoney, M. W · 2018
Later among the works it cites.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
Nacson, M. S., Srebro, N., and Soudry, D · 2018
Later among the works it cites.
Iterate averaging as regularization for stochastic gradient descent
Neu, G. and Rosasco, L · 2018
Later among the works it cites.
Statistical optimality of stochastic gradient descent on hard learning problems through multiple passes
Pillaud-Vivien, L., Rudi, A., and Bach, F · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Later among the works it cites.
Connecting optimization and regularization paths
Suggala, A., Prasad, A., and Ravikumar, P. K · 2018
Later among the works it cites.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Vaswani, S., Bach, F., and Schmidt, M · 2018
Later among the works it cites.
Theory of deep learning iib: Optimization properties of sgd
Zhang, C., Liao, Q., Rakhlin, A., Miranda, B., Golowich, N., and Poggio, T · 2018
Later among the works it cites.
Zhu, Z., Wu, J., Yu, B., Wu, L., and Ma, J · 2018
Later among the works it cites.
Cheng, X., Bartlett, P. L., and Jordan, M. I · 2019
Later among the works it cites.
Feng, Y., Gao, T., Li, L., Liu, J.-G., and Lu, Y · 2019
Later among the works it cites.
The implicit bias of gradient descent on nonseparable data
Ji, Z. and Telgarsky, M · 2019
Later among the works it cites.
Stochastic gradient descent escapes saddle points efficiently
Jin, C., Netrapalli, P., Ge, R., Kakade, S. M., and Jordan, M. I · 2019
Later among the works it cites.
Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations
Li, Q., Tai, C., and Weinan, E · 2019
Later among the works it cites.
Sgd on neural networks learns functions of increasing complexity
Nakkiran, P., Kaplun, G., Kalimeris, D., Yang, T., Edelman, B. L., Zhang, F., and Barak, B · 2019
Later among the works it cites.
Nguyen, V. A., Shafieezadeh-Abadeh, S., Kuhn, D., and Esfahani, P. M · 2019
Later among the works it cites.
Theoretical issues in deep networks: Approximation, optimization and generalization
Poggio, T., Banburski, A., and Liao, Q · 2019
Later among the works it cites.
A mathematical theory of semantic development in deep neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2019
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2019
Later among the works it cites.
Implicit regularization for optimal sparse recovery
Vaskevicius, T., Kanade, V., and Rebeschini, P · 2019
Later among the works it cites.