Fetching the paper…
Reading the bibliography…
In this article we study the stochastic gradient descent (SGD) optimization method in the training of fully-connected feedforward artificial neural networks with ReLU activation.
A method for solving the convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
Bertsekas, D. P., and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Introductory Lectures on Convex Optimization
Nesterov, Y · 2004
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E., and Bach, F · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K · 2012
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / n ) O(1/n)
Bach, F., and Moulines, E · 2013
Earlier work this paper cites.
Escaping from saddle points — online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Earlier work this paper cites.
Gradient descent only converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Earlier work this paper cites.
Gradient Descent Only Converges to Minimizers: Non-Isolated Critical Points and Invariant Regions
Panageas, I., and Piliouras, G · 2017
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Ruder, S · 2017
Earlier work this paper cites.
Solving stochastic differential equations and Kolmogorov equations by means of deep learning
Beck, C., Becker, S., Grohs, P., Jaafari, N., and Jentzen, A · 2018
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L., and Bach, F · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczós, B., and Singh, A · 2018
Earlier work this paper cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Hanin, B · 2018
Earlier work this paper cites.
How to start training: The effect of initialization and architecture
Hanin, B., and Rolnick, D · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y., and Liang, Y · 2018
Cited alongside, same era.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Wu, L., Ma, C., and E, W · 2018
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z., Li, Y., and Liang, Y · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Cited alongside, same era.
Full error analysis for the training of deep neural networks
Lower error bounds for the stochastic gradient descent optimization algorithm: Sharp convergence rates for slowly and fast decaying learning rates
Jentzen, A., and von Wurstemberger, P · 2020
Later among the works it cites.
Jentzen, A., and Welti, T · 2020
Later among the works it cites.
Stochastic gradient descent for nonconvex learning without bounded gradient assumptions
Lei, Y., Hu, T., Li, G., and Tang, K · 2020
Later among the works it cites.
Lovas, A., Lytras, I., Rásonyi, M., and Sabanis, S · 2020
Later among the works it cites.
Dying ReLU and initialization: Theory and numerical examples
Lu, L., Shin, Y., Su, Y., and Karniadakis, G. E · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beck, C., Jentzen, A., and Kuckuck, B · 2019
Cited alongside, same era.
General multilevel adaptations for stochastic approximation algorithms of Robbins-Monro and Polyak-Ruppert type
Dereich, S., and Müller-Gronbach, T · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Cited alongside, same era.
Non-asymptotic analysis of biased stochastic approximation scheme
Karimi, B., Miasojedow, B., Moulines, E., and Wai, H.-T · 2019
Cited alongside, same era.
First-order methods almost always avoid strict saddle points
Lee, J. D., Panageas, I., Piliouras, G., Simchowitz, M., Jordan, M. I., and Recht, B · 2019
Cited alongside, same era.
First-order methods almost always avoid saddle points: the case of vanishing step-sizes
Panageas, I., Piliouras, G., and Wang, X · 2019
Cited alongside, same era.
Exponential convergence time of gradient descent for one-dimensional deep linear neural networks
Shamir, O · 2019
Cited alongside, same era.
Later among the works it cites.
Sankararaman, K. A., De, S., Xu, Z., Huang, W. R., and Goldstein, T · 2020
Later among the works it cites.
Trainability of ReLU networks and data-dependent initialization
Shin, Y., and Karniadakis, G. E · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep ReLU networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2020
Later among the works it cites.
Akyildiz, Ö. D., and Sabanis, S · 2021
Closest in time.
Cheridito, P., Jentzen, A., Riekert, A., and Rossmannek, F · 2021
Closest in time.
Cheridito, P., Jentzen, A., and Rossmannek, F · 2021
Closest in time.
Convergence of stochastic gradient descent schemes for Lojasiewicz-landscapes
Dereich, S., and Kassing, S · 2021
Closest in time.
Jentzen, A., and Kröger, T · 2021
Closest in time.
Strong error analysis for stochastic gradient descent optimization algorithms
Jentzen, A., Kuckuck, B., Neufeld, A., and von Wurstemberger, P · 2021
Closest in time.