Fetching the paper…
Reading the bibliography…
Gradient descent optimization algorithms are the standard ingredients that are used to train artificial neural networks (ANNs).
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E., and Bach, F · 2011
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / n ) O(1/n)
Bach, F., and Moulines, E · 2013
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Bach, F · 2017
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Ruder, S · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L., and Bach, F · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczós, B., and Singh, A · 2018
Earlier work this paper cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Hanin, B · 2018
Earlier work this paper cites.
How to start training: The effect of initialization and architecture
Hanin, B., and Rolnick, D · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y., and Liang, Y · 2018
Cited alongside, same era.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Wu, L., Ma, C., and E, W · 2018
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z., Li, Y., and Liang, Y · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Cited alongside, same era.
Non-convergence of stochastic gradient descent in the training of deep neural networks
Stochastic gradient descent for nonconvex learning without bounded gradient assumptions
Lei, Y., Hu, T., Li, G., and Tang, K · 2020
Later among the works it cites.
Lovas, A., Lytras, I., Rásonyi, M., and Sabanis, S · 2020
Later among the works it cites.
Dying ReLU and initialization: Theory and numerical examples
Lu, L., Shin, Y., Su, Y., and Karniadakis, G. E · 2020
Later among the works it cites.
Sankararaman, K. A., De, S., Xu, Z., Huang, W. R., and Goldstein, T · 2020
Later among the works it cites.
Trainability of ReLU networks and data-dependent initialization
Shin, Y., and Karniadakis, G. E · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cheridito, P., Jentzen, A., and Rossmannek, F · 2020
Cited alongside, same era.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
E, W., Ma, C., and Wu, L · 2020
Cited alongside, same era.
Convergence rates for the stochastic gradient descent method for non-convex objective functions
Fehrman, B., Gess, B., and Jentzen, A · 2020
Cited alongside, same era.
Lower error bounds for the stochastic gradient descent optimization algorithm: Sharp convergence rates for slowly and fast decaying learning rates
Jentzen, A., and von Wurstemberger, P · 2020
Cited alongside, same era.
Gradient descent optimizes over-parameterized deep ReLU networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2020
Later among the works it cites.
Akyildiz, Ö. D., and Sabanis, S · 2021
Closest in time.
Strong error analysis for stochastic gradient descent optimization algorithms
Jentzen, A., Kuckuck, B., Neufeld, A., and von Wurstemberger, P · 2021
Closest in time.