Fetching the paper…
Reading the bibliography…
In recent years, artificial neural networks have developed into a powerful tool for addressing a multitude of problems for which classical solution approaches reach their limits.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Hornik, K · 1991
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Leshno, M., Lin, V. Y., Pinkus, A., and Schocken, S · 1993
Earlier work this paper cites.
Non-Asymptotic Analysis of Stochastic Approximation Algorithms for Machine Learning
Moulines, E., and Bach, F · 2011
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / n ) O(1/n)
Bach, F., and Moulines, E · 2013
Earlier work this paper cites.
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Breaking the Curse of Dimensionality with Convex Neural Networks
Bach, F · 2017
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Ruder, S · 2017
Earlier work this paper cites.
Automatic Differentiation in Machine Learning: a Survey
Baydin, A. G., Pearlmutter, B. A., Radul, A. A., and Siskind, J. M · 2018
Earlier work this paper cites.
On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport
Chizat, L., and Bach, F · 2018
Earlier work this paper cites.
Which Neural Net Architectures Give Rise to Exploding and Vanishing Gradients?
Hanin, B · 2018
Earlier work this paper cites.
How to Start Training: The Effect of Initialization and Architecture
Hanin, B., and Rolnick, D · 2018
Earlier work this paper cites.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured Data
Li, Y., and Liang, Y · 2018
Cited alongside, same era.
How SGD Selects the Global Minima in Over-parameterized Learning: A Dynamical Stability Perspective
Wu, L., Ma, C., and E, W · 2018
Cited alongside, same era.
Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Cited alongside, same era.
Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two Layers
Allen-Zhu, Z., Li, Y., and Liang, Y · 2019
Cited alongside, same era.
Stochastic Gradient Descent for Nonconvex Learning Without Bounded Gradient Assumptions
Lei, Y., Hu, T., Li, G., and Tang, K · 2020
Later among the works it cites.
Dying ReLU and Initialization: Theory and Numerical Examples
Lu, L., Shin, Y., Su, Y., and Karniadakis, G. E · 2020
Later among the works it cites.
The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
Sankararaman, K. A., De, S., Xu, Z., Huang, W. R., and Goldstein, T · 2020
Later among the works it cites.
Trainability of ReLU Networks and Data-Dependent Initialization
Shin, Y., and Karniadakis, G. E · 2020
Later among the works it cites.
Non-convergence of stochastic gradient descent in the training of deep neural networks
Cheridito, P., Jentzen, A., and Rossmannek, F · 2021
Closest in time.
Full error analysis for the training of deep neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A Convergence Theory for Deep Learning via Over-Parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Cited alongside, same era.
Gradient Descent Finds Global Minima of Deep Neural Networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Cited alongside, same era.
Gradient Descent Provably Optimizes Over-parameterized Neural Networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2019
Cited alongside, same era.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
E, W., Ma, C., and Wu, L · 2020
Cited alongside, same era.
Convergence rates for the stochastic gradient descent method for non-convex objective functions
Fehrman, B., Gess, B., and Jentzen, A · 2020
Cited alongside, same era.
Lower error bounds for the stochastic gradient descent optimization algorithm: Sharp convergence rates for slowly and fast decaying learning rates
Jentzen, A., and von Wurstemberger, P · 2020
Cited alongside, same era.
Beck, C., Jentzen, A., and Kuckuck, B · 2022
Closest in time.
A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions
Cheridito, P., Jentzen, A., Riekert, A., and Rossmannek, F · 2022
Closest in time.
Overall error analysis for the training of deep neural networks via stochastic gradient descent with random initialisation
Jentzen, A., and Welti, T · 2023
Closest in time.
Taming Neural Networks with TUSLA: Nonconvex Learning via Adaptive Stochastic Gradient Langevin Algorithms
Lovas, A., Lytras, I., Rásonyi, M., and Sabanis, S · 2023
Closest in time.
Nonasymptotic analysis of Stochastic Gradient Hamiltonian Monte Carlo under local conditions for nonconvex optimization
Akyildiz, O. D., and Sabanis, S · 2024
Closest in time.
Convergence of stochastic gradient descent schemes for Łojasiewicz-landscapes
Dereich, S., and Kassing, S · 2024
Closest in time.