Fetching the paper…
Reading the bibliography…
Deep learning empirically achieves high performance in many applications, but its training dynamics has not been fully understood theoretically.
Duality and stability in extremum problems involving convex functions
Rockafeller, R. T · 1967
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Barron, A. R · 1993
Earlier work this paper cites.
Techique of Variational Analysis
Borwein, J. M. and Zhu, Q. J · 2005
Earlier work this paper cites.
Optimization algorithms on matrix manifolds
Absil, P. A., Mahony, R., and Sepulchre, R · 2009
Earlier work this paper cites.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Welling, M. and Teh, Y.-W · 2011
Earlier work this paper cites.
Exact reconstruction using Beurling minimal extrapolation
De Castro, Y. and Gamboa, F · 2012
Earlier work this paper cites.
Inverse problems in spaces of measures
Bredies, K. and Pikkarainen, H. K · 2013
Earlier work this paper cites.
Distributions of angles in random packing on spheres
Cai, T. T., Fan, J., and Jiang, T · 2013
Earlier work this paper cites.
Super-resolution from noisy data
Candès, E. J. and Fernandez-Granda, C · 2013
Earlier work this paper cites.
Exact support recovery for sparse spikes deconvolution
Duval, V. and Peyré, G · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Neyshabur, B., Tomioka, R., and Srebro, N · 2015
Earlier work this paper cites.
On the rate of convergence of empirical measures in transportation distance
Trillos, N. G. and Slepčev, D · 2015
Earlier work this paper cites.
An Introduction to Matrix Concentration Inequalities , volume 8 of Foundations and Trends in Machine Learning
Tropp, J. A · 2015
Earlier work this paper cites.
Risk bounds for high-dimensional ridge function combinations including neural networks
Klusowski, J. M. and Barron, A. R · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Bach, F · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with ReLU activation
Li, Y. and Yuan, Y · 2017
Earlier work this paper cites.
Stochastic particle gradient descent for infinite ensembles
Nitanda, A. and Suzuki, T · 2017
Cited alongside, same era.
Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis
Raginsky, M., Rakhlin, A., and Telgarsky, M · 2017
Cited alongside, same era.
An analytical formula of population gradient for two-layered ReLU network and its applications in convergence and critical point analysis
Tian, Y · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Zhong, K., Song, Z., Jain, P., Bartlett, P. L., and Dhillon, I. S · 2017
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L. and Bach, F · 2018
Cited alongside, same era.
A priori estimates of the population risk for two-layer neural networks
Weinan, E., Ma, C., and Wu, L · 2019
Later among the works it cites.
Learning one-hidden-layer relu networks via gradient descent
Zhang, X., Yu, Y., Wang, L., and Gu, Q · 2019
Later among the works it cites.
A generalized neural tangent kernel analysis for two-layer neural networks
Chen, Z., Cao, Y., Gu, Q., and Zhang, T · 2020
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L. and Bach, F · 2020
Later among the works it cites.
On sparsity in overparametrised shallow ReLU networks
de Dios, J. and Bruna, J · 2020
Later among the works it cites.
On the linear convergence rates of exchange and continuous methods for total variation minimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Global non-convex optimization with discretized diffusions
Erdogdu, M. A., Mackey, L., and Shamir, O · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Cited alongside, same era.
The geometry of off-the-grid compressed sensing
Poon, C., Keriven, N., and Peyré, G · 2018
Cited alongside, same era.
Spurious local minima are common in two-layer ReLU neural networks
Safran, I. and Shamir, O · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Cited alongside, same era.
Flinth, A., de Gournay, F., and Weiss, P · 2020
Later among the works it cites.
Learning over-parametrized two-layer neural networks beyond NTK
Li, Y., Ma, T., and Zhang, H. R · 2020
Later among the works it cites.
Safran, I., Yehudai, G., and Shamir, O · 2020
Later among the works it cites.
Student specialization in deep rectified networks with finite width and input dimension
Tian, Y · 2020
Later among the works it cites.
Tzen, B. and Raginsky, M · 2020
Later among the works it cites.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
Weinan, E., Ma, C., and Wu, L · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N · 2020
Later among the works it cites.
Learning a single neuron with gradient methods
Yehudai, G. and Shamir, O · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2020
Later among the works it cites.
Sparse optimization on measures with over-parameterized gradient descent
Chizat, L · 2021
Closest in time.
Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods
Suzuki, T. and Akiyama, S · 2021
Closest in time.
A local convergence theory for mildly over-parameterized two-layer neural network
Zhou, M., Ge, R., and Jin, C · 2021
Closest in time.