Fetching the paper…
Reading the bibliography…
Many supervised machine learning methods are naturally cast as optimization problems.
V. N. Vapnik and A. Y. Chervonenkis, On a perceptron class. Avtomat. i Telemekh. 25
1964
Earlier work this paper cites.
R. M. Gower, M. Schmidt, F. Bach, and P. Richtárik, Variance-reduced methods for machine learning. Proceedings of the IEEE 108
1983
Earlier work this paper cites.
A. R. Barron, Universal approximation bounds for superpositions of a sigmoidal function. IEEE Transactions on Information Theory 39
1993
Earlier work this paper cites.
R. M. Neal, Bayesian learning for neural networks . Ph.D. thesis, University of Toronto, 1995
1995
Earlier work this paper cites.
A. W. Van der Vaart, Asymptotic statistics . 3, Cambridge University Press, 2000
2000
Earlier work this paper cites.
V. Kurkova and M. Sanguineti, Bounds on rates of variable-basis and neural-network approximation. IEEE Transactions on Information Theory 47
2001
Earlier work this paper cites.
B. Schölkopf and A. J. Smola, Learning with kernels . MIT Press, 2001
2001
Earlier work this paper cites.
V. Koltchinskii and D. Panchenko, Empirical margin distributions and bounding the generalization error of combined classifiers. The Annals of Statistics 30
2002
Earlier work this paper cites.
H. J. Kushner and G. G. Yin, Stochastic approximation and recursive algorithms and applications . Second edn., Springer-Verlag, 2003
2003
Earlier work this paper cites.
E. Suli and D. F. Mayers, An introduction to numerical analysis . Cambridge University Press, 2003, 2003
2003
Earlier work this paper cites.
S. Maniglia, Probabilistic representation and uniqueness results for measure-valued solutions of transport equations. Journal de mathématiques pures et appliquées 87
2007
Earlier work this paper cites.
A. Rahimi and B. Recht, Random features for large-scale kernel machines. Advances in neural information processing systems 20
2007
Earlier work this paper cites.
L. Ambrosio, N. Gigli, and G. Savaré, Gradient flows: in metric spaces and in the space of probability measures . Springer Science & Business Media, 2008
2008
Earlier work this paper cites.
Y. Cho and L. K. Saul, Kernel methods for deep learning. In Advances in neural information processing systems , pp. 342–350, 2009
2009
Earlier work this paper cites.
S. Bubeck, Convex optimization: Algorithms and complexity. Foundations and Trends in Machine Learning 8
2015
Earlier work this paper cites.
F. Santambrogio, Optimal transport for applied mathematicians . Springer, 2015
2015
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep learning . MIT Press, 2016
2016
Earlier work this paper cites.
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht, Gradient descent only converges to minimizers. In Conference on learning theory , pp. 1246–1257, 2016
2016
Cited alongside, same era.
F. Bach, Breaking the curse of dimensionality with convex neural networks. Journal of Machine Learning Research 18
2017
Cited alongside, same era.
S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro, Implicit regularization in matrix factorization. In Advances in neural information processing systems , pp. 6151–6159, 2017
2017
Cited alongside, same era.
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan, How to escape saddle points efficiently. In International conference on machine learning , pp. 1724–1732, PMLR, 2017
2017
Cited alongside, same era.
A. Nitanda and T. Suzuki, Stochastic particle gradient descent for infinite ensembles. Tech. Rep. 1712.05438, arXiv, 2017
S. Mei, T. Misiakiewicz, and A. Montanari, Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit. In Conference on learning theory , pp. 2388–2464, PMLR, 2019
2019
Later among the works it cites.
S. Mei and A. Montanari, The generalization error of random features regression: Precise asymptotics and the double descent curve. Communications on Pure and Applied Mathematics (2019)
2019
Later among the works it cites.
A. Nowak-Vila, F. Bach, and A. Rudi, A general theory for structured prediction with smooth convex surrogates. Tech. Rep. 1902.01958, arXiv, 2019
2019
Later among the works it cites.
L. Chizat and F. Bach, Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss. In Conference on learning theory , pp. 1305–1338, PMLR, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
L. Chizat and F. Bach, On the global convergence of gradient descent for over-parameterized models using optimal transport. Advances in Neural Information Processing Systems 31
2018
Cited alongside, same era.
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro, Characterizing implicit bias in terms of optimization geometry. In International conference on machine learning , pp. 1832–1841, 2018
2018
Cited alongside, same era.
A. Jacot, F. Gabriel, and C. Hongler, Neural tangent kernel: Convergence and generalization in neural networks. In Advances in neural information processing systems , pp. 8571–8580, 2018
2018
Cited alongside, same era.
S. Mei, A. Montanari, and P.-M. Nguyen, A mean field view of the landscape of two-layer neural networks. Proceedings of the National Academy of Sciences 115
2018
Cited alongside, same era.
M. Mohri, A. Rostamizadeh, and A. Talwalkar, Foundations of machine learning . MIT Press, 2018
2018
Cited alongside, same era.
Y. Nesterov, Lectures on convex optimization . 137, Springer, 2018
2018
Cited alongside, same era.
G. M. Rotskoff and E. Vanden-Eijnden, Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks. In Advances in neural information processing systems , pp. 7146–7155, 31, 2018
2018
Cited alongside, same era.
2020
Later among the works it cites.
Y. Li, T. Ma, and H. R. Zhang, Learning over-parametrized two-layer neural networks beyond NTK. In Conference on learning theory , pp. 2613–2682, PMLR, 2020
2020
Later among the works it cites.
P.-M. Nguyen and H. T. Pham, A rigorous framework for the mean field limit of multilayer neural networks. Tech. Rep. 2001.11443, arXiv, 2020
2020
Later among the works it cites.
J. Sirignano and K. Spiliopoulos, Mean field analysis of neural networks: A law of large numbers. SIAM Journal on Applied Mathematics 80
2020
Later among the works it cites.
S. Wojtowytsch, On the convergence of gradient descent training for two-layer ReLU-networks in the mean field regime. Tech. Rep. 2005.13530, arXiv, 2020
2020
Later among the works it cites.
G. Yang and E. J. Hu, Feature learning in infinite-width neural networks. Tech. Rep. 2011.14522, arXiv, 2020
2020
Later among the works it cites.
S. Akiyama and T. Suzuki, On learnability via gradient method for two-layer relu neural networks in teacher-student setting. Tech. Rep. 2106.06251, arXiv, 2021
2021
Closest in time.
L. Chizat, Sparse optimization on measures with over-parameterized gradient descent. Mathematical Programming (2021), 1–46
2021
Closest in time.
C. Fang, J. Lee, P. Yang, and T. Zhang, Modeling from features: a mean-field framework for over-parameterized deep neural networks. In Conference on learning theory , pp. 1887–1936, PMLR, 2021
2021
Closest in time.
C. Giraud, Introduction to high-dimensional statistics . Chapman and Hall/CRC, 2021
2021
Closest in time.
J. Sirignano and K. Spiliopoulos, Mean field analysis of deep neural networks. Mathematics of Operations Research (2021)
2021
Closest in time.
M. Zhou, R. Ge, and C. Jin, A local convergence theory for mildly over-parameterized two-layer neural network. In Conference on learning theory , PMLR, 2021
2021
Closest in time.