Fetching the paper…
Reading the bibliography…
Understanding the properties of neural networks trained via stochastic gradient descent (SGD) is at the heart of the theory of deep learning.
Cong Fang, Jason Lee, Pengkun Yang, and Tong Zhang, Modeling from features: a mean-field framework for over-parameterized deep neural networks , Conference on Learning Theory, PMLR, 2021, pp. 1887–1936
1936
Earlier work this paper cites.
Corinna Cortes and Vladimir Vapnik, Support-vector networks , Machine learning 20
1995
Earlier work this paper cites.
Peter L. Bartlett, The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network , IEEE Transactions on Information Theory 44
1998
Earlier work this paper cites.
Richard Jordan, David Kinderlehrer, and Felix Otto, The variational formulation of the Fokker–Planck equation , SIAM Journal on Mathematical Analysis 29
1998
Earlier work this paper cites.
Véronique Gayrard, Anton Bovier, Michael Eckhoff, and Markus Klein, Metastability in reversible diffusion processes I: Sharp asymptotics for capacities and exit times , Journal of the European Mathematical Society 6
2004
Earlier work this paper cites.
Cédric Villani, Optimal transport: old and new , vol. 338, Springer, 2009
2009
Earlier work this paper cites.
Tien-Chung Hu and Andrew Rosalsky, A note on the de La Vallée Poussin criterion for uniform integrability , Statistics & Probability Letters 81
2011
Earlier work this paper cites.
Ulrike Von Luxburg and Bernhard Schölkopf, Statistical learning theory: Models, concepts, and results , Handbook of the History of Logic, vol. 10, Elsevier, 2011, pp. 651–706
2011
Earlier work this paper cites.
Philippe Laurençot, Weak compactness techniques and coagulation equations , Evolutionary Equations with Applications in Natural Sciences, Springer, 2015, pp. 199–253
2015
Earlier work this paper cites.
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro, In search of the real inductive bias: On the role of implicit regularization in deep learning , Workshop Contribution at International Conference on Learning Representations, 2015
2015
Earlier work this paper cites.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, Deep residual learning for image recognition , Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778
2016
Earlier work this paper cites.
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al., Mastering the game of go with deep neural networks and tree search , Nature 529
2016
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, Attention is all you need , Advances in Neural Information Processing Systems, vol. 30, 2017
2017
Earlier work this paper cites.
Randall Balestriero and Richard G. Baraniuk, A spline theory of deep learning , International Conference on Machine Learning, PMLR, 2018, pp. 374–383
2018
Earlier work this paper cites.
Lenaic Chizat and Francis Bach, On the global convergence of gradient descent for over-parameterized models using optimal transport , Advances in Neural Information Processing Systems, vol. 32, 2018
2018
Earlier work this paper cites.
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht, Essentially no barriers in neural network energy landscape , International Conference on Machine Learning, PMLR, 2018, pp. 1309–1318
2018
Earlier work this paper cites.
Arthur Jacot, Franck Gabriel, and Clément Hongler, Neural tangent kernel: Convergence and generalization in neural networks , Advances in Neural Information Processing Systems, vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Song Mei, Andrea Montanari, and Phan-Minh Nguyen, A mean field view of the landscape of two-layer neural networks , Proceedings of the National Academy of Sciences 115
2018
Cited alongside, same era.
Grant M. Rotskoff and Eric Vanden-Eijnden, Neural networks as interacting particle systems: Asymptotic convexity of the loss landscape and universal scaling of the approximation error , Advances in Neural Information Processing Systems, vol. 32, 2018
2018
Cited alongside, same era.
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro, The implicit bias of gradient descent on separable data , The Journal of Machine Learning Research 19
2018
Cited alongside, same era.
2019
Cited alongside, same era.
Adel Javanmard, Marco Mondelli, and Andrea Montanari, Analysis of a two-layer neural network via displacement convexity , The Annals of Statistics 48
2020
Later among the works it cites.
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever, Deep double descent: Where bigger models and more data hurt , International Conference on Learning Representations, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro, A function space view of bounded norm infinite width ReLU nets: The multivariate case , International Conference on Learning Representations, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrew Brock, Jeff Donahue, and Karen Simonyan, Large scale GAN training for high fidelity natural image synthesis , International Conference on Learning Representations, 2019
2019
Cited alongside, same era.
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off , Proceedings of the National Academy of Sciences 116
2019
Cited alongside, same era.
Lenaic Chizat, Edouard Oyallon, and Francis Bach, On lazy training in differentiable programming , Advances in Neural Information Processing Systems, vol. 33, 2019
2019
Cited alongside, same era.
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry Vetrov, and Andrew Gordon Wilson, Loss surfaces, mode connectivity, and fast ensembling of DNNs , Advances in Neural Information Processing Systems, vol. 32, 2019
2019
Cited alongside, same era.
Rohith Kuditipudi, Xiang Wang, Holden Lee, Yi Zhang, Zhiyuan Li, Wei Hu, Sanjeev Arora, and Rong Ge, Explaining landscape connectivity of low-cost solutions for multilayer nets , Advances in Neural Information Processing Systems, vol. 33, 2019
2019
Cited alongside, same era.
Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Mean-field theory of two-layers neural networks: Dimension-free bounds and kernel limit , Conference on Learning Theory, PMLR, 2019, pp. 2388–2464
2019
Cited alongside, same era.
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro, Towards understanding the role of over-parametrization in generalization of neural networks , International Conference on Learning Representations, 2019
2019
Cited alongside, same era.
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro, How do infinite width bounded norm networks look in function space? , Conference on Learning Theory, PMLR, 2019, pp. 2667–2690
2019
Cited alongside, same era.
Rahul Parhi and Robert D Nowak, Banach space representer theorems for neural networks and ridge splines , The Journal of Machine Learning Research 22
2020
Later among the works it cites.
Alexander Shevchenko and Marco Mondelli, Landscape connectivity and dropout stability of sgd solutions for over-parameterized neural networks , International Conference on Machine Learning, PMLR, 2020, pp. 8773–8784
2020
Later among the works it cites.
2020
Later among the works it cites.
Justin Sirignano and Konstantinos Spiliopoulos, Mean field analysis of neural networks: A law of large numbers , SIAM Journal on Applied Mathematics 80
2020
Later among the works it cites.
Yaoyu Zhang, Zhi-Qin John Xu, Tao Luo, and Zheng Ma, A type of generalization error induced by initialization in deep neural networks , Mathematical and Scientific Machine Learning, PMLR, 2020, pp. 144–164
2020
Later among the works it cites.
Peter L. Bartlett, Andrea Montanari, and Alexander Rakhlin, Deep learning: A statistical viewpoint , Acta Numerica 30
2021
Closest in time.
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu, Towards understanding the spectral bias of deep learning , International Joint Conference on Artificial Intelligence, 2021
2021
Closest in time.
Tolga Ergen and Mert Pilanci, Convex geometry and duality of over-parameterized neural networks , The Journal of Machine Learning Research 22
2021
Closest in time.
2021
Closest in time.
Jingfeng Wu, Difan Zou, Vladimir Braverman, and Quanquan Gu, Direction matters: On the implicit bias of stochastic gradient descent with moderate learning rate , International Conference on Learning Representations, 2021
2021
Closest in time.
2022
Closest in time.
2022
Closest in time.
Kaitong Hu, Zhenjie Ren, David Šiška, and Łukasz Szpruch, Mean-field langevin dynamics and energy landscape of neural networks , Annales de l’Institut Henri Poincaré, Probabilités et Statistiques, vol. 57, Institut Henri Poincaré, 2021, pp. 2043–2065
2065
Closest in time.