Fetching the paper…
Reading the bibliography…
In these six lectures, we examine what can be learnt about the behavior of multi-layer neural networks from the analysis of linear models.
Henry P McKean Jr, A class of markov processes associated with nonlinear parabolic equations , Proceedings of the National Academy of Sciences 56
1911
Earlier work this paper cites.
Gabor Szeg, Orthogonal polynomials , vol. 23, American Mathematical Soc., 1939
1939
Earlier work this paper cites.
Mark Kac, Foundations of kinetic theory , Proceedings of The third Berkeley symposium on mathematical statistics and probability, vol. 3, 1956, pp. 171–197
1956
Earlier work this paper cites.
BF Logan and Larry A Shepp, Optimal reconstruction of a function from its projections , Duke Math. J. 42
1975
Earlier work this paper cites.
Ronald A DeVore, Ralph Howard, and Charles Micchelli, Optimal nonlinear approximation , Manuscripta mathematica 63
1989
Earlier work this paper cites.
Kurt Hornik, Approximation capabilities of multilayer feedforward networks , Neural networks 4
1991
Earlier work this paper cites.
Andrew R Barron, Universal approximation bounds for superpositions of a sigmoidal function , IEEE Transactions on Information theory 39
1993
Earlier work this paper cites.
Radford M Neal, Bayesian learning for neural networks , Ph.D. thesis, Citeseer, 1995
1995
Earlier work this paper cites.
Christopher Williams, Computing with infinite networks , Advances in neural information processing systems 9
1996
Earlier work this paper cites.
Allan Pinkus, Approximation theory of the mlp model in neural networks , Acta numerica 8
1999
Earlier work this paper cites.
Maria-Florina Balcan, Avrim Blum, and Santosh Vempala, Kernels as features: On kernels, margins, and low-dimensional mappings , Machine Learning 65
2006
Earlier work this paper cites.
Andrea Caponnetto and Ernesto De Vito, Optimal rates for the regularized least-squares algorithm , Foundations of Computational Mathematics 7
2007
Earlier work this paper cites.
Ali Rahimi and Benjamin Recht, Random features for large-scale kernel machines , Advances in neural information processing systems 20
2007
Earlier work this paper cites.
Youngmin Cho and Lawrence Saul, Kernel methods for deep learning , Advances in neural information processing systems 22
2009
Earlier work this paper cites.
Noureddine El Karoui, The spectrum of kernel random matrices , The Annals of Statistics 38
2010
Earlier work this paper cites.
Alain Berlinet and Christine Thomas-Agnan, Reproducing kernel hilbert spaces in probability and statistics , Springer Science & Business Media, 2011
2011
Earlier work this paper cites.
Theodore S Chihara, An introduction to orthogonal polynomials , Courier Corporation, 2011
2011
Earlier work this paper cites.
Costas Efthimiou and Christopher Frye, Spherical harmonics in p dimensions , World Scientific, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Amit Daniely, Roy Frostig, and Yoram Singer, Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity , Advances in neural information processing systems 29
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
Francis Bach, Breaking the curse of dimensionality with convex neural networks , The Journal of Machine Learning Research 18
2017
Earlier work this paper cites.
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro, Implicit regularization in matrix factorization , Advances in neural information processing systems 30
2017
Earlier work this paper cites.
Antti Knowles and Jun Yin, Anisotropic local laws for random matrices , Probability Theory and Related Fields 169
2017
Earlier work this paper cites.
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro, Exploring generalization in deep learning , Advances in neural information processing systems 30
2017
Earlier work this paper cites.
Alessandro Rudi and Lorenzo Rosasco, Generalization properties of learning with random features. , NIPS, 2017, pp. 3215–3225
2017
Earlier work this paper cites.
Mikhail Belkin, Siyuan Ma, and Soumik Mandal, To understand deep learning we need to understand kernel learning , International Conference on Machine Learning, PMLR, 2018, pp. 541–549
2018
Earlier work this paper cites.
Lénaïc Chizat and Francis Bach, On the global convergence of gradient descent for over-parameterized models using optimal transport , Advances in Neural Information Processing Systems 31
2018
Earlier work this paper cites.
AGG De Matthews, J Hron, M Rowland, RE Turner, and Z Ghahramani, Gaussian process behaviour in wide deep neural networks , 6th International Conference on Learning Representations, ICLR 2018-Conference Track Proceedings, 2018
2018
Earlier work this paper cites.
Edgar Dobriban and Stefan Wager, High-dimensional asymptotics of prediction: Ridge regression and classification , The Annals of Statistics 46
2018
Earlier work this paper cites.
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh, Gradient descent provably optimizes over-parameterized neural networks , International Conference on Learning Representations, 2018
2018
Earlier work this paper cites.
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro, Characterizing implicit bias in terms of optimization geometry , International Conference on Machine Learning, PMLR, 2018, pp. 1832–1841
2018
Earlier work this paper cites.
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro, Implicit bias of gradient descent on linear convolutional networks , Advances in neural information processing systems 31
2018
Earlier work this paper cites.
Arthur Jacot, Franck Gabriel, and Clément Hongler, Neural tangent kernel: Convergence and generalization in neural networks , Advances in neural information processing systems, 2018, pp. 8571–8580
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein, Deep neural networks as gaussian processes , International Conference on Learning Representations, 2018
2018
Earlier work this paper cites.
Yuanzhi Li and Yingyu Liang, Learning overparameterized neural networks via stochastic gradient descent on structured data , Advances in Neural Information Processing Systems, 2018, pp. 8157–8166
2018
Earlier work this paper cites.
Song Mei, Yu Bai, and Andrea Montanari, The landscape of empirical risk for nonconvex losses , The Annals of Statistics 46
2018
Earlier work this paper cites.
Song Mei, Andrea Montanari, and Phan-Minh Nguyen, A mean field view of the landscape of two-layer neural networks , Proceedings of the National Academy of Sciences 115
2018
Earlier work this paper cites.
Roman Novak, Lechao Xiao, Yasaman Bahri, Jaehoon Lee, Greg Yang, Jiri Hron, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-dickstein, Bayesian deep convolutional networks with many channels are gaussian processes , International Conference on Learning Representations, 2018
2018
Earlier work this paper cites.
Grant M Rotskoff and Eric Vanden-Eijnden, Neural networks as interacting particle systems: Asymptotic convexity of the loss landscape and universal scaling of the approximation error , stat 1050
2018
Earlier work this paper cites.
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro, The implicit bias of gradient descent on separable data , The Journal of Machine Learning Research 19
2018
Earlier work this paper cites.
Roman Vershynin, High-dimensional probability: An introduction with applications in data science , vol. 47, Cambridge university press, 2018
2018
Earlier work this paper cites.
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang, Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks , International Conference on Machine Learning, PMLR, 2019, pp. 322–332
2019
Cited alongside, same era.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang, On exact computation with an infinitely wide neural net , Advances in Neural Information Processing Systems 32
2019
Cited alongside, same era.
Sanjeev Arora, Simon S Du, Zhiyuan Li, Ruslan Salakhutdinov, Ruosong Wang, and Dingli Yu, Harnessing the power of infinitely wide deep nets on small-data tasks , International Conference on Learning Representations, 2019
2019
Cited alongside, same era.
Zeyuan Allen-Zhu and Yuanzhi Li, What can resnet learn efficiently, going beyond kernels? , Proceedings of the 33rd International Conference on Neural Information Processing Systems, 2019, pp. 9017–9028
2019
Cited alongside, same era.
Yiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu, and Lexing Ying, A mean field analysis of deep resnet and beyond: Towards provably optimization via overparameterization from depth , International Conference on Machine Learning, PMLR, 2020, pp. 6426–6436
2020
Later among the works it cites.
Tengyuan Liang and Alexander Rakhlin, Just interpolate: Kernel “ridgeless” regression can generalize , The Annals of Statistics 48
2020
Later among the works it cites.
Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai, On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels , Conference on Learning Theory, PMLR, 2020, pp. 2683–2711
2020
Later among the works it cites.
Chaoyue Liu, Libin Zhu, and Mikhail Belkin, On the linearity of large non-linear models: when and why the tangent kernel is constant , Advances in Neural Information Processing Systems 33
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song, A convergence theory for deep learning via over-parameterization , International Conference on Machine Learning, PMLR, 2019, pp. 242–252
2019
Cited alongside, same era.
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off , Proceedings of the National Academy of Sciences 116
2019
Cited alongside, same era.
Yu Bai and Jason D Lee, Beyond linearization: On quadratic and higher-order approximation of wide neural networks , International Conference on Learning Representations, 2019
2019
Cited alongside, same era.
Mikhail Belkin, Alexander Rakhlin, and Alexandre B Tsybakov, Does data interpolation contradict statistical optimality? , The 22nd International Conference on Artificial Intelligence and Statistics, PMLR, 2019, pp. 1611–1619
2019
Cited alongside, same era.
Yuan Cao and Quanquan Gu, Generalization bounds of stochastic gradient descent for wide and deep neural networks , Advances in Neural Information Processing Systems 32
2019
Cited alongside, same era.
Lenaic Chizat, Edouard Oyallon, and Francis Bach, On lazy training in differentiable programming , NeurIPS 2019-33rd Conference on Neural Information Processing Systems, 2019, pp. 2937–2947
2019
Cited alongside, same era.
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai, Gradient descent finds global minima of deep neural networks , International Conference on Machine Learning, PMLR, 2019, pp. 1675–1685
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai, Harmless interpolation of noisy data in regression , IEEE Journal on Selected Areas in Information Theory 1
2020
Later among the works it cites.
2020
Later among the works it cites.
Vaishaal Shankar, Alex Fang, Wenshuo Guo, Sara Fridovich-Keil, Jonathan Ragan-Kelley, Ludwig Schmidt, and Benjamin Recht, Neural kernels without tangents , International Conference on Machine Learning, PMLR, 2020, pp. 8614–8623
2020
Later among the works it cites.
Justin Sirignano and Konstantinos Spiliopoulos, Mean field analysis of neural networks: A central limit theorem , Stochastic Processes and their Applications 130
2020
Later among the works it cites.
2020
Later among the works it cites.
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro, Kernel and rich regimes in overparametrized models , Conference on Learning Theory, PMLR, 2020, pp. 3635–3673
2020
Later among the works it cites.
Denny Wu and Ji Xu, On the optimal weighted ℓ 2 \ell_{2} regularization in overparameterized linear regression , Advances in Neural Information Processing Systems 33
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu, Gradient descent optimizes over-parameterized deep relu networks , Machine Learning 109
2020
Later among the works it cites.
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath, Online stochastic gradient descent on non-convex losses from high-dimensional inference , The Journal of Machine Learning Research 22
2021
Later among the works it cites.
Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin, Deep learning: a statistical viewpoint , Acta numerica 30
2021
Later among the works it cites.
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan, Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks , Nature communications 12
2021
Later among the works it cites.
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu, Towards understanding the spectral bias of deep learning , Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, 2021, pp. 2205–2211
2021
Later among the works it cites.
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová, Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime , Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
Frederic Koehler, Lijia Zhou, Danica J Sutherland, and Nathan Srebro, Uniform convergence of interpolators: Gaussian width, norm bounds and benign overfitting , Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
Fanghui Liu, Zhenyu Liao, and Johan Suykens, Kernel regression in high dimensions: Refined analysis beyond double descent , International Conference on Artificial Intelligence and Statistics, PMLR, 2021, pp. 649–657
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco, Asymptotics of ridge (less) regression under general source condition , International Conference on Artificial Intelligence and Statistics, PMLR, 2021, pp. 3889–3897
2021
Later among the works it cites.
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals, Understanding deep learning (still) requires rethinking generalization , Communications of the ACM 64
2021
Later among the works it cites.
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz, The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks , Conference on Learning Theory, PMLR, 2022, pp. 4782–4887
2022
Later among the works it cites.
Ben Adlam, Jake A Levinson, and Jeffrey Pennington, A random matrix perspective on mixtures of nonlinearities in high dimensions , International Conference on Artificial Intelligence and Statistics, PMLR, 2022, pp. 3434–3457
2022
Later among the works it cites.
2022
Later among the works it cites.
Alberto Bietti, Joan Bruna, Clayton Sanford, and Min Jae Song, Learning single-index models with shallow neural networks , Advances in Neural Information Processing Systems 35
2022
Later among the works it cites.
Chen Cheng and Andrea Montanari, Dimension free ridge regression , arXiv:2210.08571 (2022)
2022
Later among the works it cites.
Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi, Neural networks can learn representations with gradient descent , Conference on Learning Theory, PMLR, 2022, pp. 5413–5452
2022
Later among the works it cites.
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mézard, and Lenka Zdeborová, The gaussian equivalence of generative models for learning with shallow neural networks , Mathematical and Scientific Machine Learning, PMLR, 2022, pp. 426–471
2022
Later among the works it cites.
2022
Later among the works it cites.
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani, Surprises in high-dimensional ridgeless least squares interpolation , The Annals of Statistics 50
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Andrea Montanari and Basil N Saeed, Universality of empirical risk minimization , Conference on Learning Theory, PMLR, 2022, pp. 4310–4312
2022
Later among the works it cites.
Andrea Montanari and Yiqiao Zhong, The interpolation phase transition in neural networks: Memorization and generalization under lazy training , The Annals of Statistics 50
2022
Later among the works it cites.
Ohad Shamir, The implicit bias of benign overfitting , Conference on Learning Theory, PMLR, 2022, pp. 448–478
2022
Later among the works it cites.
Lechao Xiao, Hong Hu, Theodor Misiakiewicz, Yue Lu, and Jeffrey Pennington, Precise learning curves and higher-order scalings for dot-product kernel regression , Advances in Neural Information Processing Systems 35
2022
Later among the works it cites.
Lechao Xiao, Eigenspace restructuring: a principle of space and frequency in neural networks , Conference on Learning Theory, PMLR, 2022, pp. 4888–4944
2022
Later among the works it cites.
Emmanuel Abbe, Enric Boix Adserà, and Theodor Misiakiewicz, Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics , The Thirty Sixth Annual Conference on Learning Theory, PMLR, 2023, pp. 2552–2623
2023
Closest in time.