Fetching the paper…
Reading the bibliography…
For a certain scaling of the initialization of stochastic gradient descent (SGD), wide neural networks (NN) have been shown to be well approximated by reproducing kernel Hilbert space (RKHS) methods.
David L Donoho and Iain M Johnstone, Adapting to unknown smoothness via wavelet shrinkage , Journal of the American Statistical Association 90
1995
Earlier work this paper cites.
Ali Rahimi and Benjamin Recht, Random features for large-scale kernel machines , Advances in neural information processing systems, 2008, pp. 1177–1184
2008
Earlier work this paper cites.
Alexandre B Tsybakov, Introduction to nonparametric estimation , Springer Science & Business Media, 2008
2008
Earlier work this paper cites.
2014
Earlier work this paper cites.
Sergey Ioffe and Christian Szegedy, Batch normalization: Accelerating deep network training by reducing internal covariate shift , International Conference on Machine Learning, 2015, pp. 448–456
2015
Earlier work this paper cites.
Francis Bach, Breaking the curse of dimensionality with convex neural networks , The Journal of Machine Learning Research 18
2017
Earlier work this paper cites.
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, and Skye Wanderman-Milne, JAX: composable transformations of Python+NumPy programs , 2018
2018
Earlier work this paper cites.
Lenaic Chizat and Francis Bach, On the global convergence of gradient descent for over-parameterized models using optimal transport , Advances in neural information processing systems, 2018, pp. 3036–3046
2018
Earlier work this paper cites.
AGG De Matthews, J Hron, M Rowland, RE Turner, and Z Ghahramani, Gaussian process behaviour in wide deep neural networks , 6th International Conference on Learning Representations, ICLR 2018-Conference Track Proceedings, 2018
2018
Earlier work this paper cites.
Arthur Jacot, Franck Gabriel, and Clément Hongler, Neural tangent kernel: Convergence and generalization in neural networks , Advances in neural information processing systems, 2018, pp. 8571–8580
2018
Earlier work this paper cites.
Jaehoon Lee, Jascha Sohl-dickstein, Jeffrey Pennington, Roman Novak, Sam Schoenholz, and Yasaman Bahri, Deep neural networks as gaussian processes , International Conference on Learning Representations, 2018
2018
Earlier work this paper cites.
Song Mei, Yu Bai, and Andrea Montanari, The landscape of empirical risk for nonconvex losses , The Annals of Statistics 46
2018
Earlier work this paper cites.
Song Mei, Andrea Montanari, and Phan-Minh Nguyen, A mean field view of the landscape of two-layer neural networks , Proceedings of the National Academy of Sciences 115
2018
Earlier work this paper cites.
David Page, Myrtle.ai , https://myrtle.ai/how-to-train-your-resnet-4-architecture/, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang, On exact computation with an infinitely wide neural net , Advances in Neural Information Processing Systems, 2019, pp. 8139–8148
2019
Later among the works it cites.
2019
Later among the works it cites.
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington, Wide neural networks of any depth evolve as linear models under gradient descent , Advances in neural information processing systems, 2019, pp. 8570–8581
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Zeyuan Allen-Zhu and Yuanzhi Li, What can resnet learn efficiently, going beyond kernels? , Advances in Neural Information Processing Systems, 2019, pp. 9017–9028
2019
Cited alongside, same era.
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song, A convergence theory for deep learning via over-parameterization , Proceedings of the 36th International Conference on Machine Learning (Long Beach, California, USA) (Kamalika Chaudhuri and Ruslan Salakhutdinov, eds.), Proceedings of Machine Learning Research, vol. 97, PMLR, 09–15 Jun 2019, pp. 242–252
2019
Cited alongside, same era.
Lenaic Chizat, Edouard Oyallon, and Francis Bach, On lazy training in differentiable programming , Advances in Neural Information Processing Systems, 2019, pp. 2933–2943
2019
Cited alongside, same era.
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai, Gradient descent finds global minima of deep neural networks , Proceedings of the 36th International Conference on Machine Learning (Long Beach, California, USA) (Kamalika Chaudhuri and Ruslan Salakhutdinov, eds.), Proceedings of Machine Learning Research, vol. 97, PMLR, 09–15 Jun 2019, pp. 1675–1685
2019
Cited alongside, same era.
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh, Gradient descent provably optimizes over-parameterized neural networks , International Conference on Learning Representations, 2019
2019
Cited alongside, same era.
Adrià Garriga-Alonso, Carl Edward Rasmussen, and Laurence Aitchison, Deep convolutional networks as shallow gaussian processes , International Conference on Learning Representations, 2019
2019
Cited alongside, same era.
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao, An investigation into neural net optimization via hessian eigenvalue density , International Conference on Machine Learning, 2019, pp. 2232–2241
2019
Cited alongside, same era.
Roman Novak, Lechao Xiao, Yasaman Bahri, Jaehoon Lee, Greg Yang, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-dickstein, Bayesian deep convolutional networks with many channels are gaussian processes , International Conference on Learning Representations, 2019
2019
Later among the works it cites.
Dong Yin, Raphael Gontijo Lopes, Jon Shlens, Ekin Dogus Cubuk, and Justin Gilmer, A fourier perspective on model robustness in computer vision , Advances in Neural Information Processing Systems, 2019, pp. 13255–13265
2019
Later among the works it cites.
Gilad Yehudai and Ohad Shamir, On the power and limitations of random features for understanding neural networks , Advances in Neural Information Processing Systems, 2019, pp. 6594–6604
2019
Later among the works it cites.
Sanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov, Ruosong Wang, and Dingli Yu, Harnessing the power of infinitely wide deep nets on small-data tasks , International Conference on Learning Representations, 2020
2020
Closest in time.
2020
Closest in time.
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz, Neural tangents: Fast and easy infinite neural networks in python , International Conference on Learning Representations, 2020
2020
Closest in time.
Samet Oymak and Mahdi Soltanolkotabi, Towards moderate overparameterization: global convergence guarantees for training shallow neural networks , IEEE Journal on Selected Areas in Information Theory (2020)
2020
Closest in time.