Fetching the paper…
Reading the bibliography…
The goal of this work is to shed light on the remarkable phenomenon of transition to linearity of certain neural networks as their width approaches infinity.
“On Riemannian manifolds admitting a function whose gradient is of constant norm”
Takashi Sakai · 1996
Earlier work this paper cites.
“Adaptive estimation of a quadratic functional by model selection”
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
“Distributional and L q L^{q} norm inequalities for polynomials over convex bodies in ℝ n \mathbb{R}^{n} ”
Anthony Carbery and James Wright · 2001
Earlier work this paper cites.
“Introduction to the non-asymptotic analysis of random matrices”
Roman Vershynin · 2010
Earlier work this paper cites.
“Efficient backprop”
Yann LeCun, Léon Bottou, Genevieve Orr and Klaus-Robert Müller · 2012
Earlier work this paper cites.
“Topics in random matrix theory”
Terence Tao · 2012
Earlier work this paper cites.
“Stochastic gradient descent optimizes over-parameterized deep relu networks”
Difan Zou, Yuan Cao, Dongruo Zhou and Quanquan Gu · 2012
Earlier work this paper cites.
Mathematics Exchange URL:https://math.stackexchange.com/q/868044 (version: 2014-07-15)
2014
Earlier work this paper cites.
“Deep residual learning for image recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Cited alongside, same era.
“Searching for activation functions”
Prajit Ramachandran, Barret Zoph and Quoc Le · 2017
Cited alongside, same era.
“Gradient descent provably optimizes over-parameterized neural networks”
Simon Du, Xiyu Zhai, Barnabas Poczos and Aarti Singh · 2018
Cited alongside, same era.
“Neural tangent kernel: Convergence and generalization in neural networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Cited alongside, same era.
“A Convergence Theory for Deep Learning via Over-Parameterization”
Zeyuan Allen-Zhu, Yuanzhi Li and Zhao Song · 2019
Cited alongside, same era.
“Gradient Descent Finds Global Minima of Deep Neural Networks”
Simon Du, Jason Lee, Haochuan Li, Liwei Wang and Xiyu Zhai · 2019
Later among the works it cites.
“Linearized two-layers neural networks in high dimension”
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz and Andrea Montanari · 2019
Later among the works it cites.
Ziwei Ji and Matus Telgarsky · 2019
Later among the works it cites.
“Wide neural networks of any depth evolve as linear models under gradient descent”
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein and Jeffrey Pennington · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks”
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li and Ruosong Wang · 2019
Cited alongside, same era.
“Gradient descent with identity initialization efficiently learns positive-definite linear transformations by deep residual networks”
Peter Bartlett, David Helmbold and Philip Long · 2019
Cited alongside, same era.
“On lazy training in differentiable programming”
Lenaic Chizat, Edouard Oyallon and Francis Bach · 2019
Cited alongside, same era.
Ruoyu Sun · 2019
Later among the works it cites.
Shun-ichi Amari · 2020
Closest in time.
Chaoyue Liu, Libin Zhu and Mikhail Belkin · 2020
Closest in time.