Fetching the paper…
Reading the bibliography…
How initialization and loss function affect the learning of a deep neural network (DNN), specifically its generalization error, is an important problem in practice.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
Analysis of the gradient descent algorithm for a deep neural network model with skip-connections
Weinan E, Chao Ma, Qingcan Wang, and Lei Wu · 1904
Earlier work this paper cites.
Weinan E, Chao Ma, and Lei Wu · 1904
Earlier work this paper cites.
Xvi. functions of positive and negative type, and their connection the theory of integral equations
James Mercer · 1909
Earlier work this paper cites.
Hedonic housing prices and the demand for clean air
David Harrison Jr and Daniel L Rubinfeld · 1978
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Reproducing kernel hilbert spaces in probability and, 2004
A Berlinet and C Thomas-Agnan · 2004
Earlier work this paper cites.
Partial differential equations
Lawrence C Evans · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Ensemble deep learning for regression and time series forecasting
Xueheng Qiu, Le Zhang, Ye Ren, Ponnuthurai N Suganthan, and Gehan Amaratunga · 2014
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
A note on lazy training in supervised differentiable programming
Lenaic Chizat and Francis Bach · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
The deep ritz method: A deep learning-based numerical algorithm for solving variational problems
Deep potential molecular dynamics: a scalable model with the accuracy of quantum mechanics
Linfeng Zhang, Jiequn Han, Han Wang, Roberto Car, and E Weinan · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Later among the works it cites.
Theory iii: Dynamics and generalization in deep networks
Andrzej Banburski, Qianli Liao, Brando Miranda, Lorenzo Rosasco, Bob Liang, Jack Hidary, and Tomaso Poggio · 2019
Closest in time.
A generalization theory of gradient descent for learning over-parameterized deep relu networks
Yuan Cao and Quanquan Gu · 2019
Closest in time.
Scaling description of generalization with number of parameters in deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weinan E and Bing Yu · 2018
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2018
Cited alongside, same era.
Theory of deep learning iii: the non-overfitting puzzle
T Poggio, K Kawaguchi, Q Liao, B Miranda, L Rosasco, X Boix, J Hidary, and HN Mhaskar · 2018
Cited alongside, same era.
On the spectral bias of deep neural networks
Nasim Rahaman, Devansh Arpit, Aristide Baratin, Felix Draxler, Min Lin, Fred A Hamprecht, Yoshua Bengio, and Aaron Courville · 2018
Cited alongside, same era.
Training behavior of deep neural network in frequency domain
Zhi-Qin J Xu, Yaoyu Zhang, and Yanyang Xiao · 2018
Cited alongside, same era.
Frequency principle in deep learning with general loss functions and its potential application
Zhi-Qin John Xu
Cited in the paper.
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Closest in time.
Meta-learners’ learning dynamics are unlike learners’
Neil C Rabinowitz · 2019
Closest in time.
Karthik A Sankararaman, Soham De, Zheng Xu, W Ronny Huang, and Tom Goldstein · 2019
Closest in time.
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma · 2019
Closest in time.