Fetching the paper…
Reading the bibliography…
Substantial work indicates that the dynamics of neural networks (NNs) is closely related to their initialization of parameters.
Reflections after refereeing papers for nips,
L. Breiman, · 1995
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results,
P. L. Bartlett, S. Mendelson, · 2002
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks,
X. Glorot, Y. Bengio, · 2010
Earlier work this paper cites.
Efficient backprop,
Y. A. LeCun, L. Bottou, G. B. Orr, K.-R. Müller, · 2012
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,
K. He, X. Zhang, S. Ren, J. Sun, · 2015
Earlier work this paper cites.
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, J. Sun, · 2016
Earlier work this paper cites.
Neural tangent kernel: convergence and generalization in neural networks,
A. Jacot, F. Gabriel, C. Hongler, · 2018
Earlier work this paper cites.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks,
G. M. Rotskoff, E. Vanden-Eijnden, · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport,
L. Chizat, F. Bach, · 2018
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization,
S. Arora, N. Cohen, E. Hazan, · 2018
Earlier work this paper cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks,
D. Zou, Y. Cao, D. Zhou, Q. Gu, · 2018
Earlier work this paper cites.
Gradient descent quantizes relu network features,
H. Maennel, O. Bousquet, S. Gelly, · 2018
Earlier work this paper cites.
A note on lazy training in supervised differentiable programming,
L. Chizat, F. Bach, · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net,
S. Arora, S. S. Du, W. Hu, Z. Li, R. R. Salakhutdinov, R. Wang, · 2019
Cited alongside, same era.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit,
S. Mei, T. Misiakiewicz, A. Montanari, · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization,
Z. Allen-Zhu, Y. Li, Z. Song, · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks,
S. S. Du, X. Zhai, B. Póczos, A. Singh, · 2019
Cited alongside, same era.
Training behavior of deep neural network in frequency domain,
Z.-Q. J. Xu, Y. Zhang, Y. Xiao, · 2019
Tensor programs ii: Neural tangent kernel for any architecture,
G. Yang, · 2020
Later among the works it cites.
Frequency principle: Fourier analysis sheds light on deep neural networks,
Z.-Q. J. Xu, Y. Zhang, T. Luo, Y. Xiao, Z. Ma, · 2020
Later among the works it cites.
An analytic theory of shallow networks dynamics for hinge loss classification,
F. Pellegrini, G. Biroli, · 2020
Later among the works it cites.
Phase diagram for two-layer relu neural networks at infinite-width limit,
T. Luo, Z.-Q. J. Xu, Z. Ma, Y. Zhang, · 2021
Later among the works it cites.
Deep frequency principle towards understanding why deeper learning is faster,
Z.-Q. J. Xu, H. Zhou, · 2021
Later among the works it cites.
Tensor programs iib: Architectural universality of neural tangent kernel training dynamics,
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the spectral bias of deep neural networks,
N. Rahaman, D. Arpit, A. Baratin, F. Draxler, M. Lin, F. A. Hamprecht, Y. Bengio, A. Courville, · 2019
Cited alongside, same era.
A type of generalization error induced by initialization in deep neural networks,
Y. Zhang, Z.-Q. J. Xu, T. Luo, Z. Ma, · 2020
Cited alongside, same era.
Mean field analysis of neural networks: A central limit theorem,
J. Sirignano, K. Spiliopoulos, · 2020
Cited alongside, same era.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics,
W. E, C. Ma, L. Wu, · 2020
Cited alongside, same era.
Disentangling feature and lazy training in deep neural networks,
M. Geiger, S. Spigler, A. Jacot, M. Wyart, · 2020
Cited alongside, same era.
Embedding principle of loss landscape of deep neural networks,
Y. Zhang, Z. Zhang, T. Luo, Z. Xu,
Cited in the paper.
G. Yang, E. Littwin, · 2021
Later among the works it cites.
On feature learning in shallow and multi-layer neural networks with global convergence guarantees,
Z. Chen, E. Vanden-Eijnden, J. Bruna, · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization,
C. Zhang, S. Bengio, M. Hardt, B. Recht, O. Vinyals, · 2021
Later among the works it cites.
Gradient descent on two-layer nets: Margin maximization and simplicity bias,
K. Lyu, Z. Li, R. Wang, S. Arora, · 2021
Later among the works it cites.
Towards Understanding the Condensation of Neural Networks at Initial Training,
H. Zhou, Q. Zhou, T. Luo, Y. Zhang, Z.-Q. J. Xu, · 2021
Later among the works it cites.
Overview frequency principle/spectral bias in deep learning,
Z.-Q. J. Xu, Y. Zhang, T. Luo, · 2022
Closest in time.