Fetching the paper…
Reading the bibliography…
Empirical works show that for ReLU neural networks (NNs) with small initialization, input weights of hidden neurons (the input weight of a hidden neuron consists of the weight from its input layer to the hidden neuron and its bias term) condense on isolated orientations.
A multi-scale dnn algorithm for nonlinear elliptic equations with multiple scales,
X.-A. Li, Z.-Q. J. Xu, L. Zhang, · 1906
Earlier work this paper cites.
Reflections after refereeing papers for nips,
L. Breiman, · 1995
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons,
K. Fukumizu, S.-i. Amari, · 2000
Earlier work this paper cites.
Multi-scale deep neural network (mscalednn) for solving poisson-boltzmann equation in complex domains,
Z. Liu, W. Cai, Z.-Q. J. Xu, · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results,
P. L. Bartlett, S. Mendelson, · 2002
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks,
X. Glorot, Y. Bengio, · 2010
Earlier work this paper cites.
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, J. Sun, · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks,
D. Arpit, S. Jastrzebski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. Courville, Y. Bengio, et al., · 2017
Earlier work this paper cites.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks,
G. M. Rotskoff, E. Vanden-Eijnden, · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport,
L. Chizat, F. Bach, · 2018
Earlier work this paper cites.
Visualizing the loss landscape of neural nets,
H. Li, Z. Xu, G. Taylor, C. Studer, T. Goldstein, · 2018
Earlier work this paper cites.
Gradient descent quantizes relu network features,
H. Maennel, O. Bousquet, S. Gelly, · 2018
Earlier work this paper cites.
Neural tangent kernel: convergence and generalization in neural networks,
A. Jacot, F. Gabriel, C. Hongler, · 2018
Earlier work this paper cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit,
S. Mei, T. Misiakiewicz, A. Montanari, · 2019
Earlier work this paper cites.
On the spectral bias of neural networks,
N. Rahaman, A. Baratin, D. Arpit, F. Draxler, M. Lin, F. Hamprecht, Y. Bengio, A. Courville, · 2019
Cited alongside, same era.
Fantastic generalization measures and where to find them,
Y. Jiang, B. Neyshabur, H. Mobahi, D. Krishnan, S. Bengio, · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net,
S. Arora, S. S. Du, W. Hu, Z. Li, R. R. Salakhutdinov, R. Wang, · 2019
Cited alongside, same era.
A note on lazy training in supervised differentiable programming,
L. Chizat, F. Bach, · 2019
Cited alongside, same era.
Semi-flat minima and saddle points by embedding neural networks to overparameterization,
K. Fukumizu, S. Yamaguchi, Y.-i. Mototake, M. Tanaka, · 2019
Cited alongside, same era.
Training behavior of deep neural network in frequency domain,
An analytic theory of shallow networks dynamics for hinge loss classification,
F. Pellegrini, G. Biroli, · 2020
Later among the works it cites.
A type of generalization error induced by initialization in deep neural networks,
Y. Zhang, Z.-Q. J. Xu, T. Luo, Z. Ma, · 2020
Later among the works it cites.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics,
W. E, C. Ma, L. Wu, · 2020
Later among the works it cites.
Quantifying the generalization error in deep learning in terms of data distribution and neural network smoothness,
P. Jin, L. Lu, Y. Tang, G. E. Karniadakis, · 2020
Later among the works it cites.
The slow deterioration of the generalization error of the random feature model,
C. Ma, L. Wu, E. Weinan, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z.-Q. J. Xu, Y. Zhang, Y. Xiao, · 2019
Cited alongside, same era.
On the spectral bias of deep neural networks,
N. Rahaman, D. Arpit, A. Baratin, F. Draxler, M. Lin, F. A. Hamprecht, Y. Bengio, A. Courville, · 2019
Cited alongside, same era.
Sgd on neural networks learns functions of increasing complexity,
D. Kalimeris, G. Kaplun, P. Nakkiran, B. Edelman, T. Yang, B. Barak, H. Zhang, · 2019
Cited alongside, same era.
Mean field analysis of neural networks: A central limit theorem,
J. Sirignano, K. Spiliopoulos, · 2020
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks,
M. S. Advani, A. M. Saxe, H. Sompolinsky, · 2020
Cited alongside, same era.
The inductive bias of relu networks on orthogonally separable data,
M. Phuong, C. H. Lampert, · 2020
Cited alongside, same era.
Frequency principle: Fourier analysis sheds light on deep neural networks,
Z.-Q. J. Xu, Y. Zhang, T. Luo, Y. Xiao, Z. Ma, · 2020
Cited alongside, same era.
A phase shift deep neural network for high frequency approximation and wave problems,
W. Cai, X. Li, L. Liu, · 2020
Later among the works it cites.
Fourier features let networks learn high frequency functions in low dimensional domains,
M. Tancik, P. Srinivasan, B. Mildenhall, S. Fridovich-Keil, N. Raghavan, U. Singhal, R. Ramamoorthi, J. Barron, R. Ng, · 2020
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization,
C. Zhang, S. Bengio, M. Hardt, B. Recht, O. Vinyals, · 2021
Closest in time.
Phase diagram for two-layer relu neural networks at infinite-width limit,
T. Luo, Z.-Q. J. Xu, Z. Ma, Y. Zhang, · 2021
Closest in time.
Gradient descent on two-layer nets: Margin maximization and simplicity bias,
K. Lyu, Z. Li, R. Wang, S. Arora, · 2021
Closest in time.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances,
B. Simsek, F. Ged, A. Jacot, F. Spadaro, C. Hongler, W. Gerstner, J. Brea, · 2021
Closest in time.
Deep frequency principle towards understanding why deeper learning is faster,
Z.-Q. J. Xu, H. Zhou, · 2021
Closest in time.
Subspace decomposition based dnn algorithm for elliptic-type multi-scale pdes,
X.-A. Li, Z.-Q. J. Xu, L. Zhang, · 2021
Closest in time.
Empirical phase diagram for three-layer neural networks with infinite width,
H. Zhou, Q. Zhou, Z. Jin, T. Luo, Y. Zhang, Z.-Q. J. Xu, · 2022
Closest in time.