Fetching the paper…
Reading the bibliography…
Understanding the structure of loss landscape of deep neural networks (DNNs)is obviously important.
The use of multiple measurements in taxonomic problems,
R. A. Fisher, · 1936
Earlier work this paper cites.
Reflections after refereeing papers for nips,
L. Breiman, · 1995
Earlier work this paper cites.
The loss surfaces of multilayer networks,
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, Y. LeCun, · 2015
Earlier work this paper cites.
Singularity of the hessian in deep learning,
L. Sagun, L. Bottou, Y. LeCun, · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization,
C. Zhang, S. Bengio, M. Hardt, B. Recht, O. Vinyals, · 2017
Earlier work this paper cites.
A closer look at memorization in deep networks,
D. Arpit, S. Jastrzebski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. C. Courville, Y. Bengio, S. Lacoste-Julien, · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima,
N. S. Keskar, J. Nocedal, P. T. P. Tang, D. Mudigere, M. Smelyanskiy, · 2017
Earlier work this paper cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes,
L. Wu, Z. Zhu, et al., · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport,
L. Chizat, F. R. Bach, · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks,
A. Jacot, C. Hongler, F. Gabriel, · 2018
Earlier work this paper cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks,
D. Zou, Y. Cao, D. Zhou, Q. Gu, · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks,
S. Mei, A. Montanari, P.-M. Nguyen, · 2018
Earlier work this paper cites.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks,
G. M. Rotskoff, E. Vanden-Eijnden, · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks,
J. Frankle, M. Carbin, · 2018
Earlier work this paper cites.
Loss landscape sightseeing with multi-point optimization,
I. Skorokhodov, M. Burtsev, · 2019
Cited alongside, same era.
SGD on neural networks learns functions of increasing complexity,
D. Kalimeris, G. Kaplun, P. Nakkiran, B. L. Edelman, T. Yang, B. Barak, H. Zhang, · 2019
Cited alongside, same era.
Training behavior of deep neural network in frequency domain,
Z.-Q. J. Xu, Y. Zhang, Y. Xiao, · 2019
Cited alongside, same era.
On the spectral bias of deep neural networks,
N. Rahaman, D. Arpit, A. Baratin, F. Draxler, M. Lin, F. A. Hamprecht, Y. Bengio, A. Courville, · 2019
Cited alongside, same era.
Asymmetric valleys: Beyond sharp and flat local minima,
H. He, G. Huang, Y. Yuan, · 2019
Cited alongside, same era.
Frequency principle: Fourier analysis sheds light on deep neural networks,
Z.-Q. J. Xu, Y. Zhang, T. Luo, Y. Xiao, Z. Ma, · 2020
Later among the works it cites.
C. Ma, L. Wu, W. E, · 2020
Later among the works it cites.
A type of generalization error induced by initialization in deep neural networks,
Y. Zhang, Z.-Q. J. Xu, T. Luo, Z. Ma, · 2020
Later among the works it cites.
Mean field analysis of neural networks: A central limit theorem,
J. Sirignano, K. Spiliopoulos, · 2020
Later among the works it cites.
Modeling the influence of data structure on learning in neural networks: The hidden manifold model,
S. Goldt, M. Mézard, F. Krzakala, L. Zdeborová, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Arora, S. S. Du, W. Hu, Z. Li, R. Salakhutdinov, R. Wang, · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks,
S. Du, J. Lee, H. Li, L. Wang, X. Zhai, · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization,
Z. Allen-Zhu, Y. Li, Z. Song, · 2019
Cited alongside, same era.
W. E, C. Ma, L. Wu, · 2019
Cited alongside, same era.
Neural networks are a priori biased towards boolean functions with low entropy,
C. Mingard, J. Skalse, G. Valle-Pérez, D. Martínez-Rubio, V. Mikulik, A. A. Louis, · 2019
Cited alongside, same era.
Semi-flat minima and saddle points by embedding neural networks to overparameterization,
K. Fukumizu, S. Yamaguchi, Y.-i. Mototake, M. Tanaka, · 2019
Cited alongside, same era.
Gradient dynamics of shallow univariate relu networks,
M. Trager, C. Silva, D. Panozzo, D. Zorin, J. Bruna, · 2019
Cited alongside, same era.
S. He, X. Wang, S. Shi, M. R. Lyu, Z. Tu, · 2020
Later among the works it cites.
A dynamical central limit theorem for shallow neural networks,
Z. Chen, G. M. Rotskoff, J. Bruna, E. Vanden-Eijnden, · 2020
Later among the works it cites.
Disentangling feature and lazy training in deep neural networks,
M. Geiger, S. Spigler, A. Jacot, M. Wyart, · 2020
Later among the works it cites.
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel,
S. Fort, G. K. Dziugaite, M. Paul, S. Kharaghani, D. M. Roy, S. Ganguli, · 2020
Later among the works it cites.
Phase diagram for two-layer relu neural networks at infinite-width limit,
T. Luo, Z.-Q. J. Xu, Z. Ma, Y. Zhang, · 2021
Closest in time.
Global minima of overparameterized neural networks,
Y. Cooper, · 2021
Closest in time.
Embedding principle: a hierarchical structure of loss landscape of deep neural networks,
Y. Zhang, Y. Li, Z. Zhang, T. Luo, Z. J. Xu, · 2021
Closest in time.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances,
B. Simsek, F. Ged, A. Jacot, F. Spadaro, C. Hongler, W. Gerstner, J. Brea, · 2021
Closest in time.