Fetching the paper…
Reading the bibliography…
We prove a general Embedding Principle of loss landscape of deep neural networks (NNs) that unravels a hierarchical structure of the loss landscape of NNs, i.e., loss landscape of an NN contains all critical points of all the narrower NNs.
Reflections after refereeing papers for nips
Leo Breiman · 1995
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
Kenji Fukumizu and Shun-ichi Amari · 2000
Earlier work this paper cites.
Matrix analysis
Roger A Horn and Charles R Johnson · 2012
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Singularity of the hessian in deep learning
Levent Sagun, Léon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Jorge Nocedal, Ping Tak Peter Tang, Dheevatsa Mudigere, and Mikhail Smelyanskiy · 2017
Earlier work this paper cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, and Weinan E · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis R. Bach · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
Neural tangent kernel: convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Earlier work this paper cites.
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Sub-optimal local minima exist for neural networks with almost all non-linear activations
Tian Ding, Dawei Li, and Ruoyu Sun · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Spurious valleys in one-hidden-layer neural network optimization landscapes
Luca Venturi, Afonso S Bandeira, and Joan Bruna · 2019
Later among the works it cites.
Training behavior of deep neural network in frequency domain
Zhi-Qin J Xu, Yaoyu Zhang, and Yanyang Xiao · 2019
Later among the works it cites.
Towards a mathematical understanding of neural network-based machine learning: What we know and what we don’t
Weinan E, Chao Ma, Lei Wu, and Stephan Wojtowytsch · 2020
Later among the works it cites.
Modeling the influence of data structure on learning in neural networks: The hidden manifold model
Sebastian Goldt, Marc Mézard, Florent Krzakala, and Lenka Zdeborová · 2020
Later among the works it cites.
Assessing the bilingual knowledge learned by neural machine translation models
Shilin He, Xing Wang, Shuming Shi, Michael R Lyu, and Zhaopeng Tu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Semi-flat minima and saddle points by embedding neural networks to overparameterization
Kenji Fukumizu, Shoichiro Yamaguchi, Yoh-ichi Mototake, and Mirai Tanaka · 2019
Cited alongside, same era.
Asymmetric valleys: beyond sharp and flat local minima
Haowei He, Gao Huang, and Yang Yuan · 2019
Cited alongside, same era.
Sgd on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Cited alongside, same era.
First-order methods almost always avoid strict saddle points
Jason D Lee, Ioannis Panageas, Georgios Piliouras, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2019
Cited alongside, same era.
Neural networks are a priori biased towards boolean functions with low entropy
Chris Mingard, Joar Skalse, Guillermo Valle-Pérez, David Martínez-Rubio, Vladimir Mikulik, and Ard A Louis · 2019
Cited alongside, same era.
On the spectral bias of deep neural networks
Nasim Rahaman, Devansh Arpit, Aristide Baratin, Felix Draxler, Min Lin, Fred A Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Cited alongside, same era.
Loss landscape sightseeing with multi-point optimization
Ivan Skorokhodov and Mikhail Burtsev · 2019
Cited alongside, same era.
Quantifying the generalization error in deep learning in terms of data distribution and neural network smoothness
Pengzhan Jin, Lu Lu, Yifa Tang, and George Em Karniadakis · 2020
Later among the works it cites.
An analytic theory of shallow networks dynamics for hinge loss classification
Franco Pellegrini and Giulio Biroli · 2020
Later among the works it cites.
The global landscape of neural networks: An overview
Ruoyu Sun, Dawei Li, Shiyu Liang, Tian Ding, and Rayadurgam Srikant · 2020
Later among the works it cites.
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma · 2020
Later among the works it cites.
Global minima of overparameterized neural networks
Yaim Cooper · 2021
Closest in time.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances
Berfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro, Clement Hongler, Wulfram Gerstner, and Johanni Brea · 2021
Closest in time.
Towards understanding the condensation of two-layer neural networks at initial training
Zhi-Qin John Xu, Hanxu Zhou, Tao Luo, and Yaoyu Zhang · 2021
Closest in time.