Fetching the paper…
Reading the bibliography…
We study the optimization problem associated with fitting two-layer ReLU neural networks with respect to the squared loss, where labels are generated by a target network.
Minima of higgs-landau polynomials
L Michel · 1979
Earlier work this paper cites.
The benard problem, symmetry and the lattice of isotropy subgroups
Martin Golubitsky · 1983
Earlier work this paper cites.
Spin glass theory and beyond: An Introduction to the Replica Method and Its Applications
Marc Mézard, Giorgio Parisi, and Miguel Angel Virasoro · 1987
Earlier work this paper cites.
Symmetry breaking and the maximal isotropy subgroup conjecture for reflection groups
M. J. Field and R. W. Richardson · 1989
Earlier work this paper cites.
Three unfinished works on the optimal storage capacity of networks
Elizabeth Gardner and Bernard Derrida · 1989
Earlier work this paper cites.
Improving a network generalization ability by selecting examples
W Kinzel and Pal Rujan · 1990
Earlier work this paper cites.
Stochastic gradient learning in neural networks
Léon Bottou · 1991
Earlier work this paper cites.
Representation theory, volume 129 of
William Fulton and Joe Harris · 1991
Earlier work this paper cites.
Training a 3-node neural network is np-complete
Avrim L Blum and Ronald L Rivest · 1992
Earlier work this paper cites.
Statistical mechanics of learning from examples
Hyunjune Sebastian Seung, Haim Sompolinsky, and Naftali Tishby · 1992
Earlier work this paper cites.
The statistical mechanics of learning a rule
Timothy LH Watkin, Albrecht Rau, and Michael Biehl · 1993
Earlier work this paper cites.
Learning by on-line gradient descent
Michael Biehl and Holm Schwarze · 1995
Earlier work this paper cites.
On-line backpropagation in two-layered neural networks
Peter Riegler and Michael Biehl · 1995
Earlier work this paper cites.
Exact solution for on-line learning in multilayer neural networks
David Saad and Sara A Solla · 1995
Earlier work this paper cites.
On-line learning in soft committee machines
David Saad and Sara A Solla · 1995
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
Representations of finite and Lie groups
Charles Benedict Thomas · 2004
Earlier work this paper cites.
Dynamics and symmetry
Michael J. Field · 2007
Earlier work this paper cites.
Information, physics, and computation
Marc Mezard and Andrea Montanari · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Learnability, stability and uniform convergence
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2010
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Cited alongside, same era.
Proof of the satisfiability conjecture for large k
Jian Ding, Allan Sly, and Nike Sun · 2015
Cited alongside, same era.
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size
Vardan Papyan · 2018
Later among the works it cites.
The spectrum of the fisher information matrix of a single-hidden-layer neural network
Jeffrey Pennington and Pratik Worah · 2018
Later among the works it cites.
Spurious local minima are common in two-layer relu neural networks
Itay Safran and Ohad Shamir · 2018
Later among the works it cites.
Distribution-specific hardness of learning neural networks
Ohad Shamir · 2018
Later among the works it cites.
Hessian-based analysis of large batch training and robustness to adversaries
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Porcupine neural networks:(almost) all local optima are global
Soheil Feizi, Hamid Javadi, Jesse Zhang, and David Tse · 2017
Cited alongside, same era.
Three factors influencing minima in sgd
Stanis 𝒍 {\bm{l}} aw Jastrz 𝒌 {\bm{k}} ebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Nonlinear random matrix theory for deep learning
Jeffrey Pennington and Pratik Worah · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Cited alongside, same era.
Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W Mahoney · 2018
Later among the works it cites.
On the principle of least symmetry breaking in shallow relu models
Yossi Arjevani and Michael Field · 2019
Later among the works it cites.
The committee machine: Computational to statistical gaps in learning a two-layers neural network
Benjamin Aubin, Antoine Maillard, Jean Barbier, Florent Krzakala, Nicolas Macris, and Lenka Zdeborová · 2019
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2019
Later among the works it cites.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Later among the works it cites.
Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
Sebastian Goldt, Madhu S Advani, Andrew M Saxe, Florent Krzakala, and Lenka Zdeborová · 2019
Later among the works it cites.
Generalisation dynamics of online learning in over-parameterised neural networks
Sebastian Goldt, Madhu S Advani, Andrew M Saxe, Florent Krzakala, and Lenka Zdeborová · 2019
Later among the works it cites.
Universal statistics of fisher information in deep neural networks: Mean field approach
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
Analytic characterization of the hessian in shallow relu models: A tale of symmetry
Yossi Arjevani and Michael Field · 2020
Later among the works it cites.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Later among the works it cites.
Optimization and generalization of shallow neural networks with quadratic activation functions
Stefano Sarao Mannelli, Eric Vanden-Eijnden, and Lenka Zdeborová · 2020
Later among the works it cites.
On learnability via gradient method for two-layer relu neural networks in teacher-student setting
Shunta Akiyama and Taiji Suzuki · 2021
Closest in time.
Symmetry breaking in symmetric tensor decomposition
Yossi Arjevani, Joan Bruna, Michael Field, Joe Kileel, Matthew Trager, and Francis Williams · 2021
Closest in time.
Yossi Arjevani and Michael Field · 2021
Closest in time.
Symmetry & critical points for a model shallow neural network
Yossi Arjevani and Michael Field · 2021
Closest in time.
Hidden unit specialization in layered neural networks: Relu vs. sigmoidal activation
Elisa Oostwal, Michiel Straat, and Michael Biehl · 2021
Closest in time.