Fetching the paper…
Reading the bibliography…
We consider the optimization problem associated with fitting two-layer ReLU networks with $k$ hidden neurons, where labels are assumed to be generated by a (teacher) neural network.
1901
Earlier work this paper cites.
1909
Earlier work this paper cites.
1912
Earlier work this paper cites.
F Rosenblatt. ‘The Perceptron: A Probabilistic Model For Information Storage And Organization In The Brain’. Psych. Rev. 65
1958
Earlier work this paper cites.
M Artin. ‘On the solutions of analytic equations’, Invent. Math. 5
1968
Earlier work this paper cites.
J-C Tougeron. ‘Idéaux de fonctions différentiable’, Ann. Inst. Fourier 18
1968
Earlier work this paper cites.
M Minsky and S Papert. Perceptrons: An introduction to Computational Geometry (MIT press, 1969)
1969
Earlier work this paper cites.
L Michel. ‘Minima of Higgs-Landau polynomials’, Regards sur la Physique contemporaine , CNRS, Paris (1980), 157–203
1980
Earlier work this paper cites.
L Michel. ‘Symmetry defects and broken symmetry’, Rev. in Mod. Phys 52
1980
Earlier work this paper cites.
M Golubitsky. ‘The Bénard problem, symmetry and the lattice of isotropy subgroups’, Bifurcation Theory, Mechanics and Physics (eds C P Bruter et al. ) (D Reidel, Dordrecht-Boston-Lancaster, 1983), 225–257
1983
Earlier work this paper cites.
M Aschbacher and L L Scott. ‘Maximal subgroups of finite groups’, J. Algebra 92
1985
Earlier work this paper cites.
T Bröcker and T Tom Dieck. Representations of Compact Lie Groups
1985
Earlier work this paper cites.
M W Liebeck, C E Praeger, & J Saxl. ‘A classification of the Maximal Subgroups of the Finite Alternating and Symmetric Groups’, J. Algebra 111
1987
Earlier work this paper cites.
M J Field. ‘Equivariant Bifurcation Theory and Symmetry Breaking’, J. Dynamics and Diff. Eqns. 1
1989
Earlier work this paper cites.
M J Field and R W Richardson. ‘Symmetry breaking in equivariant bifurcation problems’, Bull. Am. Math. Soc. 22
1990
Earlier work this paper cites.
Y LeCun, B E Boser, J S Denker, D Henerson, R E Howard, W E Hubbard, & L D Jackel. ‘Handwritten digit recognition with a back-propagation network’, Advances in neural information processing systems (1990), 396–404
1990
Earlier work this paper cites.
M J Field and R W Richardson. ‘Symmetry breaking and branching patterns in equivariant bifurcation theory II’, Arch. Rational Mech. and Anal
1992
Earlier work this paper cites.
S G Krantz and H R Parks. A Primer of Real Analytic Functions (Basler Lehrbücher, vol. 4, Birkhäuser Verlag, Basel, Boston, Berlin, 1992)
1992
Earlier work this paper cites.
H S Seung, H Sompolinsky, & N Tishby. ‘Statistical mechanics of learning from examples’, Phys. Rev. A 45
1992
Earlier work this paper cites.
J J Rotman. An introduction to the theory of groups (Springer-Verlag, Graduate Texts in Mathematics, 148
1995
Earlier work this paper cites.
J D Dixon and B Mortimer. Permutation Groups (Graduate texts in mathematics 163
1996
Cited alongside, same era.
A Pinkus, ‘Approximation theory of the MLP model in neural networks’, Acta Numer. 8
1999
Cited alongside, same era.
2003
Cited alongside, same era.
C B Thomas. Representations of Finite and Lie groups (Imperial College Press, 2004)
2004
Cited alongside, same era.
B Newton and B Benesh. ‘A classification of certain maximal subgroups of symmetric groups’, J. of Algebra 304
2006
Cited alongside, same era.
H Hauser. ‘The classical Artin approximation theorems’, Bull. AMS 54
2017
Later among the works it cites.
Y Li and Y Yuan. ‘Convergence analysis of two-layer neural networks with relu activation’, Advances in Neural Information Processing Systems (2017), 597–607
2017
Later among the works it cites.
2017
Later among the works it cites.
S Sonoda and N Murata. ‘Neural network with unbounded activation functions is universal approximator’, Applied and Computational Harmonic Analysis 43
2017
Later among the works it cites.
G Świrszcz, W M Czarnecki & R Pascanu. ‘Local minima in training of neural Networks’, preprint, 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2007
Cited alongside, same era.
2008
Cited alongside, same era.
Y Cho and L K Saul. ‘Kernel Methods for Deep Learning’, Advances in neural information processing systems (2009), 342–350
2009
Cited alongside, same era.
X Glorot and Y Bengio. ‘Understanding the difficulty of training deep feedforward neural networks’, In Proc. AISTATS 9
2010
Cited alongside, same era.
Y N Dauphin, R Pascanu, C Gulcehre, K Cho, S Ganguli, & Y Bengio. ‘Identifying and attacking the saddle point problem in high-dimensional non-convex optimization’, Advances in neural information processing systems (2014), 2933–2941
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2015
Cited alongside, same era.
Y Tian. ‘An analytical formula of population gradient for two-layered ReLU network and its applications in convergence and critical point analysis’, Proc. of the 34th Int. Conf. on Machine Learning 70
2017
Later among the works it cites.
B Xie, Y Liang, & L Song. ‘Diverse Neural Network Learns True Target Functions’, Proc. of the 20th Int. Conf. on Artificial Intelligence and Statistics (2017), 1216–1224
2017
Later among the works it cites.
2017
Later among the works it cites.
A Brutzkus, A Globerson, E Malach, & S Shalev-Shwartz. ‘SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data’, (in 6th International Conference on Learning Representations , ICLR 2018, Vancouver, BC, Canada, April 30– May 3, 2018, Conf. Track Proc., 2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
S S Du, J D Lee, Y Tian, A Singh, & B Póczos. ‘Gradient descent learns one-hidden-layer CNN: don’t be afraid of Spurious Local Minima’, Proc. of the 35th International Conference on Machine Learning (2018), 1338–1347
2018
Later among the works it cites.
R Ge, J D Lee, & T Ma. ‘Learning one-hidden-layer neural networks with landscape design’, (in 6th Int. Conf. on Learning Representations, ICLR 2018 , Conf. Track Proc., 2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
Y Li and Y Liang. ‘Learning overparameterized neural networks via stochastic gradient descent on structured data’, Advances in Neural Information Processing Systems (2018), 8157–8166
2018
Later among the works it cites.
2018
Later among the works it cites.
R Panigrahy, A Rahimi, S Sachdeva, & Q Zhang. ‘Convergence Results for Neural Networks via Electrodynamics’, (in 9th Innovations in Theoretical Computer Science Conference, ITCS 2018 , January 11-14, 2018, Cambridge, MA, USA, 2018), 22:1–22:19
2018
Later among the works it cites.
I Safran and O Shamir. ‘Spurious Local Minima are Common in Two-Layer ReLU Neural Networks’, Proc. of the 35th Int. Conf. on Machine Learning 80
2018
Later among the works it cites.
M Soltanolkotabi, A Javanmard, & D Jason. ‘Theoretical insights into the optimization landscape of over-parameterized shallow neural networks’, IEEE Trans. on Inform. Th. 65
2018
Later among the works it cites.
S S Du, X Zhai, & B Póczos.‘Gradient descent provably optimizes over-parameterized neural networks’, (in 7th Int. Conf. on Learning Representations, ICLR 2019 , New Orleans, LA, USA, May 6–9, 2019),
2019
Later among the works it cites.