Fetching the paper…
Reading the bibliography…
We consider learning two layer neural networks using stochastic gradient descent.
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, Dropout: a simple way to prevent neural networks from overfitting , The Journal of Machine Learning Research 15
1958
Earlier work this paper cites.
George Cybenko, Approximation by superpositions of a sigmoidal function , Mathematics of control, signals and systems 2
1989
Earlier work this paper cites.
Alain-Sol Sznitman, Topics in propagation of chaos , Ecole d’été de probabilités de Saint-Flour XIX—1989, Springer, 1991, pp. 165–251
1991
Earlier work this paper cites.
Andrew R Barron, Universal approximation bounds for superpositions of a sigmoidal function , IEEE Transactions on Information theory 39
1993
Earlier work this paper cites.
Peter L Bartlett, The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network , IEEE transactions on Information Theory 44
1998
Earlier work this paper cites.
Richard Jordan, David Kinderlehrer, and Felix Otto, The variational formulation of the fokker–planck equation , SIAM journal on mathematical analysis 29
1998
Earlier work this paper cites.
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner, Gradient-based learning applied to document recognition , Proceedings of the IEEE 86
1998
Earlier work this paper cites.
Yoshua Bengio, Nicolas L Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte, Convex neural networks , Advances in neural information processing systems, 2006, pp. 123–130
2006
Earlier work this paper cites.
Martin Anthony and Peter L Bartlett, Neural network learning: Theoretical foundations , cambridge university press, 2009
2009
Earlier work this paper cites.
Lawrence C. Evans, Partial differential equations , Springer, 2009
2009
Cited alongside, same era.
Filippo Santambrogio, Optimal transport for applied mathematicians: Calculus of variations, pdes, and modeling , vol. 87, Birkhäuser, 2015
2015
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Yuanzhi Li and Yingyu Liang, Learning overparameterized neural networks via stochastic gradient descent on structured data , Advances in Neural Information Processing Systems, 2018, pp. 8168–8177
2018
Later among the works it cites.
Song Mei, Andrea Montanari, and Phan-Minh Nguyen, A mean field view of the landscape of two-layer neural networks , Proceedings of the National Academy of Sciences (2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Later among the works it cites.
2019
Closest in time.
2019
Closest in time.