Fetching the paper…
Reading the bibliography…
While deep learning is successful in a number of applications, it is not yet well understood theoretically.
Shpigel Nacson M, Gunasekar S, Lee JD, Srebro N, Soudry D (2019) Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models · 1905
Earlier work this paper cites.
Daubechies I, DeVore R, Foucart S, Hanin B, Petrova G (2019) Nonlinear approximation and (deep) relu networks · 1905
Earlier work this paper cites.
Advances in Computational Mathematics
Mhaskar H (1993) Approximation properties of a multilayered feedforward artificial neural network · 1993
Earlier work this paper cites.
(IEEE), pp. 190–196
Mhaskar HN (1993) Neural networks for localized approximation of real functions in Neural Networks for Processing [1993] III. Proceedings of the 1993 IEEE-SP Workshop · 1993
Earlier work this paper cites.
Mathematics of Computation
Chui C, Li X, Mhaskar H (1994) Neural networks for localized approximation · 1994
Earlier work this paper cites.
Advances in Computational Mathematics
Chui CK, Li X, Mhaskar HN (1996) Limitations of the approximation capabilities of neural networks with one hidden layer · 1996
Earlier work this paper cites.
Ferreira PJSG (1996) The existence and uniqueness of the minimum norm solution to certain linear and nonlinear problems · 1996
Earlier work this paper cites.
Acta Numerica
Pinkus A (1999) Approximation theory of the mlp model in neural networks · 1999
Earlier work this paper cites.
Donoho DL (2000) High-dimensional data analysis: The curses and blessings of dimensionality in AMS CONFERENCE ON MATH CHALLENGES OF THE 21ST CENTURY
2000
Earlier work this paper cites.
IEEE Transactions on Signal Processing
Douglas SC, Amari S, Kung SY (2000) On gradient adaptation with unit-norm constraints · 2000
Earlier work this paper cites.
Notices of the American Mathematical Society (AMS)
Poggio T, Smale S (2003) The mathematics of learning: Dealing with data · 2003
Earlier work this paper cites.
pp. 169–207
Bousquet O, Boucheron S, Lugosi G (2003) Introduction to statistical learning theory · 2003
Earlier work this paper cites.
pp. 1237–1244
Rosset S, Zhu J, Hastie T (2003) Margin maximizing loss functions in Advances in Neural Information Processing Systems 16 [Neural Information Processing Systems, NIPS 2003, December 8-13, 2003, Vancouver and Whistler, British Columbia, Canada] · 2003
Earlier work this paper cites.
Advances in Neural Information Processing Systems
Montufar, G. F.and Pascanu R, Cho K, Bengio Y (2014) On the number of linear regions of deep neural networks · 2014
Earlier work this paper cites.
Center for Brains, Minds and Machines (CBMM) Memo No. 1. arXiv:1311.4158v5
Anselmi F, et al. (2014) Unsupervised learning of invariant representations with low sample complexity: the magic of sensory cortex or a new framework for machine learning? · 2014
Earlier work this paper cites.
Center for Brains, Minds and Machines (CBMM) Memo No. 35, also in arXiv
Anselmi F, Rosasco L, Tan C, Poggio T (2015) Deep convolutional network are hierarchical kernel machines · 2015
Cited alongside, same era.
Poggio T, Rosasco L, Shashua A, Cohen N, Anselmi F (2015) Notes on hierarchical splines, dclns and i-theory, (MIT Computer Science and Artificial Intelligence Laboratory), Technical report
2015
Cited alongside, same era.
CBMM memo 041
Poggio T, Anselmi F, Rosasco L (2015) I-theory on depth vs width: hierarchical function composition · 2015
Cited alongside, same era.
Theoretical Computer Science
Anselmi F, et al. (2015) Unsupervised learning of invariant representations · 2015
Cited alongside, same era.
CBMM memo 037
Poggio T, Rosaco L, Shashua A, Cohen N, Anselmi F (2015) Notes on hierarchical splines, dclns and i-theory · 2015
Cited alongside, same era.
(Curran Associates Inc., USA), pp. 597–607
Li Y, Yuan Y (2017) Convergence analysis of two-layer neural networks with relu activation in Proceedings of the 31st International Conference on Neural Information Processing Systems · 2017
Later among the works it cites.
pp. 605–614
Brutzkus A, Globerson A (2017) Globally optimal gradient descent for a convnet with gaussian inputs in Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 · 2017
Later among the works it cites.
(JMLR.org), pp. 4140–4149
Zhong K, Song Z, Jain P, Bartlett PL, Dhillon IS (2017) Recovery guarantees for one-hidden-layer neural networks in Proceedings of the 34th International Conference on Machine Learning - Volume 70 · 2017
Later among the works it cites.
arXiv:1703.09833, CBMM Memo No. 066
Poggio T, Liao Q (2017) Theory II: Landscape of the empirical risk in deep learning · 2017
Later among the works it cites.
CBMM Memo 072
Zhang C, et al. (2017) Theory of deep learning IIb: Optimization properties of SGD · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Telgarsky M (2015) Representation benefits of deep feedforward networks · 2015
Cited alongside, same era.
Center for Brains, Minds and Machines (CBMM) Memo No. 45, also in arXiv
Mhaskar H, Liao Q, Poggio T (2016) Learning real and boolean functions: When is deep better than shallow? · 2016
Cited alongside, same era.
Center for Brains, Minds and Machines (CBMM) Memo No. 54, also in arXiv
Mhaskar H, Poggio T (2016) Deep versus shallow networks: an approximation theory perspective · 2016
Cited alongside, same era.
Center for Brains, Minds and Machines (CBMM) Memo No. 47, also in arXiv
Liao Q, Poggio T (2016) Bridging the gap between residual learning, recurrent neural networks and visual cortex · 2016
Cited alongside, same era.
Safran I, Shamir O (2016) Depth separation in relu networks for approximating smooth non-linear functions · 2016
Cited alongside, same era.
Poggio T, Mhaskar H, Rosasco L, Miranda B, Liao Q (2016) Theory I: Why and when can deep - but not shallow - networks avoid the curse of dimensionality, (CBMM Memo No. 058, MIT Center for Brains, Minds and Machines), Technical report
2016
Cited alongside, same era.
(PMLR, Columbia University, New York, New York, USA), Vol. 49, pp. 1246–1257
Lee JD, Simchowitz M, Jordan MI, Recht B (2016) Gradient descent only converges to minimizers in 29th Annual Conference on Learning Theory · 2016
Cited alongside, same era.
arXiv:180.3251 [cs, math]
Raginsky M, Rakhlin A, Telgarsky M (2017) Non-convex learning via stochastic gradient langevin dynamics: A nonasymptotic analysis · 2017
Later among the works it cites.
(Curran Associates, Inc.), pp. 2422–2430
Daniely A (2017) Sgd learns the conjugate kernel class of the network in Advances in Neural Information Processing Systems 30 · 2017
Later among the works it cites.
Du SS, Lee JD, Tian Y (2018) When is a convolutional filter easy to learn? in International Conference on Learning Representations
2018
Later among the works it cites.
(PMLR, Stockholmsmässan, Stockholm Sweden), Vol. 80, pp. 1339–1348
Du S, Lee J, Tian Y, Singh A, Poczos B (2018) Gradient descent learns one-hidden-layer CNN: Don’t be afraid of spurious local minima in Proceedings of the 35th International Conference on Machine Learning · 2018
Later among the works it cites.
arXiv e-prints
Zhang X, Yu Y, Wang L, Gu Q (2018) Learning One-hidden-layer ReLU Networks via Gradient Descent · 2018
Later among the works it cites.
(Curran Associates, Inc.), pp. 8157–8166
Li Y, Liang Y (2018) Learning overparameterized neural networks via stochastic gradient descent on structured data in Advances in Neural Information Processing Systems 31 · 2018
Later among the works it cites.
CBMM Memo No. 090
Banburski A, et al. (2019) Theory of deep learning III: Dynamics and generalization in deep networks · 2019
Closest in time.
IEEE Transactions on Information Theory
Soltanolkotabi M, Javanmard A, Lee JD (2019) Theoretical insights into the optimization landscape of over-parameterized shallow neural networks · 2019
Closest in time.
Du SS, Zhai X, Poczos B, Singh A (2019) Gradient descent provably optimizes over-parameterized neural networks in International Conference on Learning Representations
2019
Closest in time.