Fetching the paper…
Reading the bibliography…
Fitting a function by using linear combinations of a large number $N$ of `simple' components is one of the most fruitful ideas in statistical learning.
Imre Bihari, A generalization of a lemma of Bellman and its application to uniqueness problems of differential equations , Acta Mathematica Hungarica 7
1956
Earlier work this paper cites.
Frank Rosenblatt, Principles of neurodynamics , Spartan Book, 1962
1962
Earlier work this paper cites.
Roland L’vovich Dobrushin, Vlasov equations , Functional Analysis and Its Applications 13
1979
Earlier work this paper cites.
Hiroshi Tanaka, Stochastic differential equations with reflecting boundary condition in convex regions , Hiroshima Mathematical Journal 9
1979
Earlier work this paper cites.
Pierre-Louis Lions and Alain-Sol Sznitman, Stochastic differential equations with reflecting boundary conditions , Communications on Pure and Applied Mathematics 37
1984
Earlier work this paper cites.
Olga Aleksandrovna Ladyzhenskaia, Vsevolod Alekseevich Solonnikov, and Nina N. Ural’tseva, Linear and quasi-linear equations of parabolic type , vol. 23, American Mathematical Society, 1988
1988
Earlier work this paper cites.
George Cybenko, Approximation by superpositions of a sigmoidal function , Mathematics of control, signals and systems 2
1989
Earlier work this paper cites.
Jooyoung Park and Irwin W. Sandberg, Universal approximation using radial-basis-function networks , Neural computation 3
1991
Earlier work this paper cites.
Alain-Sol Sznitman, Topics in propagation of chaos , Ecole d’été de probabilités de Saint-Flour XIX—1989, Springer, 1991, pp. 165–251
1991
Earlier work this paper cites.
David L. Donoho, Superresolution via sparsity constraints , SIAM Journal on Mathematical Analysis 23
1992
Earlier work this paper cites.
Andrew R. Barron, Universal approximation bounds for superpositions of a sigmoidal function , IEEE Transactions on Information theory 39
1993
Earlier work this paper cites.
L. Chris G. Rogers and David Williams, Diffusions, Markov processes and martingales: Volume 2, Itô calculus , vol. 2, Cambridge university press, 1994
1994
Earlier work this paper cites.
Slomiński, Leszek, On approximation of solutions of multidimensional SDE’s with reflecting boundary conditions , Stochastic processes and their Applications 50
1994
Earlier work this paper cites.
Robert J McCann, A convexity principle for interacting gases , Advances in mathematics 128
1997
Earlier work this paper cites.
Peter L. Bartlett, The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network , IEEE Transactions on Information Theory 44
1998
Earlier work this paper cites.
Nello Cristianini and John Shawe-Taylor, An introduction to support vector machines and other kernel-based learning methods , Cambridge University Press, 2000
2000
Earlier work this paper cites.
José A. Carrillo, Ansgar Jüngel, Peter A. Markowich, Giuseppe Toscani, and Andreas Unterreiter, Entropy dissipation methods for degenerate parabolicproblems and generalized sobolev inequalities , Monatshefte für Mathematik 133
2001
Earlier work this paper cites.
Jerome H. Friedman, Greedy function approximation: a gradient boosting machine , Annals of Statistics (2001), 1189–1232
2001
Earlier work this paper cites.
Karl Oelschläger, A sequence of integro-differential equations approximating a viscous porous medium equation , Zeitschrift für Analysis und ihre Anwendungen 20
2001
Earlier work this paper cites.
Peter Bühlmann and Bin Yu, Boosting with the L2 loss: regression and classification , Journal of the American Statistical Association 98
2003
Cited alongside, same era.
José A. Carrillo, Robert J. McCann, and Cédric Villani, Kinetic equilibration rates for granular media and related equations: entropy dissipation and mass transportation estimates , Revista Matematica Iberoamericana 19
2003
Cited alongside, same era.
Robert E. Schapire, The boosting approach to machine learning: An overview , Nonlinear estimation and classification, Springer, 2003, pp. 149–171
2003
Cited alongside, same era.
Yoshua Bengio, Nicolas L. Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte, Convex neural networks , Advances in Neural Information Processing Systems, 2006, pp. 123–130
2006
Cited alongside, same era.
Thomas M Cover and Joy A Thomas, Elements of information theory , John Wiley & Sons, 2006
Emmanuel J. Candès and Carlos Fernandez-Granda, Towards a mathematical theory of super-resolution , Communications on Pure and Applied Mathematics 67
2014
Later among the works it cites.
Filippo Santambrogio, Optimal transport for applied mathematicians: Calculus of variations, PDEs, and modeling , vol. 87, Birkhäuser, 2015
2015
Later among the works it cites.
Yining Chen and Richard J Samworth, Generalized additive and index models with shape constraints , Journal of the Royal Statistical Society: Series B (Statistical Methodology) 78
2016
Later among the works it cites.
Francis Bach, Breaking the curse of dimensionality with convex neural networks , The Journal of Machine Learning Research 18
2017
Later among the works it cites.
Yuanzhi Li and Yang Yuan, Convergence analysis of two-layer neural networks with relu activation , Advances in Neural Information Processing Systems, 2017, pp. 597–607
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2006
Cited alongside, same era.
Robert Philipowski, Interacting diffusions approximating the porous medium equation and propagation of chaos , Stochastic Processes and their Applications 117
2007
Cited alongside, same era.
Juan Luis Vázquez, The porous medium equation: mathematical theory , Oxford University Press, 2007
2007
Cited alongside, same era.
Luigi Ambrosio, Nicola Gigli, and Giuseppe Savaré, Gradient flows: in metric spaces and in the space of probability measures , Springer Science & Business Media, 2008
2008
Cited alongside, same era.
Alessio Figalli and Robert Philipowski, Convergence to the viscous porous medium equation and propagation of chaos , ALEA Lat. Am. J. Probab. Math. Stat 4
2008
Cited alongside, same era.
Ali Rahimi and Benjamin Recht, Random features for large-scale kernel machines , Advances in neural information processing systems, 2008, pp. 1177–1184
2008
Cited alongside, same era.
Cédric Villani, Optimal transport: old and new , vol. 338, Springer Science & Business Media, 2008
2008
Cited alongside, same era.
Martin Anthony and Peter L. Bartlett, Neural network learning: Theoretical foundations , Cambridge University Press, 2009
2009
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
Yuandong Tian, Symmetry-breaking convergence analysis of certain two-layered neural networks with ReLU nonlinearity , Workshop at International Conference on Learning Representation (ICLR), 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
2018
Later among the works it cites.
Lenaic Chizat and Francis Bach, On the global convergence of gradient descent for over-parameterized models using optimal transport , Advances in Neural Information Processing Systems, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
Song Mei, Yu Bai, Andrea Montanari, et al., The landscape of empirical risk for nonconvex losses , The Annals of Statistics 46
2018
Later among the works it cites.
Song Mei, Andrea Montanari, and Phan-Minh Nguyen, A mean field view of the landscape of two-layer neural networks , Proceedings of the National Academy of Sciences (2018)
2018
Later among the works it cites.
Grant M. Rotskoff and Eric Vanden-Eijnden, Neural networks as interacting particle systems: Asymptotic convexity of the loss landscape and universal scaling of the approximation error , Advances in Neural Information Processing Systems, 2018
2018
Later among the works it cites.
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D. Lee, Theoretical insights into the optimization landscape of over-parameterized shallow neural networks , IEEE Transactions on Information Theory (2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit , Conference on Learning Theory (COLT), 2019
2019
Closest in time.