Fetching the paper…
Reading the bibliography…
Multi-layer neural networks are among the most powerful models in machine learning, yet the fundamental reasons for this success defy mathematical understanding.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Principles of neurodynamics
Frank Rosenblatt · 1962
Earlier work this paper cites.
Linear and quasi-linear equations of parabolic type
Olga Aleksandrovna Ladyzhenskaia, Vsevolod Alekseevich Solonnikov, and Nina N Ural’tseva · 1988
Earlier work this paper cites.
Topics in propagation of chaos
Alain-Sol Sznitman · 1991
Earlier work this paper cites.
Efficient agnostic learning of neural networks with bounded fan-in
Wee Sun Lee, Peter L Bartlett, and Robert C Williamson · 1996
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L Bartlett · 1998
Earlier work this paper cites.
The variational formulation of the fokker–planck equation
Richard Jordan, David Kinderlehrer, and Felix Otto · 1998
Earlier work this paper cites.
Thermodynamics of glasses: A first principles computation
Marc Mézard and Giorgio Parisi · 1999
Earlier work this paper cites.
On the trend to equilibrium for the fokker-planck equation: an interplay between physics and functional analysis
Peter A Markowich and Cédric Villani · 2000
Earlier work this paper cites.
Kinetic equilibration rates for granular media and related equations: entropy dissipation and mass transportation estimates
José A Carrillo, Robert J McCann, Cédric Villani, et al · 2003
Earlier work this paper cites.
Convex neural networks
Yoshua Bengio, Nicolas L Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte · 2006
Earlier work this paper cites.
Contractions in the 2-wasserstein length space and thermalization of granular media
José A Carrillo, Robert J McCann, and Cédric Villani · 2006
Earlier work this paper cites.
Gradient flows: in metric spaces and in the space of probability measures
Luigi Ambrosio, Nicola Gigli, and Giuseppe Savaré · 2008
Earlier work this paper cites.
Partial Differential Equations
Lawrence C. Evans · 2009
Cited alongside, same era.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Cited alongside, same era.
Differential topology
Victor Guillemin and Alan Pollack · 2010
Cited alongside, same era.
Global-in-time weak measure solutions and finite-time aggregation for nonlocal interaction equations
José A Carrillo, Marco DiFrancesco, Alessio Figalli, Thomas Laurent, Dejan Slepčev, et al · 2011
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Complex analysis
Serge Lang · 2013
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Later among the works it cites.
The landscape of empirical risk for non-convex losses
Song Mei, Yu Bai, and Andrea Montanari · 2016
Later among the works it cites.
L1-regularized neural networks are improperly learnable in polynomial time
Yuchen Zhang, Jason D. Lee, and Michael I. Jordan · 2016
Later among the works it cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Later among the works it cites.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sanjeev Arora, Aditya Bhaskara, Rong Ge, and Tengyu Ma · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Cited alongside, same era.
The zero set of a real analytic function
Boris Mityagin · 2015
Cited alongside, same era.
Provable methods for training neural networks with sparse connectivity
Hanie Sedghi and Anima Anandkumar · 2015
Cited alongside, same era.
Optimal Transport for Applied Mathematicians: Calculus of Variations, PDEs, and Modeling
Filippo Santambrogio · 2015
Cited alongside, same era.
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D. Lee · 2017
Later among the works it cites.
Symmetry-breaking convergence analysis of certain two-layered neural networks with ReLU nonlinearity
Yuandong Tian · 2017
Later among the works it cites.
Chuang Wang, Jonathan Mattingly, and Yue M Lu · 2017
Later among the works it cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L. Bartlett, and Inderjit S. Dhillon · 2017
Later among the works it cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Closest in time.
Grant M Rotskoff and Eric Vanden-Eijnden · 2018
Closest in time.
Mean field analysis of neural networks
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Closest in time.