Fetching the paper…
Reading the bibliography…
We describe the new field of mathematical analysis of deep learning.
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei, Language models are few-shot learners , Advances in Neural Information Processing Systems, 2020, pp. 1877–1901
1901
Earlier work this paper cites.
1901
Earlier work this paper cites.
1902
Earlier work this paper cites.
1902
Earlier work this paper cites.
1903
Earlier work this paper cites.
1904
Earlier work this paper cites.
1906
Earlier work this paper cites.
1908
Earlier work this paper cites.
1908
Earlier work this paper cites.
1909
Earlier work this paper cites.
1912
Earlier work this paper cites.
1912
Earlier work this paper cites.
Paul Adrien Maurice Dirac, Quantum mechanics of many-electron systems , Proceedings of the Royal Society of London. Series A, Containing Papers of a Mathematical and Physical Character 123
1929
Earlier work this paper cites.
Hassler Whitney, Analytic extensions of differentiable functions defined in closed sets , Transactions of the American Mathematical Society 36
1934
Earlier work this paper cites.
Kurt Lewin, Psychology and the process of group living , The Journal of Social Psychology 17
1943
Earlier work this paper cites.
Warren S McCulloch and Walter Pitts, A logical calculus of the ideas immanent in nervous activity , The Bulletin of Mathematical Biophysics 5
1943
Earlier work this paper cites.
Herbert Robbins and Sutton Monro, A stochastic approximation method , The Annals of Mathematical Statistics (1951), 400–407
1951
Earlier work this paper cites.
Richard Bellman, On the theory of dynamic programming , Proceedings of the National Academy of Sciences 38
1952
Earlier work this paper cites.
Jack Kiefer and Jacob Wolfowitz, Stochastic estimation of the maximum of a regression function , The Annals of Mathematical Statistics 23
1952
Earlier work this paper cites.
Frank Rosenblatt, The perceptron: a probabilistic model for information storage and organization in the brain , Psychological review 65
1958
Earlier work this paper cites.
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, Dropout: a simple way to prevent neural networks from overfitting , Journal of Machine Learning Research 15
1958
Earlier work this paper cites.
Henry J Kelley, Gradient theory of optimal flight paths , Ars Journal 30
1960
Earlier work this paper cites.
Stuart Dreyfus, The numerical solution of variational problems , Journal of Mathematical Analysis and Applications 5
1962
Earlier work this paper cites.
Wassily Hoeffding, Probability inequalities for sums of bounded random variables , Journal of the American Statistical Association 58
1963
Earlier work this paper cites.
Richard M Dudley, The sizes of compact subsets of hilbert space and continuity of Gaussian processes , Journal of Functional Analysis 1
1967
Earlier work this paper cites.
William F Donoghue, Distributions and fourier transforms , Pure and Applied Mathematics, Academic Press, 1969
1969
Earlier work this paper cites.
Marvin Minsky and Seymour A Papert, Perceptrons , MIT Press, 1969
1969
Earlier work this paper cites.
Seppo Linnainmaa, Alogritmin kumulatiivinen pyöristysvirhe yksittäisten pyöristysvirheiden Taylor-kehitelmänä , Master’s thesis, University of Helsinki, 1970
1970
Earlier work this paper cites.
Vladimir Vapnik and Alexey Chervonenkis, On the uniform convergence of relative frequencies of events to their probabilities , Theory of Probability & Its Applications 16
1971
Earlier work this paper cites.
Thomas Zaslavsky, Facing up to arrangements: Face-count formulas for partitions of space by hyperplanes: Face-count formulas for partitions of space by hyperplanes , Memoirs of the American Mathematical Society, American Mathematical Society, 1975
1975
Earlier work this paper cites.
John J Hopfield, Neural networks and physical systems with emergent collective computational abilities , Proceedings of the National Academy of Sciences 79
1982
Earlier work this paper cites.
Arkadi Semenovich Nemirovsky and David Borisovich Yudin, Problem complexity and method efficiency in optimization , Wiley-Interscience Series in Discrete Mathematics, Wiley, 1983
1983
Earlier work this paper cites.
Evarist Giné and Joel Zinn, Some limit theorems for empirical processes , The Annals of Probability (1984), 929–989
1984
Earlier work this paper cites.
David H Ackley, Geoffrey E Hinton, and Terrence J Sejnowski, A learning algorithm for Boltzmann machines , Cognitive Science 9
1985
Earlier work this paper cites.
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams, Learning representations by back-propagating errors , Nature 323
1986
Earlier work this paper cites.
Paul J Werbos, Generalization of backpropagation with application to a recurrent gas market model , Neural Networks 1
1988
Earlier work this paper cites.
Eric B Baum and David Haussler, What size net gives valid generalization? , Neural Computation 1
1989
Earlier work this paper cites.
Avrim Blum and Ronald L Rivest, Training a 3-node neural network is NP-complete , Advances in Neural Information Processing Systems, 1989, pp. 494–501
1989
Earlier work this paper cites.
George Cybenko, Approximation by superpositions of a sigmoidal function , Mathematics of Control, Signals and Systems 2
1989
Earlier work this paper cites.
Ken-Ichi Funahashi, On the approximate realization of continuous mappings by neural networks , Neural Networks 2
1989
Earlier work this paper cites.
Kurt Hornik, Maxwell Stinchcombe, and Halbert White, Multilayer feedforward networks are universal approximators , Neural Networks 2
1989
Earlier work this paper cites.
Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel, Backpropagation applied to handwritten zip code recognition , Neural Computation 1
1989
Earlier work this paper cites.
Yann LeCun, John S Denker, and Sara A Solla, Optimal brain damage , Advances in Neural Information Processing Systems, 1989, pp. 598–605
1989
Earlier work this paper cites.
Colin McDiarmid, On the method of bounded differences , Surveys in Combinatorics 141
1989
Earlier work this paper cites.
Jeffrey L Elman, Finding structure in time , Cognitive Science 14
1990
Earlier work this paper cites.
Michael I Jordan, Attractor dynamics and parallelism in a connectionist sequential machine , Artificial neural networks: concept learning, IEEE Press, 1990, pp. 112–127
1990
Earlier work this paper cites.
Stephen J Judd, Neural network design and the complexity of learning , MIT Press, 1990
1990
Earlier work this paper cites.
Michel Ledoux and Michel Talagrand, Probability in Banach spaces: Isoperimetry and processes , vol. 23, Springer Science & Business Media, 1991
1991
Earlier work this paper cites.
Andrew R Barron, Neural net approximation , Yale Workshop on Adaptive and Learning Systems, vol. 1, 1992, pp. 69–72
1992
Earlier work this paper cites.
Etienne Pardoux and Shige Peng, Backward stochastic differential equations and quasilinear parabolic partial differential equations , Stochastic partial differential equations and their applications, Springer, 1992, pp. 200–217
1992
Earlier work this paper cites.
Moshe Leshno, Vladimir Ya Lin, Allan Pinkus, and Shimon Schocken, Multilayer feedforward networks with a nonpolynomial activation function can approximate any function , Neural Networks 6
1993
Earlier work this paper cites.
Charles K Chui, Xin Li, and Hrushikesh N Mhaskar, Neural networks for localized approximation , Mathematics of Computation 63
1994
Earlier work this paper cites.
Geoffrey E Hinton and Richard S Zemel, Autoencoders, minimum description length, and helmholtz free energy , Advances in Neural Information Processing Systems 6
1994
Earlier work this paper cites.
Michel Talagrand, Sharper bounds for Gaussian and empirical processes , The Annals of Probability (1994), 28–76
1994
Earlier work this paper cites.
David Haussler, Sphere packing numbers for subsets of the boolean n-cube with bounded vapnik-chervonenkis dimension , Journal of Combinatorial Theory, Series A 2
1995
Earlier work this paper cites.
Ronald J Williams and David Zipser, Gradient-based learning algorithms for recurrent , Backpropagation: Theory, Architectures, and Applications 433
1995
Earlier work this paper cites.
Peter Auer, Mark Herbster, and Manfred K Warmuth, Exponentially many local minima for single neurons , Advances in Neural Information Processing Systems, 1996, p. 316–322
1996
Earlier work this paper cites.
Luc Devroye, László Györfi, and Gábor Lugosi, A probabilistic theory of pattern recognition , Springer, 1996
1996
Earlier work this paper cites.
Hrushikesh N Mhaskar, Neural networks for optimal approximation of smooth and analytic functions , Neural Computation 8
1996
Earlier work this paper cites.
Bruno A Olshausen and David J Field, Sparse coding of natural images produces localized, oriented, bandpass receptive fields , Nature 381
1996
Earlier work this paper cites.
Sepp Hochreiter and Jürgen Schmidhuber, Long short-term memory , Neural Computation 9
1997
Earlier work this paper cites.
Marek Karpinski and Angus Macintyre, Polynomial bounds for VC dimension of sigmoidal and general Pfaffian neural networks , Journal of Computer and System Sciences 54
1997
Earlier work this paper cites.
Aad W van der Vaart and Jon A Wellner, Weak convergence and empirical processes with applications to statistics , Journal of the Royal Statistical Society-Series A Statistics in Society 160
1997
Earlier work this paper cites.
Peter L Bartlett, The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network , IEEE Transactions on Information Theory 44
1998
Earlier work this paper cites.
Peter L Bartlett, Vitaly Maiorov, and Ron Meir, Almost linear VC-dimension bounds for piecewise polynomial networks , Neural Computation 10
1998
Earlier work this paper cites.
Emmanuel J Candès, Ridgelets: Theory and applications , Ph.D. thesis, Stanford University, 1998
1998
Earlier work this paper cites.
Ronald A DeVore, Nonlinear approximation , Acta Numerica 7
1998
Earlier work this paper cites.
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner, Gradient-based learning applied to document recognition , Proceedings of the IEEE 86
1998
Earlier work this paper cites.
Genevieve B Orr and Klaus-Robert Müller, Neural networks: tricks of the trade , Springer, 1998
1998
Earlier work this paper cites.
Martin Anthony and Peter L Bartlett, Neural network learning: Theoretical foundations , Cambridge University Press, 1999
1999
Earlier work this paper cites.
David A McAllester, Pac-bayesian model averaging , Conference on Learning Theory, 1999, pp. 164–170
1999
Earlier work this paper cites.
Vitaly Maiorov and Allan Pinkus, Lower bounds for approximation by MLP neural networks , Neurocomputing 25
1999
Earlier work this paper cites.
Akito Sakurai, Tight bounds for the VC-dimension of piecewise polynomial networks , Advances in Neural Information Processing Systems, 1999, pp. 323–329
1999
Earlier work this paper cites.
Vladimir Vapnik, An overview of statistical learning theory , IEEE Transactions on Neural Networks 10
1999
Earlier work this paper cites.
2001
Earlier work this paper cites.
Trevor Hastie, Robert Tibshirani, and Jerome Friedman, The elements of statistical learning: Data mining, inference, and prediction , Springer Series in Statistics, Springer, 2001
2001
Earlier work this paper cites.
Olivier Bousquet and André Elisseeff, Stability and generalization , Journal of Machine Learning Research 2
2002
Earlier work this paper cites.
Felipe Cucker and Steve Smale, On the mathematical foundations of learning , Bulletin of the American Mathematical Society 39
2002
Earlier work this paper cites.
Jiří Šíma, Training a single sigmoidal neuron is hard , Neural Computation 14
2002
Earlier work this paper cites.
Olivier Bousquet, Stéphane Boucheron, and Gábor Lugosi, Introduction to statistical learning theory , Summer School on Machine Learning, 2003, pp. 169–207
2003
Earlier work this paper cites.
Shahar Mendelson and Roman Vershynin, Entropy and the combinatorial dimension , Inventiones mathematicae 152
2003
Earlier work this paper cites.
Tomaso Poggio, Ryan Rifkin, Sayan Mukherjee, and Partha Niyogi, General conditions for predictivity in learning theory , Nature 428
2004
Earlier work this paper cites.
Peter L Bartlett, Olivier Bousquet, and Shahar Mendelson, Local Rademacher complexities , The Annals of Statistics 33
2005
Earlier work this paper cites.
P López Ríos, Ao Ma, Neil D Drummond, Michael D Towler, and Richard J Needs, Inhomogeneous backflow transformations in quantum Monte Carlo calculations , Physical Review E 74
2006
Earlier work this paper cites.
Walter Rudin, Real and complex analysis , McGraw-Hill Series in Higher Mathematics, Tata McGraw-Hill, 2006
2006
Earlier work this paper cites.
2007
Earlier work this paper cites.
Felipe Cucker and Ding-Xuan Zhou, Learning theory: an approximation theory viewpoint , vol. 24, Cambridge University Press, 2007
2007
Earlier work this paper cites.
2007
Earlier work this paper cites.
2007
Earlier work this paper cites.
2007
Earlier work this paper cites.
Ali Rahimi, Benjamin Recht, et al., Random features for large-scale kernel machines , Advances in Neural Information Processing Systems, 2007, pp. 1177–1184
2007
Earlier work this paper cites.
2008
Earlier work this paper cites.
2008
Earlier work this paper cites.
Andreas Griewank and Andrea Walther, Evaluating derivatives: principles and techniques of algorithmic differentiation , SIAM, 2008
2008
Earlier work this paper cites.
2009
Earlier work this paper cites.
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, Imagenet: A large-scale hierarchical image database , Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255
2009
Earlier work this paper cites.
2009
Earlier work this paper cites.
Alex Krizhevsky and Geoffrey Hinton, Learning multiple layers of features from tiny images , Tech. report, University of Toronto, 2009
2009
Earlier work this paper cites.
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro, Robust stochastic approximation approach to stochastic programming , SIAM Journal on Optimization 19
2009
Cited alongside, same era.
Erich Novak and Henryk Woźniakowski, Approximation of infinitely differentiable multivariate functions is intractable , Journal of Complexity 25
2009
Cited alongside, same era.
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan, Stochastic convex optimization , Conference on Learning Theory, 2009
2009
Cited alongside, same era.
Harry Yserentant, Regularity and approximability of electronic wave functions , Springer, 2010
2010
Cited alongside, same era.
Alexander G de G Matthews, Jiri Hron, Mark Rowland, Richard E Turner, and Zoubin Ghahramani, Gaussian process behaviour in wide deep neural networks , International Conference on Learning Representations, 2018
2018
Later among the works it cites.
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro, A PAC-Bayesian approach to spectrally-normalized margin bounds for neural networks , International Conference on Learning Representations, 2018
2018
Later among the works it cites.
Hieu Pham, Melody Guan, Barret Zoph, Quoc Le, and Jeff Dean, Efficient neural architecture search via parameters sharing , International Conference on Machine Learning, 2018, pp. 4095–4104
2018
Later among the works it cites.
Vardan Papyan, Yaniv Romano, Jeremias Sulam, and Michael Elad, Theoretical foundations of deep learning via sparse representations: A multilayer sparse model and its connection to convolutional neural networks , IEEE Signal Processing Magazine 35
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
Peter G Casazza, Gitta Kutyniok, and Friedrich Philipp, Introduction to finite frame theory , Finite Frames: Theory and Applications, Birkhäuser Boston, 2012, pp. 1–53
2012
Cited alongside, same era.
2012
Cited alongside, same era.
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, Imagenet classification with deep convolutional neural networks , Advances in Neural Information Processing Systems, 2012, pp. 1097–1105
2012
Cited alongside, same era.
Stéphane Mallat, Group invariant scattering , Communications on Pure and Applied Mathematics 65
2012
Cited alongside, same era.
Attila Szabo and Neil S Ostlund, Modern quantum chemistry: introduction to advanced electronic structure theory , Courier Corporation, 2012
2012
Cited alongside, same era.
Huan Xu and Shie Mannor, Robustness and generalization , Machine learning 86
2012
Cited alongside, same era.
Antonio Auffinger, Gérard Ben Arous, and Jiří Černỳ, Random matrices and complexity of spin glasses , Communications on Pure and Applied Mathematics 66
2013
Cited alongside, same era.
Philipp Petersen and Felix Voigtlaender, Optimal approximation of piecewise smooth functions using deep ReLU neural networks , Neural Networks 108
2018
Later among the works it cites.
Uri Shaham, Alexander Cloninger, and Ronald R Coifman, Provable approximation properties for deep neural networks , Applied and Computational Harmonic Analysis 44
2018
Later among the works it cites.
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli, Analysing mathematical reasoning abilities of neural models , International Conference on Learning Representations, 2018
2018
Later among the works it cites.
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro, The implicit bias of gradient descent on separable data , 2018
2018
Later among the works it cites.
Jeremias Sulam, Vardan Papyan, Yaniv Romano, and Michael Elad, Multilayer convolutional sparse modeling: Pursuit and dictionary learning , IEEE Transactions on Signal Processing 66
2018
Later among the works it cites.
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry, How does batch normalization help optimization? , Advances in Neural Information Processing Systems, 2018, pp. 2488–2498
2018
Later among the works it cites.
2018
Later among the works it cites.
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky, Deep image prior , Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 9446–9454
2018
Later among the works it cites.
Roman Vershynin, High-dimensional probability: An introduction with applications in data science , vol. 47, Cambridge University Press, 2018
2018
Later among the works it cites.
Jong Chul Ye, Yoseob Han, and Eunju Cha, Deep convolutional framelets: A general deep learning framework for inverse problems , SIAM Journal on Imaging Sciences 11
2018
Later among the works it cites.
Tom Young, Devamanyu Hazarika, Soujanya Poria, and Erik Cambria, Recent trends in deep learning based natural language processing , IEEE Computational Intelligence Magazine 13
2018
Later among the works it cites.
2018
Later among the works it cites.
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu, A convergence analysis of gradient descent for deep linear neural networks , International Conference on Learning Representations, 2019
2019
Later among the works it cites.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang, On exact computation with an infinitely wide neural net , Advances in Neural Information Processing Systems, 2019, pp. 8139–8148
2019
Later among the works it cites.
Simon Arridge, Peter Maass, Ozan Öktem, and Carola-Bibiane Schönlieb, Solving inverse problems using data-driven models , Acta Numerica 28
2019
Later among the works it cites.
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song, A convergence theory for deep learning via over-parameterization , International Conference on Machine Learning, 2019, pp. 242–252
2019
Later among the works it cites.
Julius Berner, Dennis Elbrächter, and Philipp Grohs, How degenerate is the parametrization of neural networks with the ReLU activation function? , Advances in Neural Information Processing Systems, 2019, pp. 7790–7801
2019
Later among the works it cites.
Peter L Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian, Nearly-tight VC-dimension and pseudodimension bounds for piecewise linear neural networks , Journal of Machine Learning Research 20
2019
Later among the works it cites.
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off , Proceedings of the National Academy of Sciences 116
2019
Later among the works it cites.
Mikhail Belkin, Alexander Rakhlin, and Alexandre B Tsybakov, Does data interpolation contradict statistical optimality? , International Conference on Artificial Intelligence and Statistics, 2019, pp. 1611–1619
2019
Later among the works it cites.
Minshuo Chen, Haoming Jiang, Wenjing Liao, and Tuo Zhao, Efficient approximation of deep ReLU networks for functions on low dimensional manifolds , Advances in Neural Information Processing Systems, 2019, pp. 8174–8184
2019
Later among the works it cites.
Wojciech Czaja and Weilin Li, Analysis of time-frequency scattering transforms , Applied and Computational Harmonic Analysis 47
2019
Later among the works it cites.
Lenaic Chizat, Edouard Oyallon, and Francis Bach, On lazy training in differentiable programming , Advances in Neural Information Processing Systems, 2019, pp. 2937–2947
2019
Later among the works it cites.
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai, Gradient descent finds global minima of deep neural networks , International Conference on Machine Learning, 2019, pp. 1675–1685
2019
Later among the works it cites.
Weinan E, Martin Hutzenthaler, Arnulf Jentzen, and Thomas Kruse, On multilevel picard numerical approximations for high-dimensional nonlinear parabolic partial differential equations and high-dimensional nonlinear backward stochastic differential equations , Journal of Scientific Computing 79
2019
Later among the works it cites.
Weinan E, Jiequn Han, and Qianxiao Li, A mean-field optimal control formulation of deep learning , Research in the Mathematical Sciences 6
2019
Later among the works it cites.
Davis Gilton, Greg Ongie, and Rebecca Willett, Neumann networks for linear inverse problems in imaging , IEEE Transactions on Computational Imaging 6
2019
Later among the works it cites.
Boris Hanin, Universal function approximation by deep neural nets with bounded width and ReLU activations , Mathematics 7
2019
Later among the works it cites.
Catherine F Higham and Desmond J Higham, Deep learning: An introduction for applied mathematicians , SIAM Review 61
2019
Later among the works it cites.
Boris Hanin and David Rolnick, Deep ReLU networks have surprisingly few activation patterns , Advances in Neural Information Processing Systems, 2019, pp. 359–368
2019
Later among the works it cites.
Peter Hinz and Sara van de Geer, A framework for the construction of upper bounds on the number of affine linear regions of ReLU feed-forward neural networks , IEEE Transactions on Information Theory 65
2019
Later among the works it cites.
Jiequn Han, Linfeng Zhang, and Weinan E, Solving many-electron Schrödinger equation using deep neural networks , Journal of Computational Physics 399
2019
Later among the works it cites.
Yiding Jiang, Dilip Krishnan, Hossein Mobahi, and Samy Bengio, Predicting the generalization gap in deep networks with margin distributions , International Conference on Learning Representations, 2019
2019
Later among the works it cites.
Ziwei Ji and Matus Telgarsky, Gradient descent aligns the layers of deep linear networks , International Conference on Learning Representations, 2019
2019
Later among the works it cites.
Guillaume Lample and François Charton, Deep learning for symbolic mathematics , International Conference on Learning Representations, 2019
2019
Later among the works it cites.
Kaifeng Lyu and Jian Li, Gradient descent maximizes the margin of homogeneous neural networks , International Conference on Learning Representations, 2019
2019
Later among the works it cites.
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes, Fisher–Rao metric, geometry, and complexity of neural networks , International Conference on Artificial Intelligence and Statistics, 2019, pp. 888–896
2019
Later among the works it cites.
Bo Li, Shanshan Tang, and Haijun Yu, Better approximations of high dimensional smooth functions by deep neural networks with rectified power units , Communications in Computational Physics 27
2019
Later among the works it cites.
Vaishnavh Nagarajan and J Zico Kolter, Uniform convergence may be unable to explain generalization in deep learning , Advances in Neural Information Processing Systems, 2019, pp. 11615–11626
2019
Later among the works it cites.
Mor Shpigel Nacson, Jason D Lee, Suriya Gunasekar, Pedro Henrique Pamplona Savarese, Nathan Srebro, and Daniel Soudry, Convergence of gradient descent on separable data , International Conference on Artificial Intelligence and Statistics, 2019, pp. 3420–3428
2019
Later among the works it cites.
Kenta Oono and Taiji Suzuki, Approximation and non-parametric estimation of ResNet-type convolutional neural networks , International Conference on Machine Learning, 2019, pp. 4922–4931
2019
Later among the works it cites.
Lars Ruthotto and Eldad Haber, Deep neural networks motivated by partial differential equations , Journal of Mathematical Imaging and Vision (2019), 1–13
2019
Later among the works it cites.
Maziar Raissi, Paris Perdikaris, and George E Karniadakis, Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations , Journal of Computational Physics 378
2019
Later among the works it cites.
Christoph Schwab and Jakob Zech, Deep learning in high dimension: Neural network expression rates for generalized polynomial chaos expansions in uq , Analysis and Applications 17
2019
Later among the works it cites.
Luca Venturi, Afonso S Bandeira, and Joan Bruna, Spurious valleys in one-hidden-layer neural network optimization landscapes , Journal of Machine Learning Research 20
2019
Later among the works it cites.
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, and Petko Georgiev, Grandmaster level in StarCraft II using multi-agent reinforcement learning , Nature 575
2019
Later among the works it cites.
Julius Berner, Markus Dablander, and Philipp Grohs, Numerically solving parametric families of high-dimensional Kolmogorov partial differential equations via deep learning , Advances in Neural Information Processing Systems, 2020, pp. 16615–16627
2020
Later among the works it cites.
Julius Berner, Philipp Grohs, and Arnulf Jentzen, Analysis of the generalization error: Empirical risk minimization over deep artificial neural networks overcomes the curse of dimensionality in the numerical approximation of black–scholes partial differential equations , SIAM Journal on Mathematics of Data Science 2
2020
Later among the works it cites.
Mikhail Belkin, Daniel Hsu, and Ji Xu, Two models of double descent for weak features , SIAM Journal on Mathematics of Data Science 2
2020
Later among the works it cites.
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler, Benign overfitting in linear regression , Proceedings of the National Academy of Sciences 117
2020
Later among the works it cites.
Lenaic Chizat and Francis Bach, Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss , Conference on Learning Theory, 2020, pp. 1305–1338
2020
Later among the works it cites.
Sören Dittmer, Tobias Kluth, Peter Maass, and Daniel Otero Baguer, Regularization by architecture: A deep prior approach for inverse problems , Journal of Mathematical Imaging and Vision 62
2020
Later among the works it cites.
Philipp Grohs, Fabian Hornung, Arnulf Jentzen, and Philippe Von Wurstemberger, A proof that artificial neural networks overcome the curse of dimensionality in the numerical approximation of Black-Scholes partial differential equations , Memoirs of the American Mathematical Society (2020)
2020
Later among the works it cites.
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart, Scaling description of generalization with number of parameters in deep learning , Journal of Statistical Mechanics: Theory and Experiment (2020), no. 2, 023401
2020
Later among the works it cites.
Ingo Gühring, Gitta Kutyniok, and Philipp Petersen, Error bounds for approximations with deep ReLU neural networks in W s , p {W}^{s,p} norms , Analysis and Applications 18
2020
Later among the works it cites.
Philipp Grohs, Sarah Koppensteiner, and Martin Rathmair, Phase retrieval: Uniqueness and stability , SIAM Review 62
2020
Later among the works it cites.
Lukas Gonon and Christoph Schwab, Deep ReLU network expression rates for option prices in high-dimensional, exponential Lévy models , 2020, ETH Zurich SAM Research Report
2020
Later among the works it cites.
Martin Hutzenthaler, Arnulf Jentzen, Thomas Kruse, and Tuan Anh Nguyen, A proof that rectified deep neural networks overcome the curse of dimensionality in the numerical approximation of semilinear heat equations , SN Partial Differential Equations and Applications 1
2020
Later among the works it cites.
Juncai He, Lin Li, Jinchao Xu, and Chunyue Zheng, ReLU deep neural networks and linear finite elements , Journal of Computational Mathematics 38
2020
Later among the works it cites.
Jan Hermann, Zeno Schätzle, and Frank Noé, Deep-neural-network solution of the electronic Schrödinger equation , Nature Chemistry 12
2020
Later among the works it cites.
Arnulf Jentzen, Benno Kuckuck, Ariel Neufeld, and Philippe von Wurstemberger, Strong error analysis for stochastic gradient descent optimization algorithms , IMA Journal of Numerical Analysis 41
2020
Later among the works it cites.
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio, Fantastic generalization measures and where to find them , International Conference on Learning Representations, 2020
2020
Later among the works it cites.
Patrick Kidger and Terry Lyons, Universal approximation with deep narrow networks , Conference on Learning Theory, 2020, pp. 2306–2327
2020
Later among the works it cites.
Yiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu, and Lexing Ying, A mean field analysis of deep ResNet and beyond: Towards provably optimization via overparameterization from depth , International Conference on Machine Learning, 2020, pp. 6426–6436
2020
Later among the works it cites.
Tengyuan Liang and Alexander Rakhlin, Just interpolate: Kernel “ridgeless” regression can generalize , The Annals of Statistics 48
2020
Later among the works it cites.
Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai, On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels , Conference on Learning Theory, 2020, pp. 2683–2711
2020
Later among the works it cites.
Housen Li, Johannes Schwab, Stephan Antholzer, and Markus Haltmeier, NETT: Solving inverse problems with deep neural networks , Inverse Problems 36
2020
Later among the works it cites.
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington, Wide neural networks of any depth evolve as linear models under gradient descent , Journal of Statistical Mechanics: Theory and Experiment 2020
2020
Later among the works it cites.
Carlo Marcati, Joost Opschoor, Philipp Petersen, and Christoph Schwab, Exponential ReLU neural network approximation rates for point and edge singularities , 2020, ETH Zurich SAM Research Report
2020
Later among the works it cites.
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai, Harmless interpolation of noisy data in regression , IEEE Journal on Selected Areas in Information Theory 1
2020
Later among the works it cites.
Ryumei Nakada and Masaaki Imaizumi, Adaptive approximation and generalization of deep neural network with intrinsic dimensionality , Journal of Machine Learning Research 21
2020
Later among the works it cites.
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever, Deep double descent: Where bigger models and more data hurt , International Conference on Learning Representations, 2020
2020
Later among the works it cites.
Joost Opschoor, Philipp Petersen, and Christoph Schwab, Deep ReLU networks and high-order finite element methods , Analysis and Applications (2020), no. 0, 1–56
2020
Later among the works it cites.
Philipp Petersen, Mones Raslan, and Felix Voigtlaender, Topological properties of the set of functions generated by neural networks of fixed size , Foundations of Computational Mathematics (2020), 1–70
2020
Later among the works it cites.
David Pfau, James S Spencer, Alexander GDG Matthews, and W Matthew C Foulkes, Ab initio solution of the many-electron schrödinger equation with deep neural networks , Physical Review Research 2
2020
Later among the works it cites.
Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi, Ali Farhadi, and Mohammad Rastegari, What’s hidden in a randomly weighted neural network? , Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 11893–11902
2020
Later among the works it cites.
Andrew W Senior, Richard Evans, John Jumper, James Kirkpatrick, Laurent Sifre, Tim Green, Chongli Qin, Augustin Žídek, Alexander WR Nelson, and Alex Bridgland, Improved protein structure prediction using potentials from deep learning , Nature 577
2020
Later among the works it cites.
Zuowei Shen, Deep network approximation characterized by number of neurons , Communications in Computational Physics 28
2020
Later among the works it cites.
Dmitry Yarotsky and Anton Zhevnerchuk, The phase diagram of approximation rates for deep neural networks , Advances in Neural Information Processing Systems, vol. 33, 2020
2020
Later among the works it cites.
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Michael C Mozer, and Yoram Singer, Identity crisis: Memorization and generalization under extreme overparameterization , International Conference on Learning Representations, 2020
2020
Later among the works it cites.
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu, Gradient descent optimizes over-parameterized deep ReLU networks , Machine Learning 109
2020
Later among the works it cites.
Ding-Xuan Zhou, Theory of deep convolutional neural networks: Downsampling , Neural Networks 124
2020
Later among the works it cites.
Christian Beck, Sebastian Becker, Philipp Grohs, Nor Jaafari, and Arnulf Jentzen, Solving the kolmogorov pde by means of deep learning , Journal of Scientific Computing 88
2021
Closest in time.
2021
Closest in time.
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Linearized two-layers neural networks in high dimension , The Annals of Statistics 49
2021
Closest in time.
2021
Closest in time.
Licong Lin and Edgar Dobriban, What causes the test error? Going beyond bias-variance via ANOVA , Journal of Machine Learning Research 22
2021
Closest in time.
Weilin Li, Generalization error of minimum weighted norm and kernel interpolation , SIAM Journal on Mathematics of Data Science 3
2021
Closest in time.
Fabian Laakmann and Philipp Petersen, Efficient approximation of solutions of parametric linear transport equations by ReLU DNNs , Advances in Computational Mathematics 47
2021
Closest in time.
Vishal Monga, Yuelong Li, and Yonina C Eldar, Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing , IEEE Signal Processing Magazine 38
2021
Closest in time.
Tenavi Nakamura-Zimmerer, Qi Gong, and Wei Kang, Adaptive deep learning for high-dimensional Hamilton–Jacobi–Bellman Equations , SIAM Journal on Scientific Computing 43
2021
Closest in time.
2021
Closest in time.
Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and S Yu Philip, A comprehensive survey on graph neural networks , IEEE Transactions on Neural Networks and Learning Systems 32
2021
Closest in time.