Fetching the paper…
Reading the bibliography…
In the past decade the mathematical theory of machine learning has lagged far behind the triumphs of deep neural networks on practical challenges.
Angenaherte auflosung von systemen linearer gleichungen
Stefan Kaczmarz · 1937
Earlier work this paper cites.
A topological property of real analytic subsets
Stanislaw Lojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for minimizing functionals
Boris Teodorovich Polyak · 1963
Earlier work this paper cites.
On estimating regression
Elizbar A Nadaraya · 1964
Earlier work this paper cites.
Smooth regression analysis
Geoffrey S Watson · 1964
Earlier work this paper cites.
Nearest neighbor pattern classification
Thomas Cover and Peter Hart · 1967
Earlier work this paper cites.
A two-dimensional interpolation function for irregularly-spaced data
Donald Shepard · 1968
Earlier work this paper cites.
A correspondence between bayesian estimation on stochastic processes and smoothing by splines
George S Kimeldorf and Grace Wahba · 1970
Earlier work this paper cites.
Gauss and the Invention of Least Squares
Stephen M. Stigler · 1981
Earlier work this paper cites.
Occam’s razor
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K Warmuth · 1987
Earlier work this paper cites.
Simplicial multivariable linear interpolation
John H Halton · 1991
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
Stuart Geman, Elie Bienenstock, and René Doursat · 1992
Earlier work this paper cites.
Reflections after refereeing papers for nips
Leo Breiman · 1995
Earlier work this paper cites.
The Nature of Statistical Learning Theory
Vladimir N. Vapnik · 1995
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
The hilbert kernel regression estimate
Luc Devroye, Laszlo Györfi, and Adam Krzyżak · 1998
Earlier work this paper cites.
Boosting the margin: a new explanation for the effectiveness of voting methods
Robert E. Schapire, Yoav Freund, Peter Bartlett, and Wee Sun Lee · 1998
Earlier work this paper cites.
Pert-perfect random tree ensembles
Adele Cutler and Guohua Zhao · 2001
Earlier work this paper cites.
The Elements of Statistical Learning
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
A Distribution-Free Theory of Nonparametric Regression
László Györfi, Michael Kohler, Adam Krzyzak, and Harro Walk · 2002
Earlier work this paper cites.
Everything old is new again: a fresh look at historical approaches in machine learning
Ryan Michael Rifkin · 2002
Earlier work this paper cites.
Introduction to statistical learning theory
Olivier Bousquet, Stéphane Boucheron, and Gábor Lugosi · 2003
Earlier work this paper cites.
Scattered Data Approximation
Holger Wendland · 2004
Earlier work this paper cites.
Beyond the point cloud: from transductive to semi-supervised learning
Vikas Sindhwani, Partha Niyogi, and Mikhail Belkin · 2005
Earlier work this paper cites.
Leaving the span
Manfred K Warmuth and SVN Vishwanathan · 2005
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Comment: Boosting algorithms: Regularization, prediction and model fitting
Andreas Buja, David Mease, Abraham J Wyner, et al · 2007
Earlier work this paper cites.
On early stopping in gradient descent learning
Yuan Yao, Lorenzo Rosasco, and Andrea Caponnetto · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
A randomized kaczmarz algorithm with exponential convergence
Thomas Strohmer and Roman Vershynin · 2009
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R Bach · 2011
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas L Roux, Mark Schmidt, and Francis R Bach · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Deanna Needell, Rachel Ward, and Nati Srebro · 2014
Cited alongside, same era.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2019
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
A jamming transition from under- to over-parametrization affects generalization in deep learning
S Spigler, M Geiger, S d’Ascoli, L Sagun, G Biroli, and M Wyart · 2019
Later among the works it cites.
One pixel attack for fooling deep neural networks
Jiawei Su, Danilo Vasconcellos Vargas, and Kouichi Sakurai · 2019
Later among the works it cites.
On the number of variables to use in principal component regression
Ji Xu and Daniel Hsu · 2019
Later among the works it cites.
Backward feature correction: How deep learning performs deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An analysis of deep neural network models for practical applications
Alfredo Canziani, Adam Paszke, and Eugenio Culurciello · 2016
Cited alongside, same era.
Robustness of classifiers: from adversarial to random noise
Alhussein Fawzi, Seyed-Mohsen Moosavi-Dezfooli, and Pascal Frossard · 2016
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-lojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
Reflections on random kitchen sinks
Ali Rahimi and Ben. Recht · 2017
Cited alongside, same era.
Tutorial on deep learning
Ruslan Salakhutdinov · 2017
Cited alongside, same era.
Zeyuan Allen-Zhu and Yuanzhi Li · 2020
Later among the works it cites.
Harnessing the power of infinitely wide deep nets on small-data tasks
Sanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov, Ruosong Wang, and Dingli Yu · 2020
Later among the works it cites.
Failures of model-dependent generalization bounds for least-norm interpolation
Peter L Bartlett and Philip M Long · 2020
Later among the works it cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Later among the works it cites.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2020
Later among the works it cites.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Later among the works it cites.
Finite versus infinite neural networks: an empirical study
Jaehoon Lee, Samuel S Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2020
Later among the works it cites.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
Mingchen Li, Mahdi Soltanolkotabi, and Samet Oymak · 2020
Later among the works it cites.
Just interpolate: Kernel ridgeless regression can generalize
Tengyuan Liang, Alexander Rakhlin, et al · 2020
Later among the works it cites.
Accelerating sgd with momentum for over-parameterized learning
Chaoyue Liu and Mikhail Belkin · 2020
Later among the works it cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2020
Later among the works it cites.
On the linearity of large non-linear models: when and why the tangent kernel is constant
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2020
Later among the works it cites.
A brief prehistory of double descent
Marco Loog, Tom Viering, Alexander Mey, Jesse H Krijthe, and David MJ Tax · 2020
Later among the works it cites.
Kernel methods through the roof: handling billions of points efficiently
Giacomo Meanti, Luigi Carratino, Lorenzo Rosasco, and Alessandro Rudi · 2020
Later among the works it cites.
Classification vs regression in overparameterized regimes: Does the loss function matter?, 2020
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu, and Anant Sahai · 2020
Later among the works it cites.
Harmless interpolation of noisy data in regression
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai · 2020
Later among the works it cites.
In defense of uniform convergence: Generalization via derandomization with an application to interpolating predictors
Jeffrey Negrea, Gintare Karolina Dziugaite, and Daniel Roy · 2020
Later among the works it cites.
Do deeper convolutional networks perform better?
Eshaan Nichani, Adityanarayanan Radhakrishnan, and Caroline Uhler · 2020
Later among the works it cites.
On the expressive power of kernel methods and the efficiency of kernel learning by association schemes
Kothari K Pravesh and Livni Roi · 2020
Later among the works it cites.
Overparameterized neural networks implement associative memory
Adityanarayanan Radhakrishnan, Mikhail Belkin, and Caroline Uhler · 2020
Later among the works it cites.
Improved protein structure prediction using potentials from deep learning
Andrew Senior, Richard Evans, John Jumper, James Kirkpatrick, Laurent Sifre, Tim Green, Chongli Qin, Augustin Zidek, Alexander WR Nelson, Alex Bridgland, et al · 2020
Later among the works it cites.
Neural kernels without tangents
Vaishaal Shankar, Alex Fang, Wenshuo Guo, Sara Fridovich-Keil, Jonathan Ragan-Kelley, Ludwig Schmidt, and Benjamin Recht · 2020
Later among the works it cites.
Theoretical insights into multiclass classification: A high-dimensional asymptotic view
Christos Thrampoulidis, Samet Oymak, and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
On uniform convergence and low-norm interpolation learning
Lijia Zhou, Danica J Sutherland, and Nati Srebro · 2020
Later among the works it cites.
Deep learning: a statistical viewpoint, 2021
Peter L. Bartlett, Andrea Montanari, and Alexander Rakhlin · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2021
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Like Hui and Mikhail Belkin · 2021
Closest in time.