Fetching the paper…
Reading the bibliography…
Breakthroughs in machine learning are rapidly changing science and society, yet our fundamental understanding of this technology has lagged far behind.
Occam’s razor
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K Warmuth · 1987
Earlier work this paper cites.
A training algorithm for optimal margin classifiers
Bernhard E Boser, Isabelle M Guyon, and Vladimir N Vapnik · 1992
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
Stuart Geman, Elie Bienenstock, and René Doursat · 1992
Earlier work this paper cites.
Darpa timit acoustic-phonetic continous speech corpus cd-rom
John S Garofolo, Lori F Lamel, William M Fisher, Jonathon G Fiscus, and David S Pallett · 1993
Earlier work this paper cites.
Newsweeder: Learning to filter netnews
Ken Lang · 1995
Earlier work this paper cites.
The Nature of Statistical Learning Theory
Vladimir N. Vapnik · 1995
Earlier work this paper cites.
Dynamics of training
Siegfried Bös and Manfred Opper · 1997
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L. Bartlett · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Boosting the margin: a new explanation for the effectiveness of voting methods
Robert E. Schapire, Yoav Freund, Peter Bartlett, and Wee Sun Lee · 1998
Earlier work this paper cites.
Random forests
Leo Breiman · 2001
Earlier work this paper cites.
Pert-perfect random tree ensembles
Adele Cutler and Guohua Zhao · 2001
Earlier work this paper cites.
Greedy function approximation: a gradient boosting machine
Jerome H Friedman · 2001
Earlier work this paper cites.
The Elements of Statistical Learning , volume 1
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2001
Cited alongside, same era.
Boosting with the l 2 l_{2} loss: regression and classification
Peter Bühlmann and Bin Yu · 2003
Cited alongside, same era.
Gaussian processes in machine learning
Carl Edward Rasmussen · 2004
Cited alongside, same era.
Scattered Data Approximation
Holger Wendland · 2004
Cited alongside, same era.
All of Nonparametric Statistics
Larry Wasserman · 2006
Cited alongside, same era.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Cited alongside, same era.
Kernel methods for deep learning
Youngmin Cho and Lawrence K. Saul · 2009
An analysis of deep neural network models for practical applications
Alfredo Canziani, Adam Paszke, and Eugenio Culurciello · 2016
Later among the works it cites.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Later among the works it cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Later among the works it cites.
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco · 2017
Later among the works it cites.
Deep learning tutorial at the Simons Institute, Berkeley, https://simons.berkeley.edu/talks/ruslan-salakhutdinov-01-26-2017-1, 2017
Ruslan Salakhutdinov · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Homo heuristicus: Why biased minds make better inferences
Gerd Gigerenzer and Henry Brighton · 2009
Cited alongside, same era.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Cited alongside, same era.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Ng · 2011
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Cited alongside, same era.
Explaining the success of adaboost and random forests as interpolating classifiers
Abraham J Wyner, Matthew Olson, Justin Bleich, and David Mease · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Closest in time.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Closest in time.
A modern take on the bias-variance tradeoff in neural networks
Brady Neal, Sarthak Mittal, Aristide Baratin, Vinayak Tantia, Matthew Scicluna, Simon Lacoste-Julien, and Ioannis Mitliagkas · 2018
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2018
Closest in time.
A jamming transition from under-to over-parametrization affects loss landscape and generalization
Stefano Spigler, Mario Geiger, Stéphane d’Ascoli, Levent Sagun, Giulio Biroli, and Matthieu Wyart · 2018
Closest in time.