“Structural risk minimization over data-dependent hierarchies”
John Shawe-Taylor, Peter Bartlett, Robert Williamson and Martin Anthony · 1940
Earlier work this paper cites.
“Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition”
Thomas Cover · 1965
Earlier work this paper cites.
“Estimation of dependences based on empirical data”
Vladimir Vapnik · 1982
Earlier work this paper cites.
“Learning from noisy examples”
Dana Angluin and Philip Laird · 1988
Earlier work this paper cites.
“Learning with an unreliable teacher”
Gabor Lugosi · 1992
Earlier work this paper cites.
“Toward efficient agnostic learning”
Michael Kearns, Robert Schapire and Linda Sellie · 1994
Earlier work this paper cites.
“Chernoff–Hoeffding bounds for applications with limited independence”
Jeanette Schmidt, Alan Siegel and Aravind Srinivasan · 1995
Earlier work this paper cites.
“The nature of statistical learning theory”
Vladimir Vapnik · 1995
Earlier work this paper cites.
“Balls and Bins: A Study in Negative Dependence”
Devdatt Dubhashi and Desh Ranjan · 1998
Earlier work this paper cites.
“Efficient noise-tolerant learning from statistical queries”
Michael Kearns · 1998
Earlier work this paper cites.
“Sample-efficient strategies for learning in the presence of noise”
Nicolo Cesa-Bianchi, Eli Dichterman, Paul Fischer, Eli Shamir and Hans Simon · 1999
Earlier work this paper cites.
“On PAC learning using Winnow, Perceptron, and a Perceptron-like algorithm”
Rocco Servedio · 1999
Earlier work this paper cites.
“Equitable coloring extends Chernoff-Hoeffding bounds”
Sriram Pemmaraju · 2001
Earlier work this paper cites.
“Lectures on the coupling method”
Torgny Lindvall · 2002
Earlier work this paper cites.
“On discriminative vs. generative classifiers: A comparison of logistic regression and naive Bayes”
Andrew Ng and Michael Jordan · 2002
Earlier work this paper cites.
“An introduction to multivariate statistical analysis”, Wiley Series in Probability and Statistics
T.W. Anderson · 2003
Earlier work this paper cites.
“Some theory for Fisher’s linear discriminant function, ‘naive Bayes’, and some alternatives when there are many more variables than observations”
Peter Bickel and Elizaveta Levina · 2004
Earlier work this paper cites.
“Understanding deep learning requires rethinking generalization”
Original
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht and Oriol Vinyals · 2004
Earlier work this paper cites.
“Boosting in the presence of noise”
Adam Kalai and Rocco Servedio · 2005
Earlier work this paper cites.
“Risk bounds for statistical learning”
Pascal Massart and Élodie Nédélec · 2006
Earlier work this paper cites.
“Statistical performance of support vector machines”
Gilles Blanchard, Olivier Bousquet and Pascal Massart · 2008
Earlier work this paper cites.
“Higher criticism thresholding: Optimal feature selection when useful features are rare and weak”
David Donoho and Jiashun Jin · 2008
Earlier work this paper cites.
“High dimensional classification using features annealed independence rules”
Jianqing Fan and Yingying Fan · 2008
Earlier work this paper cites.
“Agnostically learning halfspaces”
Adam Kalai, Adam Klivans, Yishay Mansour and Rocco Servedio · 2008
Earlier work this paper cites.
“Neural network learning: Theoretical foundations”
Martin Anthony and Peter Bartlett · 2009
Earlier work this paper cites.
“Introduction to algorithms”
Thomas Cormen, Charles Leiserson, Ronald Rivest and Clifford Stein · 2009
Earlier work this paper cites.
“The elements of statistical learning: data mining, inference, and prediction”
Trevor Hastie, Robert Tibshirani and Jerome Friedman · 2009
Earlier work this paper cites.