Fetching the paper…
Reading the bibliography…
It is almost always easier to find an accurate-but-complex model than an accurate-yet-simple model.
An information measure for classification
Christopher S Wallace and David M Boulton. 1968 · 1968
Earlier work this paper cites.
On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities
VN Vapnik and A Ya Chervonenkis. 1971 · 1971
Earlier work this paper cites.
Approximation of monomials by lower degree polynomials
DJ Newman and TJ Rivlin. 1976 · 1976
Earlier work this paper cites.
A finite sample distribution-free performance bound for local discrimination rules
William H Rogers and Terry J Wagner. 1978 · 1978
Earlier work this paper cites.
Learning from noisy examples
Dana Angluin and Philip Laird. 1988 · 1988
Earlier work this paper cites.
On the degree of polynomials that approximate symmetric boolean functions (preliminary version). In Proceedings of the Twenty-Fourth Annual ACM Symposium on Theory of Computing (Victoria, British Columbia, Canada) (STOC ’92) . Association for Computing Machinery, New York, NY, USA, 468–474
Ramamohan Paturi. 1992 · 1992
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik. 1995 · 1995
Earlier work this paper cites.
The Nature of Statistical Learning Theory
Vladimir N Vapnik. 1995 · 1995
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
A tutorial on support vector machines for pattern recognition
Christopher JC Burges. 1998 · 1998
Earlier work this paper cites.
Nonlinear approximation
Ronald A DeVore. 1998 · 1998
Earlier work this paper cites.
Boosting the margin: A new explanation for the effectiveness of voting methods
Robert E Schapire, Yoav Freund, Peter Bartlett, Wee Sun Lee, et al · 1998
Earlier work this paper cites.
Algorithmic stability and sanity-check bounds for leave-one-out cross-validation
Michael Kearns and Dana Ron. 1999 · 1999
Earlier work this paper cites.
Statistical modeling: The two cultures (with comments and a rejoinder by the author)
Leo Breiman et al · 2001
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson. 2002 · 2002
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff. 2002 · 2002
Earlier work this paper cites.
Empirical margin distributions and bounding the generalization error of combined classifiers
Vladimir Koltchinskii and Dmitry Panchenko. 2002 · 2002
Earlier work this paper cites.
PAC-Bayes & margins. In Proceedings of the 15th International Conference on Neural Information Processing Systems (NIPS’02) . MIT Press, Cambridge, MA, USA, 439–446
John Langford and John Shawe-Taylor. 2002 · 2002
Earlier work this paper cites.
The covering number in learning theory
Ding-Xuan Zhou. 2002 · 2002
Cited alongside, same era.
A few notes on statistical learning theory. In Advanced Lectures on Machine Learning . Springer, 40 pages
Shahar Mendelson. 2003 · 2003
Cited alongside, same era.
Complexity regularization via localized random penalties
Gábor Lugosi and Marten Wegkamp. 2004 · 2004
Cited alongside, same era.
Local Rademacher complexities
Peter L Bartlett, Olivier Bousquet, Shahar Mendelson, et al · 2005
Cited alongside, same era.
The estimate for approximation error of neural networks: A constructive approach
Feilong Cao, Tingfan Xie, and Zongben Xu. 2008 · 2008
Cited alongside, same era.
On the complexity of linear prediction: risk bounds, margin bounds, and regularization. In Proceedings of the 21st International Conference on Neural Information Processing Systems (Vancouver, British Columbia, Canada) (NIPS’08) . Curran Associates Inc., Red Hook, NY, USA, 793–800
Sharp minima can generalize for deep nets. In Proceedings of the 34th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 70) . 1019–1028
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. 2017 · 2017
Later among the works it cites.
Identifying a minimal class of models for high-dimensional data
Daniel Nevo and Ya’acov Ritov. 2017 · 2017
Later among the works it cites.
Interpretable classification models for recidivism prediction
Jiaming Zeng, Berk Ustun, and Cynthia Rudin. 2017 · 2017
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. 2019 · 2019
Closest in time.
Entropy-SGD: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina. 2019 · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sham M. Kakade, Karthik Sridharan, and Ambuj Tewari. 2008 · 2008
Cited alongside, same era.
Stability selection
Nicolai Meinshausen and Peter Bühlmann. 2010 · 2010
Cited alongside, same era.
Smoothness, Low Noise and Fast Rates. In Advances in Neural Information Processing Systems , Vol. 23. Curran Associates, Inc., 2199–2207
Nathan Srebro, Karthik Sridharan, and Ambuj Tewari. 2010 · 2010
Cited alongside, same era.
Algorithms and error bounds for multivariate piecewise constant approximation
Oleg Davydov. 2011 · 2011
Cited alongside, same era.
Interplay between concentration, complexity and geometry in learning theory with applications to high dimensional data analysis
Guillaume Lecué. 2011 · 2011
Cited alongside, same era.
Noise tolerance under risk minimization
Naresh Manwani and PS Sastry. 2013 · 2013
Cited alongside, same era.
Learning with noisy labels. In Advances in Neural Information Processing Systems , C. J. C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Q. Weinberger (Eds.), Vol. 26. Curran Associates, Inc
Nagarajan Natarajan, Inderjit S Dhillon, Pradeep K Ravikumar, and Ambuj Tewari. 2013 · 2013
Cited alongside, same era.
UCI Machine Learning Repository
Dheeru Dua and Casey Graff. 2019 · 2019
Closest in time.
All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simultaneously
Aaron Fisher, Cynthia Rudin, and Francesca Dominici. 2019 · 2019
Closest in time.
Detecting underspecification with local ensembles
David Madras, James Atwood, and Alex D’Amour. 2019 · 2019
Closest in time.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. 2019 · 2019
Closest in time.
Underspecification presents challenges for credibility in modern machine learning
Alexander D’Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D Hoffman, et al · 2020
Closest in time.
Exploring the cloud of variable importance for the set of all good models
Jiayun Dong and Cynthia Rudin. 2020 · 2020
Closest in time.
Predictive Multiplicity in Classification. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119) . PMLR, 6765–6774
Charles Marx, Flavio Calmon, and Berk Ustun. 2020 · 2020
Closest in time.
The Age of Secrecy and Unfairness in Recidivism Prediction
Cynthia Rudin, Caroline Wang, and Beau Coker. 2020 · 2020
Closest in time.
A theory of statistical inference for ensuring the robustness of scientific results
Beau Coker, Cynthia Rudin, and Gary King. 2021 · 2021
Closest in time.
Characterizing fairness over the set of good models under selective labels. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research)
Amanda Coston, Ashesh Rambachan, and Alexandra Chouldechova. 2021 · 2021
Closest in time.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. 2021 · 2021
Closest in time.
Machine learning with operational costs
Theja Tulabandhula and Cynthia Rudin. 2013 · 2028
Closest in time.