In Introduction to Data Mining
Tan, P.-N., Steinbach, M., and Kumar, V. (2005) · 2005
Later among the works it cites.
Bias in error estimation when using cross-validation for model selection
Varma, S. and Simon, R. (2006) · 2006
Later among the works it cites.
On comparison of feature selection algorithms
Refaeilzadeh, P., Tang, L., and Liu, H. (2007) · 2007
Later among the works it cites.
In The Elements of Statistical Learning: Data Mining, Inference, and Prediction
Hastie, T., Tibshirani, R., and Friedman, J. H. (2009) · 2009
Later among the works it cites.
Estimating classification error rate: Repeated cross-validation, repeated hold-out and bootstrap
Kim, J.-H. (2009) · 2009
Later among the works it cites.
Multiple McNemar tests
Westfall, P. H., Troendle, J. F., and Pennello, G. (2010) · 2010
Later among the works it cites.
Scikit-learn: Machine learning in python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al. (2011) · 2011
Later among the works it cites.
Statistical Methods for Rates and Proportions
Fleiss, J. L., Levin, B., and Paik, M. C. (2013) · 2013
Later among the works it cites.
In An Introduction to Statistical Learning: With Applications in R
James, G., Witten, D., Hastie, T., and Tibshirani, R. (2013) · 2013
Later among the works it cites.
Cross-validation failure: small sample sizes lead to large error bars
Varoquaux, G. (2017) · 2017
Later among the works it cites.
Mlxtend: Providing machine learning and data science utilities and extensions to python’s scientific computing stack
Raschka, S. (2018) · 2018
Closest in time.