Fetching the paper…
Reading the bibliography…
Probability forecasts for binary outcomes, often referred to as probabilistic classifiers or confidence scores, are ubiquitous in science and society, and methods for evaluating and comparing them are in great demand.
Verification of forecasts expressed in terms of probability
Brier, G. W. (1950) · 1950
Earlier work this paper cites.
An empirical distribution function for sampling with incomplete information
Ayer, M., Brunk, H. D., Ewing, G. M., Reid, W. T., and Silvermann, E. (1955) · 1955
Earlier work this paper cites.
A bias-corrected decomposition of the Brier score
Ferro, C. A. T. and Fricker, T. E. (2012) · 1960
Earlier work this paper cites.
Elicitation of personal probabilities and expectations
Savage, L. J. (1971) · 1971
Earlier work this paper cites.
A new vector partition of the probability score
Murphy, A. H. (1973) · 1973
Earlier work this paper cites.
The relative operating characteristic in psychology
Swets, J. A. (1973) · 1973
Earlier work this paper cites.
Signal Dectection Theory and ROC Analysis
Egan, J. P. (1975) · 1975
Earlier work this paper cites.
Reliability of subjective probability forecasts of precipitation and temperature
Murphy, A. H. and Winkler, R. L. (1977) · 1977
Earlier work this paper cites.
The meaning and use of the area under a receiver operating characteristic (ROC) curve
Hanley, A. and McNeil, J. (1982) · 1982
Earlier work this paper cites.
Comparing the areas under two or more correlated receiver operating characteristic curves: A nonparametric approach
DeLong, E. R., DeLong, D. M., and Clarke-Pearson, D. L. (1988) · 1988
Earlier work this paper cites.
A general method for comparing probability assessors
Schervish, M. J. (1989) · 1989
Earlier work this paper cites.
Fursion of detection probabilities and comparison of multisensor systems
Krzysztofowicz, R. and Long, D. (1990) · 1990
Earlier work this paper cites.
Diagnostic verification of probability forecasts
Murphy, A. H. and Winkler, R. L. (1992) · 1992
Earlier work this paper cites.
The use of the area under the ROC curve in the evaluation of machine learning algorithms
Bradley, A. P. (1997) · 1997
Earlier work this paper cites.
Axiomatic characterization of the quadratic scoring rule
Selten, R. (1998) · 1998
Earlier work this paper cites.
Fragile families: Sample and design
Reichman, N. E., Teitler, J. O., Garfinkel, I., and McLanahan, S. S. (2001) · 2001
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Zadrozny, B. and Elkan, C. (2002) · 2002
Earlier work this paper cites.
The ROC curve and the area under it as performance measures
Marzban, C. (2004) · 2004
Earlier work this paper cites.
Loss functions for binary class probability estimation and classification: Structure and application
Buja, A., Stuetzle, W., and Shen, Y. (2005) · 2005
Earlier work this paper cites.
ROCR: Visualizing classifier performance in R
Sing, T., Sander, O., Beerenwinkel, N., and Lengauer, T. (2005) · 2005
Earlier work this paper cites.
Cost curves: An improved method for visualizing classifier performance
Drummond, C. and Holte, R. (2006) · 2006
Earlier work this paper cites.
An introduction to ROC analysis
Fawcett, T. (2006) · 2006
Earlier work this paper cites.
Increasing the reliability of reliability diagrams
Bröcker, J. and Smith, L. A. (2007) · 2007
Earlier work this paper cites.
PAV and the ROC convex hull
Fawcett, T. and Niculescu-Mizil, A. (2007) · 2007
Earlier work this paper cites.
Probabilistic forecasts, calibration and sharpness
Gneiting, T., Balabdaoui, F., and Raftery, A. E. (2007) · 2007
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Gneiting, T. and Raftery, A. E. (2007) · 2007
Cited alongside, same era.
Rejoinder on: Assessing probabilistic forecasts of multivariate quantities, with an application to ensemble predictions of surface winds
Gneiting, T., Stanberry, L. A., Grimit, E. P., Held, L., and Johnson, N. A. (2008) · 2008
Cited alongside, same era.
Isotone optimization in R
De Leeuw, J., Hornik, K., and Mair, P. (2009) · 2009
Cited alongside, same era.
Measuring classifier performance: A coherent alternative to the area under the ROC curve
Hand, D. J. (2009) · 2009
Cited alongside, same era.
On the convexity of ROC curves estimated from radiological test results
Pesce, L. L., Metz, C. E., and Berbaum, K. S. (2010) · 2010
Cited alongside, same era.
Brier curves: A new cost-based visualisation of classifier performance
A comparison of flare forecasting methods. II. Benchmarks, metrics, and performance results for operational solar flare forecasting systems
Leka, K. D., Park, S.-H., Kusano, K., Andries, J., Barnes, G., Bingham, S., Bloomfield, D. S., McCloskey, A. E., Delouille, V., Falconer, D., Gallagher, P. T., Georgoulis, M. K., Kubo, Y., Lee, K., Lee, S., Lobzin, V., Mun, J., Murray, S. A., Nageem, T. A. M. H., Qahwaji, R., Sharpe, M., Steenburgh, R. A., Steward, G., and Terkildsen, M. (2019) · 2019
Later among the works it cites.
Statistical Methods in the Atmospheric Sciences
Wilks, D. S. (2019) · 2019
Later among the works it cites.
classifierplots: Generates a visualization of classifier performance as a grid of diagnostic plots
Defazio, A. and Campbell, H. (2020) · 2020
Later among the works it cites.
ROC curves for clinical prediction models part 4. Selection of the risk threshold — once chosen, always the same?
Janssens, A. C. J. W. (2020) · 2020
Later among the works it cites.
From research to applications – examples of operational ensemble post-processing in France using machine learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hernández-Orallo, J., Flach, P., and Ferri, C. (2011) · 2011
Cited alongside, same era.
A survey on graphical methods for classification predictive performance evaluation
Prati, R. C., Batista, G. E. A. P. A., and Monard, M. C. (2011) · 2011
Cited alongside, same era.
pROC: An open-source package for R
Robin, X., Turck, N., Hainard, A., Tiberti, N., Lisacek, F., Sanchez, J. C., and Müller, M. (2011) · 2011
Cited alongside, same era.
Estimating reliability and resolution of probability forecasts through decomposition of the empirical score
Bröcker, J. (2012) · 2012
Cited alongside, same era.
A unified view of performance metrics: Translating threshold choice into expected classification loss
Hernández-Orallo, J., Flach, P., and Ferri, C. (2012) · 2012
Cited alongside, same era.
Economic value and skill
Richardson, D. S. (2012) · 2012
Cited alongside, same era.
Combining predictive distributions
Gneiting, T. and Ranjan, R. (2013) · 2013
Cited alongside, same era.
Taillardat, M. and Mestre, O. (2020) · 2020
Later among the works it cites.
Stable reliability diagrams for probabilistic classifiers
Dimitriadis, T., Gneiting, T., and Jordan, A. I. (2021) · 2021
Later among the works it cites.
Classifier calibration: How to assess and improve predicted class probabilities: A survey
Filho, T. S., Song, H., Perelló-Nieto, M., Santos-Rodríguez, R., Kull, M., and Flach, P. (2021) · 2021
Later among the works it cites.
Gneiting, T. and Resin, J. (2021) · 2021
Later among the works it cites.
A low-cost post-processing technique improves weather forecasts around the world
Hewson, T. D. and Pillosu, F. M. (2021) · 2021
Later among the works it cites.
Generic conditions for forecast dominance
Krüger, F. and Ziegel, J. F. (2021) · 2021
Later among the works it cites.
Correction for Salganik et al., Measuring the predictability of life outcomes with a scientific mass collaboration
Salganik, M. J., Lundberg, I., Kindel, A. T., Ahearn, C. E., Al-Ghoneim, K., Almaatouq, A., Altschul, D. M., Brand, J. E., Carnegie, N. B., Compton, R. J., Datta, D., Davidson, T., Filippova, A., Gilroy, C., Goode, B. J., Jahani, E., Kashyap, R., Kirchner, A., McKay, S., Morgan, A. C., Pentland, A., Polimis, K., Raes, L., Rigobon, D. E., Roberts, C. V., Stanescu, D. M., Suhara, Y., Usmani, A., Wang, E. H., Adem, M., Alhajri, A., AlShebli, B., Amin, R., Amos, R. B., Argyle, L. P., Baer-Bositis, L., Büchi, M., Chung, B.-R., Eggert, W., Faletto, G., Fan, Z., Freese, J., Gadgil, T., Gagné, J., Gao, Y., Halpern-Manners, A., Hashim, S. P., Hausen, S., He, G., Higuera, K., Hogan, B., Horwitz, I. M., Hummel, L. M., Jain, N., Jin, K., Jurgens, D., Kaminski, P., Karapetyan, A., Kim, E. H., Leizman, B., Liu, N., Möser, M., Mack, A. E., Mahajan, M., Mandell, N., Marahrens, H., Mercado-Garcia, D., Mocz, V., Mueller-Gastell, K., Musse, A., Niu, Q., Nowak, W., Omidvar, H., Or, A., Ouyang, K., Pinto, K. M., Porter, E., Porter, K. E., Qian, C., Rauf, T., Sargsyan, A., Schaffner, T., Schnabel, L., Schonfeld, B., Sender, B., Tang, J. D., Tsurkov, E., van Loon, A., Varol, O., Wang, X., Wang, Z., Wang, J., Wang, F., Weissman, S., Whitaker, K., Wolters, M. K., Woon, W. L., Wu, J., Wu, C., Yang, K., Yin, J., Zhao, B., Zhu, C., Brooks-Gunn, J., Engelhardt, B. E., Hardt, M., Knox, D., Levy, K., Narayanan, A., Stewart, B. M., Watts, D. J., and McLanahan, S. (2021) · 2021
Later among the works it cites.
Metrics of calibration for probabilistic predictions
Arrieta-Ibarra, I., Gujral, P., Tannen, J., Tygert, M., and Xu, C. (2022) · 2022
Later among the works it cites.
Uniform calibration tests for forecasting systems with small lead time
Bröcker, J. (2022) · 2022
Later among the works it cites.
Honest calibration assessment for binary outcome predictions
Dimitriadis, T., Dümbgen, L., Henzi, A., Puke, M., and Ziegel, J. (2022) · 2022
Later among the works it cites.
Receiver operating characteristic (ROC) curves: Equivalences, beta model, and minimum distance estimation
Gneiting, T. and Vogel, P. (2022) · 2022
Later among the works it cites.
Receiver operating characteristic (ROC) movies, universal ROC (UROC) curves, and coefficient of predictive ability (CPA)
Gneiting, T. and Walz, E.-M. (2022) · 2022
Later among the works it cites.
Model diagnostics and forecast evaluation for quantiles
Gneiting, T., Wolffram, D., Resin, J., Kraus, K., Bracher, J., Dimitriadis, T., Hagenmeyer, V., Jordan, A. I., Lerch, S., Phipps, K., and Schienle, M. (2022) · 2022
Later among the works it cites.
Notes on the H-measure of classifier performance
Hand, D. J. and Anagnostopoulos, C. (2022) · 2022
Later among the works it cites.
Characterizing the optimal solutions to the isotonic regression problem for identifiable functionals
Jordan, A. I., Mühlemann, A., and Ziegel, J. F. (2022) · 2022
Later among the works it cites.
Introduction to the M5 forecasting competition Special Issue
Makridakis, S., Petropoulos, F., and Spiliotis, E. (2022) · 2022
Later among the works it cites.
R : A Language and Environment for Statistical Computing
R · 2022
Later among the works it cites.
Mitigating bias in calibration error estimation
Roelofs, R., Cain, N., Shlens, J., and Mozer, M. C. (2022) · 2022
Later among the works it cites.
Replication material for “Evaluating probabilistic classifiers: The triptych”
Dimitriadis, T. and Jordan, A. I. (2023) · 2023
Closest in time.
Calibrate: Interactive analysis of probabilistic model output
Xenopoulos, P., Rulff, J., Nonato, L. G., Barr, B., and Silva, C. (2023) · 2023
Closest in time.