Fetching the paper…
Reading the bibliography…
A variety of different performance metrics are commonly used in the machine learning literature for the evaluation of classification systems.
Optimal Statistical Decisions
Morris H. DeGroot, · 1970
Earlier work this paper cites.
“Comparison of the predicted and observed secondary structure of t4 phage lysozyme,”
B. W. Matthews, · 1975
Earlier work this paper cites.
“Therapeutic decision making: A cost-benefit analysis,”
Stephen G. Pauker and Jerome P. Kassirer, · 1975
Earlier work this paper cites.
Information Retrieval
C. J. Van Rijsbergen, · 1979
Earlier work this paper cites.
“The comparison and evaluation of forecasters,”
Morris H DeGroot and Stephen E Fienberg, · 1983
Earlier work this paper cites.
“Evidence-based medicine as bayesian decision-making,”
Deborah Ashby and Adrian F. M. Smith, · 2000
Earlier work this paper cites.
“Assessing the accuracy of prediction algorithms for classification: an overview,”
P. Baldi, S Brunak, Y. Chauvin, C. A. F. Andersen, and H. Nielsen, · 2000
Earlier work this paper cites.
“Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods,”
J. C. Platt, · 2000
Earlier work this paper cites.
The Elements of Statistical Learning
T. Hastie, R. Tibshirani, and J. Friedman, · 2001
Earlier work this paper cites.
“The foundations of cost-sensitive learning,”
Charles Elkan, · 2001
Earlier work this paper cites.
Pattern Classification
R. Duda, P. Hart, and D. Stork, · 2001
Earlier work this paper cites.
“Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers,”
Bianca Zadrozny and Charles Elkan, · 2001
Earlier work this paper cites.
“Comparing two k-category assignments by a k-category correlation coefficient,”
J. Gorodkin, · 2004
Earlier work this paper cites.
Pattern Recognition and Machine Learning
C. M. Bishop, · 2006
Cited alongside, same era.
“On calibration of language recognition scores,”
N. Brümmer and D. A. van Leeuwen, · 2006
Cited alongside, same era.
“An introduction to application-independent evaluation of speaker recognition systems,”
D. A Van Leeuwen and N. Brümmer, · 2007
Cited alongside, same era.
Measuring, Refining and Calibrating Speaker and Language Information Extracted from Speech
N. Brümmer, · 2010
Cited alongside, same era.
“Bayesian decision analysis for choosing between diagnostic/prognostic prediction procedures,”
John Kornak and Ying Lu, · 2011
Cited alongside, same era.
“Strictly proper scoring rules, prediction, and estimation,”
Tilmann Gneiting and Adrian E. Raftery, · 2012
Cited alongside, same era.
“Binary classifier calibration: Non-parametric approach,”
Mahdi Pakdaman Naeini, Gregory F Cooper, and Milos Hauskrecht, · 2015
Later among the works it cites.
“Net benefit approaches to the evaluation of prediction models, molecular markers, and diagnostic tests,”
Andrew J Vickers, Ben Van Calster, and Ewout W Steyerberg, · 2016
Later among the works it cites.
“On calibration of modern neural networks,”
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger, · 2017
Later among the works it cites.
“Radiomics based on adapted diffusion kurtosis imaging helps to clarify most mammographic findings suspicious for cancer,”
Sebastian Bickelhaupt, Paul Ferdinand Jaeger, Frederik Bernd Laun, Wolfgang Lederer, Heidi Daniel, Tristan Anselm Kuder, Lorenz Wuesthof, Daniel Paech, David Bonekamp, Alexander Radbruch, Stefan Delorme, Heinz-Peter Schlemmer, Franziska Hildegard Steudle, and Klaus Hermann Maier-Hein, · 2018
Later among the works it cites.
“Tied normal variance–mean mixtures for linear score calibration,”
Sandro Cumani and Pietro Laface, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An introduction to statistical learning
G. James, D. Witten, T. Hastie, and R. Tibshirani, · 2013
Cited alongside, same era.
“Likelihood-ratio calibration using prior-weighted proper scoring rules,”
N. Brümmer and G. Doddington, · 2013
Cited alongside, same era.
“Theory and applications of proper scoring rules,”
Alexander Philip Dawid and Monica Musio, · 2014
Cited alongside, same era.
“Generative modelling for unsupervised score calibration,”
Niko Brümmer and Daniel Garcia-Romero, · 2014
Cited alongside, same era.
“Novel decompositions of proper scoring rules for classification: Score adjustment as precursor to calibration,”
Meelis Kull and Peter Flach, · 2015
Cited alongside, same era.
“Obtaining well calibrated probabilities using bayesian binning,”
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht, · 2015
Cited alongside, same era.
“Measuring calibration in deep learning.,”
Jeremy Nixon, Michael W Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran, · 2019
Later among the works it cites.
“Calibration tests in multi-class classification: A unifying framework,”
David Widmann, Fredrik Lindsten, and Dave Zachariah, · 2019
Later among the works it cites.
“The advantages of the matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation,”
Davide Chicco and Giuseppe Jurman, · 2020
Later among the works it cites.
“Out of a Hundred Trials, How Many Errors Does Your Speaker Verifier Make?,”
N. Brümmer, L. Ferrer, and A. Swart, · 2021
Later among the works it cites.
“Better uncertainty calibration via proper scores for classification and beyond,”
Sebastian Gregor Gruber and Florian Buettner, · 2022
Closest in time.
“Twitter-COMMs: Detecting climate, COVID, and military multimodal misinformation,”
Giscard Biamby, Grace Luo, Trevor Darrell, and Anna Rohrbach, · 2022
Closest in time.
“Deployment of image analysis algorithms under prevalence shifts,”
Patrick Godau, Piotr Kalinowski, Evangelia Christodoulou, Annika Reinke, Minu Tizabi, Luciana Ferrer, Paul Jäger, and Lena Maier-Hein, · 2023
Closest in time.