2022

Stop Measuring Calibration When Humans Disagree

Baan, Joris, Aziz, Wilker, Plank, Barbara et al.

Understand

Calibration is a popular framework to evaluate whether a classifier knows when it does not know - i.e., its predictive probabilities are a good indication of how likely a prediction is to be correct.

  • Correctness is commonly estimated against the human majority class.
  • Recently, calibration to human majority has been measured on tasks where humans inherently disagree about which class applies.
  • We show that measuring calibration to human majority given inherent disagreements is theoretically problematic, demonstrate this empirically on the ChaosNLI dataset, and derive several instance-level measures of calibration that capture key statistical properties of human judgements - class frequency, ranking and entropy.

Reading the bibliography…