2022

Jury Learning: Integrating Dissenting Voices into Machine Learning Models

Gordon, Mitchell L., Lam, Michelle S., Park, Joon Sung et al.

Understand

Whose labels should a machine learning (ML) algorithm learn to emulate? For ML tasks ranging from online comment toxicity to misinformation detection to medical diagnosis, different groups in society may have irreconcilable disagreements about ground truth labels.

  • Supervised ML today resolves these label disagreements implicitly using majority vote, which overrides minority groups' labels.
  • We introduce jury learning, a supervised ML approach that resolves these disagreements explicitly through the metaphor of a jury: defining which people or groups, in what proportion, determine the classifier's prediction.
  • For example, a jury learning model for online toxicity might centrally feature women and Black jurors, who are commonly targets of online harassment.

Reading the bibliography…