2022

InterFair: Debiasing with Natural Language Feedback for Fair Interpretable Predictions

Majumder, Bodhisattwa Prasad, He, Zexue, McAuley, Julian

Understand

Debiasing methods in NLP models traditionally focus on isolating information related to a sensitive attribute (e.g., gender or race).

  • We instead argue that a favorable debiasing method should use sensitive information 'fairly,' with explanations, rather than blindly eliminating it.
  • This fair balance is often subjective and can be challenging to achieve algorithmically.
  • We explore two interactive setups with a frozen predictive model and show that users able to provide feedback can achieve a better and fairer balance between task performance and bias mitigation.

Reading the bibliography…