2020

DQI: Measuring Data Quality in NLP

Mishra, Swaroop, Arunkumar, Anjana, Sachdeva, Bhavdeep et al.

Understand

Neural language models have achieved human level performance across several NLP datasets.

  • However, recent studies have shown that these models are not truly learning the desired task; rather, their high performance is attributed to overfitting using spurious biases, which suggests that the capabilities of AI systems have been over-estimated.
  • We introduce a generic formula for Data Quality Index (DQI) to help dataset creators create datasets free of such unwanted biases.
  • We evaluate this formula using a recently proposed approach for adversarial filtering, AFLite.

Reading the bibliography…