2020

Measuring Massive Multitask Language Understanding

Hendrycks, Dan, Burns, Collin, Basart, Steven et al.

Understand

We propose a new test to measure a text model's multitask accuracy.

  • The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more.
  • To attain high accuracy on this test, models must possess extensive world knowledge and problem solving ability.
  • We find that while most recent models have near random-chance accuracy, the very largest GPT-3 model improves over random chance by almost 20 percentage points on average.

Reading the bibliography…