2023

Task Contamination: Language Models May Not Be Few-Shot Anymore

Li, Changmao, Flanigan, Jeffrey

Understand

Large language models (LLMs) offer impressive performance in various zero-shot and few-shot tasks.

  • However, their success in zero-shot and few-shot settings may be affected by task contamination, a potential limitation that has not been thoroughly examined.
  • This paper investigates how zero-shot and few-shot performance of LLMs has changed chronologically over time.
  • Utilizing GPT-3 series models and several other recent open-sourced LLMs, and controlling for dataset difficulty, we find that on datasets released before the LLM training data creation date, LLMs perform surprisingly better than on datasets released after.

Reading the bibliography…