2024

RULER: What's the Real Context Size of Your Long-Context Language Models?

Hsieh, Cheng-Ping, Sun, Simeng, Kriman, Samuel et al.

Understand

The needle-in-a-haystack (NIAH) test, which examines the ability to retrieve a piece of information (the "needle") from long distractor texts (the "haystack"), has been widely adopted to evaluate long-context language models (LMs).

  • However, this simple retrieval-based test is indicative of only a superficial form of long-context understanding.
  • To provide a more comprehensive evaluation of long-context LMs, we create a new synthetic benchmark RULER with flexible configurations for customized sequence length and task complexity.
  • RULER expands upon the vanilla NIAH test to encompass variations with diverse types and quantities of needles.

Reading the bibliography…