2024

One Thousand and One Pairs: A "novel" challenge for long-context language models

Karpinska, Marzena, Thai, Katherine, Lo, Kyle et al.

Understand

Synthetic long-context LLM benchmarks (e.g., "needle-in-the-haystack") test only surface-level retrieval capabilities, but how well can long-context LLMs retrieve, synthesize, and reason over information across book-length inputs? We address this question by creating NoCha, a dataset of 1,001 minimally different pairs of true and false claims about 67 recently-published English fictional books, written by human readers of those books.

  • In contrast to existing long-context benchmarks, our annotators confirm that the largest share of pairs in NoCha require global reasoning over the entire book to verify.
  • Our experiments show that while human readers easily perform this task, it is enormously challenging for all ten long-context LLMs that we evaluate: no open-weight model performs above random chance (despite their strong performance on synthetic benchmarks), while GPT-4o achieves the highest accuracy at 55.8%.
  • Further analysis reveals that (1) on average, models perform much better on pairs that require only sentence-level retrieval vs.

Reading the bibliography…