2024

Long-context LLMs Struggle with Long In-context Learning

Li, Tianle, Zhang, Ge, Do, Quy Duc et al.

Understand

Large Language Models (LLMs) have made significant strides in handling long sequences.

  • Some models like Gemini could even to be capable of dealing with millions of tokens.
  • However, their performance evaluation has largely been confined to metrics like perplexity and synthetic tasks, which may not fully capture their true abilities in more challenging, real-world scenarios.
  • We introduce a benchmark (LongICLBench) for long in-context learning in extreme-label classification using six datasets with 28 to 174 classes and input lengths from 2K to 50K tokens.

Reading the bibliography…