Fetching the paper…

When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models · Around