Fetching the paper…

Simple linear attention language models balance the recall-throughput tradeoff · Around