Fetching the paper…

Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution · Around