Fetching the paper…

Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA · Around