Fetching the paper…

Linear Transformers with Learnable Kernel Functions are Better In-Context Models · Around