Fetching the paper…

FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving · Around