Fetching the paper…

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills · Around