Fetching the paper…

FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines · Around