Fetching the paper…

Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction · Around