Fetching the paper…

Efficient LLM Inference over Heterogeneous Edge Networks with Speculative Decoding · Around