Fetching the paper…

Efficient Hybrid Inference for LLMs: Reward-Based Token Modelling with Selective Cloud Assistance · Around