Fetching the paper…

Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs · Around