Fetching the paper…

LongFuncEval: Measuring the effectiveness of long context models for function calling · Around