Fetching the paper…

Constraint-aware and Ranking-distilled Token Pruning for Efficient Transformer Inference · Around