Fetching the paper…

TCRA-LLM: Token Compression Retrieval Augmented Large Language Model for Inference Cost Reduction · Around