Fetching the paper…

ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models · Around