Fetching the paper…

GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference · Around