Fetching the paper…

HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models · Around