Fetching the paper…

DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference · Around