Fetching the paper…

TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling · Around