Fetching the paper…

Self-Supervised Representation Learning for Speech Using Visual Grounding and Masked Language Modeling · Around