Fetching the paper…

COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations · Around