Fetching the paper…

Seeing voices and hearing voices: learning discriminative embeddings using cross-modal self-supervision · Around