Fetching the paper…

Learning Explicit and Implicit Latent Common Spaces for Audio-Visual Cross-Modal Retrieval · Around