Fetching the paper…

Self-supervised Contrastive Cross-Modality Representation Learning for Spoken Question Answering · Around