Fetching the paper…

Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech Translation · Around