Fetching the paper…

MAM: Masked Acoustic Modeling for End-to-End Speech-to-Text Translation · Around