Fetching the paper…
Reading the bibliography…
This paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-translation.
“JANUS-III: Speech-to-speech translation in multiple languages,”
Alon Lavie et al, · 1997
Earlier work this paper cites.
“The ATR multilingual speech-to-speech translation system,”
Satoshi Nakamura et al, · 2006
Earlier work this paper cites.
Verbmobil: Foundations of speech-to-speech translation
Wolfgang Wahlster, · 2013
Earlier work this paper cites.
“Improving neural machine translation models with monolingual data,”
Rico Sennrich et al, · 2015
Earlier work this paper cites.
“Word translation without parallel data,”
Alexis Conneau et al, · 2017
Earlier work this paper cites.
“Unsupervised machine translation using monolingual corpora only,”
Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato, · 2017
Earlier work this paper cites.
“Learning bilingual word embeddings with (almost) no bilingual data,”
Mikel Artetxe et al, · 2017
Earlier work this paper cites.
“Unpaired image-to-image translation using cycle-consistent adversarial networks,”
Jun-Yan Zhu et al, · 2017
Earlier work this paper cites.
“Unsupervised neural machine translation,”
Mikel Artetxe et al, · 2018
Earlier work this paper cites.
“Unsupervised machine translation using monolingual corpora only,”
Guillaume Lample et al, · 2018
Earlier work this paper cites.
“A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings,”
Mikel Artetxe et al, · 2018
Earlier work this paper cites.
“Word translation without parallel data,”
Guillaume Lample et al, · 2018
Earlier work this paper cites.
“Efficient neural audio synthesis,”
N. Kalchbrenner et al, · 2018
Earlier work this paper cites.
“Direct speech-to-speech translation with a sequence-to-sequence model,”
Ye Jia et al, · 2019
Earlier work this paper cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park et al, · 2019
Earlier work this paper cites.
“Speech-to-speech translation between untranscribed unknown languages,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2019
Earlier work this paper cites.
“Towards unsupervised speech-to-text translation,”
Yu-An Chung, Wei-Hung Weng, Schrasing Tong, and James Glass, · 2019
Earlier work this paper cites.
“Leveraging weakly supervised data to improve end-to-end speech-to-text translation,”
Ye Jia et al, · 2019
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski et al, · 2020
Cited alongside, same era.
“Common Voice: A massively-multilingual speech corpus,”
Rosana Ardila et al, · 2020
Cited alongside, same era.
Jonathan Shen et al, · 2020
Cited alongside, same era.
“CoVoST 2 and massively multilingual speech-to-text translation,”
Changhan Wang et al, · 2020
Cited alongside, same era.
“Unified speech-text pre-training for speech translation and recognition,”
Yun Tang et al, · 2022
Later among the works it cites.
“Leveraging pseudo-labeled data to improve direct speech-to-speech translation,”
Qianqian Dong et al, · 2022
Later among the works it cites.
“Simple and effective unsupervised speech translation,”
Changhan Wang et al, · 2022
Later among the works it cites.
“Unity: Two-pass direct speech-to-speech translation with discrete units,”
Hirofumi Inaguma et al, · 2022
Later among the works it cites.
“Joint pre-training with speech and bilingual text for direct speech to speech translation,”
Kun Wei et al, · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ann Lee et al, · 2021
Cited alongside, same era.
“Uwspeech: Speech to speech translation for unwritten languages,”
Chen Zhang et al, · 2021
Cited alongside, same era.
“Textless speech-to-speech translation on real data,”
Ann Lee et al, · 2021
Cited alongside, same era.
“Transformer-based direct speech-to-speech translation with transcoder,”
Takatomo Kano et al, · 2021
Cited alongside, same era.
“W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,”
Yu-An Chung et al, · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu et al, · 2021
Cited alongside, same era.
“Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing,”
Junyi Ao et al, · 2021
Cited alongside, same era.
Later among the works it cites.
“Soundstream: An end-to-end neural audio codec,”
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi, · 2022
Later among the works it cites.
“High fidelity neural audio compression,”
Alexandre Défossez et al, · 2022
Later among the works it cites.
“The yitrans end-to-end speech translation system for iwslt 2022 offline shared task,”
Ziqiang Zhang et al, · 2022
Later among the works it cites.
“Speechlm: Enhanced speech pre-training with unpaired textual data,”
Ziqiang Zhang et al, · 2022
Later among the works it cites.
“WaveFit: An iterative and non-autoregressive neural vocoder based on fixed-point iteration,”
Yuma Koizumi et al, · 2022
Later among the works it cites.
“Residual adapters for few-shot text-to-speech speaker adaptation,”
Nobuyuki Morioka et al, · 2022
Later among the works it cites.
“CVSS corpus and massively multilingual speech-to-speech translation,”
Ye Jia et al, · 2022
Later among the works it cites.
“Audiopalm: A large language model that can speak and listen,”
Paul K Rubenstein et al, · 2023
Closest in time.
“Textless direct speech-to-speech translation with discrete speech representation,”
Xinjian Li et al, · 2023
Closest in time.
Minsu Kim et al, · 2023
Closest in time.
“Joint pre-training with speech and bilingual text for direct speech to speech translation,”
Kun Wei et al, · 2023
Closest in time.
“Dub: Discrete unit back-translation for speech translation,”
Dong Zhang et al, · 2023
Closest in time.