Fetching the paper…
Reading the bibliography…
Direct speech-to-speech translation (S2ST) is among the most challenging problems in the translation paradigm due to the significant scarcity of S2ST data.
“Janus-iii: Speech-to-speech translation in multiple languages,”
A.Lavie, A.Waibel, et al., · 1997
Earlier work this paper cites.
“The atr multilingual speech-to-speech translation system,”
S.Nakamura, K.Markov, et al., · 2006
Earlier work this paper cites.
“Improved signal-to-noise ratio estimation for speech enhancement,”
C.Plapous, C.Marro, et al., · 2006
Earlier work this paper cites.
“MUSAN: A Music, Speech, and Noise Corpus,” 2015,
D.Snyder, G.Chen, et al., · 2015
Earlier work this paper cites.
J. L.Ba, J. R.Kiros, et al., · 2016
Earlier work this paper cites.
“Attention is all you need,”
A.Vaswani, N.Shazeer, et al., · 2017
Earlier work this paper cites.
“Towards speech-to-text translation without speech recognition,”
S.Bansal, H.Kamper, et al., · 2017
Earlier work this paper cites.
“Understanding back-translation at scale,”
S.Edunov, M.Ott, et al., · 2018
Earlier work this paper cites.
“Ted-lium 3: twice as much data and corpus repartition for experiments on speaker adaptation,”
F.Hernandez, V.Nguyen, et al., · 2018
Earlier work this paper cites.
“Fastspeech: Fast, robust and controllable text to speech,”
Y.Ren, Y.Ruan, et al., · 2019
Earlier work this paper cites.
“Direct speech-to-speech translation with a sequence-to-sequence model,”
Y.Jia, R. J.Weiss, et al., · 2019
Earlier work this paper cites.
“Leveraging weakly supervised data to improve end-to-end speech-to-text translation,”
Y.Jia, M.Johnson, et al., · 2019
Earlier work this paper cites.
“Ccnet: Extracting high quality monolingual datasets from web crawl data,”
G.Wenzek, M.-A.Lachaux, et al., · 2019
Cited alongside, same era.
“Cross-lingual language model pretraining,”
A.Conneau and G.Lample, · 2019
Cited alongside, same era.
“Common voice: A massively-multilingual speech corpus,”
R.Ardila, M.Branson, et al., · 2019
Cited alongside, same era.
“Self-training for end-to-end speech recognition,”
J.Kahn, A.Lee, et al., · 2020
Cited alongside, same era.
“Cross-lingual retrieval for iterative self-supervised training,”
C.Tran, Y.Tang, et al., · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
J.Kim, J.Kong, et al., · 2021
Later among the works it cites.
“Direct speech-to-speech translation with discrete units,”
A.Lee, P.-J.Chen, et al., · 2021
Later among the works it cites.
“On generative spoken language modeling from raw audio,”
K.Lakhotia, E.Kharitonov, et al., · 2021
Later among the works it cites.
“Speech resynthesis from discrete disentangled self-supervised representations,”
A.Polyak, Y.Adi, et al., · 2021
Later among the works it cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
W.-N.Hsu, B.Bolte, et al., · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A.Baevski, Y.Zhou, et al., · 2020
Cited alongside, same era.
“Multilingual denoising pre-training for neural machine translation,”
Y.Liu, J.Gu, et al., · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
A.Gulati, J.Qin, et al., · 2020
Cited alongside, same era.
“Covost 2 and massively multilingual speech-to-text translation,”
C.Wang, A.Wu, et al., · 2020
Cited alongside, same era.
“Europarl-st: A multilingual corpus for speech translation of parliamentary debates,”
J.Iranzo-Sánchez, J. A.Silvestre-Cerda, et al., · 2020
Cited alongside, same era.
“Mls: A large-scale multilingual dataset for speech research,”
V.Pratap, Q.Xu, et al., · 2020
Cited alongside, same era.
“Libri-light: A benchmark for asr with limited or no supervision,”
J.Kahn, M.Rivière, et al., · 2020
Cited alongside, same era.
C.Wang, A.Wu, et al., · 2021
Later among the works it cites.
“Speech resynthesis from discrete disentangled self-supervised representations,”
A.Polyak, Y.Adi, et al., · 2021
Later among the works it cites.
“The multilingual tedx corpus for speech recognition and translation,”
S.Elizabeth, W.Matthew, et al., · 2021
Later among the works it cites.
C.Wang, M.Riviere, et al., · 2021
Later among the works it cites.
“Contrastive clustering to mine pseudo parallel data for unsupervised translation,”
X.-P.Nguyen, H.Gong, et al., · 2022
Closest in time.
S.Popuri, P.-J.Chen, et al., · 2022
Closest in time.
“Textless speech-to-speech translation on real data,”
A.Lee, H.Gong, et al., · 2022
Closest in time.