Fetching the paper…
Reading the bibliography…
End-to-end speech-to-speech translation (S2ST) without relying on intermediate text representations is a rapidly emerging frontier of research.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proc. ICML , 2006
2006
Earlier work this paper cites.
O. Bojar, C. Buck, C. Callison-Burch, C. Federmann et al. , “Findings of the 2013 Workshop on Statistical Machine Translation,” in Proc. Workshop on Statistical Machine Translation , 2013
2013
Earlier work this paper cites.
M. J. Gales, K. M. Knill, A. Ragni, and S. P. Rath, “Speech recognition and keyword spotting for low-resource languages: Babel project research at cued,” in Proc. SLTU , 2014
2014
Earlier work this paper cites.
O. Bojar, R. Chatterjee, C. Federmann, B. Haddow et al. , “Findings of the 2015 workshop on statistical machine translation,” in Proc. Workshop on Statistical Machine Translation , 2015
2015
Earlier work this paper cites.
R. J. Weiss, J. Chorowski et al. , “Sequence-to-sequence models can directly translate foreign speech,” in Proc. Interspeech , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez et al. , “Attention is all you need,” in Proc. NeurIPS , 2017
2017
Earlier work this paper cites.
O. Bojar, R. Chatterjee, C. Federmann, Y. Graham et al. , “Findings of the 2017 conference on machine translation (WMT17),” in Proc. Conference on Machine Translation , 2017
2017
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness et al. , “Overcoming catastrophic forgetting in neural networks,” PNAS , 2017
2017
Earlier work this paper cites.
A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” arXiv , 2018
2018
Earlier work this paper cites.
A. Bérard, L. Besacier et al. , “End-to-end automatic speech translation of audiobooks,” in Proc. ICASSP , 2018
2018
Earlier work this paper cites.
A. Anastasopoulos and D. Chiang, “Tied multitask learning for neural speech translation,” in Proc. NAACL , 2018
2018
Earlier work this paper cites.
O. Bojar, C. Federmann, M. Fishel, Y. Graham et al. , “Findings of the 2018 conference on machine translation (WMT18),” in Proc. Conference on Machine Translation , 2018
2018
Earlier work this paper cites.
Y. Qi, D. S. Sachan, M. Felix, S. J. Padmanabhan, and G. Neubig, “When and why are pre-trained word embeddings useful for neural machine translation?” in Proc. NAACL , 2018
2018
Earlier work this paper cites.
Y. Jia, R. J. Weiss, F. Biadsy, W. Macherey, M. Johnson, Z. Chen, and Y. Wu, “Direct speech-to-speech translation with a sequence-to-sequence model,” in Proc. Interspeech , 2019
2019
Earlier work this paper cites.
A. Tjandra, S. Sakti et al. , “Speech-to-speech translation between untranscribed unknown languages,” in Proc. ASRU , 2019
2019
Earlier work this paper cites.
S. Bansal, H. Kamper, K. Livescu, A. Lopez, and S. Goldwater, “Pre-training on high-resource speech recognition improves low-resource speech-to-text translation,” in Proc. NAACL , 2019
2019
Earlier work this paper cites.
Y. Jia, M. Johnson, W. Macherey, R. J. Weiss, Y. Cao, C.-C. Chiu et al. , “Leveraging weakly supervised data to improve end-to-end speech-to-text translation,” in Proc. ICASSP , 2019
2019
Earlier work this paper cites.
Y. Liu, H. Xiong, Z. He, J. Zhang et al. , “End-to-end speech translation with knowledge distillation,” in Proc. Interspeech , 2019
2019
Earlier work this paper cites.
A. Conneau and G. Lample, “Cross-lingual language model pretraining,” in Proc. NeurIPS , 2019
2019
Cited alongside, same era.
L. Barrault, O. Bojar, M. R. Costa-Jussa, C. Federmann et al. , “Findings of the 2019 conference on machine translation (WMT19),” in Proc. Conference on Machine Translation , 2019
2019
Cited alongside, same era.
J. Shen, P. Nguyen, Y. Wu et al. , “Lingvo: A modular and scalable framework for sequence-to-sequence modeling,” arXiv , 2019
2019
Cited alongside, same era.
N. Arivazhagan et al. , “Massively multilingual neural machine translation in the wild: Findings and challenges,” arXiv , 2019
2019
Cited alongside, same era.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. NeurIPS , 2020
2020
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia et al. , “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” TASLP , 2021
2021
Later among the works it cites.
Y.-A. Chung, Y. Zhang, W. Han, C.-C. Chiu, J. Qin et al. , “w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,” in Proc. ASRU , 2021
2021
Later among the works it cites.
X. Li, C. Wang, Y. Tang, C. Tran, Y. Tang, J. Pino, A. Baevski, A. Conneau, and M. Auli, “Multilingual speech translation from efficient finetuning of pretrained models,” in Proc. ACL , 2021
2021
Later among the works it cites.
C. Wang, A. Wu, J. Pino, A. Baevski, M. Auli, and A. Conneau, “Large-scale self-and semi-supervised learning for speech translation,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Wu, C. Wang, J. Pino et al. , “Self-supervised representations improve end-to-end speech translation,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang et al. , “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis et al. , “Common Voice: A massively-multilingual speech corpus,” in Proc. LREC , 2020
2020
Cited alongside, same era.
V. Pratap, Q. Xu, A. Sriram et al. , “MLS: A large-scale multilingual dataset for speech research,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
M. Joshi, D. Chen, Y. Liu et al. , “SpanBERT: Improving pre-training by representing and predicting spans,” TACL , 2020
2020
Cited alongside, same era.
L. Barrault, M. Biesialska, O. Bojar, M. R. Costa-jussà et al. , “Findings of the 2020 conference on machine translation (WMT20),” in Proc. Conference on Machine Translation , 2020
2020
Cited alongside, same era.
D. S. Park, Y. Zhang, Y. Jia et al. , “Improved noisy student training for automatic speech recognition,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
A. Babu, C. Wang, A. Tjandra et al. , “XLS-R: Self-supervised cross-lingual speech representation learning at scale,” arXiv , 2021
2021
Later among the works it cites.
R. Zheng, J. Chen, M. Ma, and L. Huang, “Fused acoustic and text encoding for multimodal bilingual pretraining and speech translation,” in Proc. ICML , 2021
2021
Later among the works it cites.
A. Bapna, Y.-A. Chung, N. Wu, A. Gulati, Y. Jia, J. H. Clark, M. Johnson et al. , “SLAM: A unified encoder for speech and language modeling via speech-text joint pre-training,” arXiv , 2021
2021
Later among the works it cites.
C. Wang, A. Wu, J. Gu, and J. Pino, “CoVoST 2 and massively multilingual speech translation,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
P.-A. Duquenne et al. , “Multimodal and multilingual embeddings for large-scale speech mining,” in Proc. NeurIPS , 2021
2021
Later among the works it cites.
C. Wang, M. Rivière, A. Lee, A. Wu et al. , “VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,” in Proc. ACL , 2021
2021
Later among the works it cites.
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, “mT5: A massively multilingual pre-trained text-to-text transformer,” in Proc. NAACL , 2021
2021
Later among the works it cites.
Y. Jia, H. Zen et al. , “PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
Y. Jia, M. T. Ramanovich, T. Remez, and R. Pomerantz, “Translatotron 2: High-quality direct speech-to-speech translation with voice preservation,” in Proc. ICML , 2022
2022
Closest in time.
A. Lee, P.-J. Chen, C. Wang, J. Gu, X. Ma et al. , “Direct speech-to-speech translation with discrete units,” in Proc. ACL , 2022
2022
Closest in time.
Y. Jia, M. T. Ramanovich et al. , “CVSS corpus and massively multilingual speech-to-speech translation,” in Proc. LREC , 2022
2022
Closest in time.
A. Bapna, C. Cherry, Y. Zhang et al. , “mSLAM: Massively multilingual joint pre-training for speech and text,” arXiv , 2022
2022
Closest in time.