Fetching the paper…
Reading the bibliography…
The amount of labeled data to train models for speech tasks is limited for most languages, however, the data scarcity is exacerbated for speech translation which requires labeled data covering two different languages.
Finite-state transducers in language and speech processing
Mehryar Mohri. 1997 · 1997
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, Pierre-Antoine Manzagol, and Léon Bottou. 2010 · 2010
Earlier work this paper cites.
Phonemizer
Mathieu Bernard. 2015 · 2015
Earlier work this paper cites.
Listen and translate: A proof of concept for end-to-end speech-to-text translation
Alexandre Bérard, Olivier Pietquin, Laurent Besacier, and Christophe Servan. 2016 · 2016
Earlier work this paper cites.
An attentional model for speech translation without transcription
Long Duong, Antonios Anastasopoulos, David Chiang, Steven Bird, and Trevor Cohn. 2016 · 2016
Earlier work this paper cites.
Towards speech-to-text translation without speech recognition
Sameer Bansal, Herman Kamper, Adam Lopez, and Sharon Goldwater. 2017 · 2017
Earlier work this paper cites.
Must-c: a multilingual speech translation corpus
Mattia A Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2019a · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Sequence-to-sequence models can directly translate foreign speech
Ron J Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen. 2017 · 2017
Earlier work this paper cites.
Unsupervised neural machine translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2018 · 2018
Earlier work this paper cites.
Low-resource speech-to-text translation
Sameer Bansal, Herman Kamper, Karen Livescu, Adam Lopez, and Sharon Goldwater. 2018 · 2018
Earlier work this paper cites.
Word translation without parallel data
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018 · 2018
Earlier work this paper cites.
Augmenting librispeech with french translations: A multimodal corpus for direct speech translation evaluation
Ali Can Kocabiyikoglu, Laurent Besacier, and Olivier Kraif. 2018 · 2018
Earlier work this paper cites.
Unsupervised machine translation using monolingual corpora only
Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018 · 2018
Cited alongside, same era.
Completely unsupervised phoneme recognition by adversarially learning mapping relationships from audio embeddings
Da-Rong Liu, Kuan-Yu Chen, Hung-Yi Lee, and Lin shan Lee. 2018 · 2018
Cited alongside, same era.
End-to-end speech translation with the transformer
Laura Cross Vila, Carlos Escolano, José AR Fonollosa, and Marta R Costa-Jussa. 2018 · 2018
Cited alongside, same era.
Pre-training on high-resource speech recognition improves low-resource speech-to-text translation
Sameer Bansal, Herman Kamper, Karen Livescu, Adam Lopez, and Sharon Goldwater. 2019 · 2019
Cited alongside, same era.
Completely unsupervised speech recognition by a generative adversarial network harmonized with iteratively refined hidden markov models
Kuan-Yu Chen, Che-Ping Tsai, Da-Rong Liu, Hung-Yi Lee, and Lin shan Lee. 2019 · 2019
Cited alongside, same era.
Simulspeech: End-to-end simultaneous speech to text translation
Yi Ren, Jinglin Liu, Xu Tan, Chen Zhang, Tao Qin, Zhou Zhao, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
Fairseq S2T: Fast speech-to-text modeling with fairseq
Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, and Juan Pino. 2020 · 2020
Later among the works it cites.
Xls-r: Self-supervised cross-lingual speech representation learning at scale
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, et al. 2021 · 2021
Later among the works it cites.
Unsupervised speech recognition
Alexei Baevski, Wei-Ning Hsu, Alexis Conneau, and Michael Auli. 2021 · 2021
Later among the works it cites.
Cascade versus direct speech translation: Do the differences still make a difference?
Luisa Bentivogli, Mauro Cettolo, Marco Gaido, Alina Karakanta, Alberto Martinelli, Matteo Negri, and Marco Turchi. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards unsupervised speech-to-text translation
Yu-An Chung, Wei-Hung Weng, Schrasing Tong, and James Glass. 2019 · 2019
Cited alongside, same era.
Adapting transformer to end-to-end spoken language translation
Mattia A Di Gangi, Matteo Negri, and Marco Turchi. 2019b · 2019
Cited alongside, same era.
Direct speech-to-speech translation with a sequence-to-sequence model
Ye Jia, Ron J. Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Z. Chen, and Yonghui Wu. 2019 · 2019
Cited alongside, same era.
Neural speech synthesis with transformer network
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019 · 2019
Cited alongside, same era.
wav2vec: Unsupervised pre-training for speech recognition
Setffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Unsupervised speech recognition via segmental empirical output distribution matching
Chih-Kuan Yeh, Jianshu Chen, Chengzhu Yu, and Dong Yu. 2019 · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Cited alongside, same era.
AlloST: Low-Resource Speech Translation Without Source Transcription
Yao-Fei Cheng, Hung-Shin Lee, and Hsin-Min Wang. 2021 · 2021
Later among the works it cites.
Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training
Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski, Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Jacob Kahn, Ann Lee, Ronan Collobert, Gabriel Synnaeve, and Michael Auli. 2021 · 2021
Later among the works it cites.
Transformer-based direct speech-to-speech translation with transcoder
Takatomo Kano, Sakriani Sakti, and Satoshi Nakamura. 2021 · 2021
Later among the works it cites.
Multilingual speech translation from efficient finetuning of pretrained models
Xian Li, Changhan Wang, Yun Tang, Chau Tran, Yuqing Tang, Juan Pino, Alexei Baevski, Alexis Conneau, and Michael Auli. 2021 · 2021
Later among the works it cites.
R-drop: Regularized dropout for neural networks
Lijun Wu, Juntao Li, Yue Wang, Qi Meng, Tao Qin, Wei Chen, Min Zhang, Tie-Yan Liu, et al. 2021 · 2021
Later among the works it cites.
Ethnologue: Languages of the world, 25th edition
M. Paul Lewis, Gary F. Simon, and Charles D. Fennig. 2022 · 2022
Closest in time.
Analyzing the robustness of unsupervised speech recognition
Guan-Ting Lin, Chan-Jan Hsu, Da-Rong Liu, Hung-Yi Lee, and Yu Tsao. 2022 · 2022
Closest in time.
Unsupervised text-to-speech synthesis by unsupervised automatic speech recognition
Junrui Ni, Liming Wang, Heting Gao, Kaizhi Qian, Yang Zhang, Shiyu Chang, and Mark Hasegawa-Johnson. 2022 · 2022
Closest in time.