Fetching the paper…
Reading the bibliography…
We present a direct speech-to-speech translation (S2ST) model that translates speech from one language to speech in another language without relying on intermediate text generation.
End-to-end ASR: from supervised to semi-supervised learning with modern architectures
Gabriel Synnaeve, Qiantong Xu, Jacob Kahn, Tatiana Likhomanenko, Edouard Grave, Vineel Pratap, Anuroop Sriram, Vitaliy Liptchinsky, and Ronan Collobert. 2019 · 1911
Earlier work this paper cites.
JANUS-III: Speech-to-speech translation in multiple languages
Alon Lavie, Alex Waibel, Lori Levin, Michael Finke, Donna Gates, Marsal Gavalda, Torsten Zeppenfeld, and Puming Zhan. 1997 · 1997
Earlier work this paper cites.
On the integration of speech recognition and statistical machine translation
Evgeny Matusov, Stephan Kanthak, and Hermann Ney. 2005 · 2005
Earlier work this paper cites.
Prosody generation for speech-to-speech translation
PD Aguero, Jordi Adell, and Antonio Bonafonte. 2006 · 2006
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
The ATR multilingual speech-to-speech translation system
Satoshi Nakamura, Konstantin Markov, Hiromi Nakaiwa, Gen-ichiro Kikui, Hisashi Kawai, Takatoshi Jitsuhiro, J-S Zhang, Hirofumi Yamamoto, Eiichiro Sumita, and Seiichi Yamamoto. 2006 · 2006
Earlier work this paper cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020 · 2006
Earlier work this paper cites.
UWSpeech: Speech to speech translation for unwritten languages
Chen Zhang, Xu Tan, Yi Ren, Tao Qin, Kejun Zhang, and Tie-Yan Liu. 2020 · 2006
Earlier work this paper cites.
Intent transfer in speech-to-speech machine translation
Gopala Krishna Anumanchipalli, Luis C Oliveira, and Alan W Black. 2012 · 2012
Earlier work this paper cites.
Exploring wav2vec 2.0 on speaker verification and language identification
Zhiyun Fan, Meng Li, Shiyu Zhou, and Bo Xu. 2020 · 2012
Earlier work this paper cites.
Fisher and CALLHOME Spanish–English speech translation
Matt Post, Gaurav Kumar, Adam Lopez, Damianos Karakos, Chris Callison-Burch, and Sanjeev Khudanpur. 2014 · 2014
Earlier work this paper cites.
Librispeech: an ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Listen and translate: A proof of concept for end-to-end speech-to-text translation
Alexandre Bérard, Olivier Pietquin, Christophe Servan, and Laurent Besacier. 2016 · 2016
Earlier work this paper cites.
Toward expressive speech translation: A unified sequence-to-sequence LSTMs approach for translating words and emphasis
Quoc Truong Do, Sakriani Sakti, and Satoshi Nakamura. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al. 2017 · 2017
Cited alongside, same era.
Sequence-to-sequence models can directly translate foreign speech
Ron J Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen. 2017 · 2017
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Later among the works it cites.
Libri-light: A benchmark for ASR with limited or no supervision
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al. 2020 · 2020
Later among the works it cites.
HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020 · 2020
Later among the works it cites.
fairseq S2T: Fast speech-to-text modeling with fairseq
Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, and Juan Pino. 2020b · 2020
Later among the works it cites.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matt Post. 2018 · 2018
Cited alongside, same era.
A comparative study on end-to-end speech to text translation
Parnia Bahar, Tobias Bieschke, and Hermann Ney. 2019 · 2019
Cited alongside, same era.
Leveraging weakly supervised data to improve end-to-end speech-to-text translation
Ye Jia, Melvin Johnson, Wolfgang Macherey, Ron J Weiss, Yuan Cao, Chung-Cheng Chiu, Naveen Ari, Stella Laurenzo, and Yonghui Wu. 2019a · 2019
Cited alongside, same era.
Direct speech-to-speech translation with a sequence-to-sequence model
Ye Jia, Ron J Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu. 2019b · 2019
Cited alongside, same era.
Neural speech synthesis with transformer network
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
SpecAugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le. 2019 · 2019
Cited alongside, same era.
Closest in time.
Translatotron 2: Robust direct speech-to-speech translation
Ye Jia, Michelle Tadmor Ramanovich, Tal Remez, and Roi Pomerantz. 2021 · 2021
Closest in time.
Transformer-based direct speech-to-speech translation with transcoder
Takatomo Kano, Sakriani Sakti, and Satoshi Nakamura. 2021 · 2021
Closest in time.
Generative spoken language modeling from raw audio
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, et al. 2021 · 2021
Closest in time.
Multilingual speech translation from efficient finetuning of pretrained models
Xian Li, Changhan Wang, Yun Tang, Chau Tran, Yuqing Tang, Juan Pino, Alexei Baevski, Alexis Conneau, and Michael Auli. 2021 · 2021
Closest in time.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. 2021 · 2021
Closest in time.
Changhan Wang, Morgane Rivière, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux. 2021 · 2021
Closest in time.
Self-training and pre-training are complementary for speech recognition
Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, and Michael Auli. 2021 · 2021
Closest in time.
SUPERB: Speech processing universal performance benchmark
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al. 2021 · 2021
Closest in time.