Fetching the paper…
Reading the bibliography…
Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST.
“Janus-iii: Speech-to-speech translation in multiple languages,”
Alon Lavie, Alex Waibel, Lori Levin, Michael Finke, Donna Gates, Marsal Gavalda, Torsten Zeppenfeld, and Puming Zhan, · 1997
Earlier work this paper cites.
“On the integration of speech recognition and statistical machine translation,”
Evgeny Matusov, Stephan Kanthak, and Hermann Ney, · 2005
Earlier work this paper cites.
“Europarl: A parallel corpus for statistical machine translation,”
Philipp Koehn, · 2005
Earlier work this paper cites.
“The atr multilingual speech-to-speech translation system,”
Satoshi Nakamura, Konstantin Markov, Hiromi Nakaiwa, Gen-ichiro Kikui, Hisashi Kawai, Takatoshi Jitsuhiro, J-S Zhang, Hirofumi Yamamoto, Eiichiro Sumita, and Seiichi Yamamoto, · 2006
Earlier work this paper cites.
“Listen and translate: A proof of concept for end-to-end speech-to-text translation,”
Alexandre Bérard, Olivier Pietquin, Christophe Servan, and Laurent Besacier, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“A call for clarity in reporting bleu scores,”
Matt Post, · 2018
Earlier work this paper cites.
“Direct speech-to-speech translation with a sequence-to-sequence model,”
Ye Jia, Ron J Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu, · 2019
Earlier work this paper cites.
“Speech-to-speech translation between untranscribed unknown languages,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2019
Earlier work this paper cites.
“Self-attention with structural position representations,”
Xing Wang, Zhaopeng Tu, Longyue Wang, and Shuming Shi, · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Multilingual speech translation with efficient finetuning of pretrained models,”
Xian Li, Changhan Wang, Yun Tang, Chau Tran, Yuqing Tang, Juan Pino, Alexei Baevski, Alexis Conneau, and Michael Auli, · 2020
Cited alongside, same era.
“Europarl-st: A multilingual corpus for speech translation of parliamentary debates,”
Javier Iranzo-Sánchez, Joan Albert Silvestre-Cerda, Javier Jorge, Nahuel Roselló, Adria Giménez, Albert Sanchis, Jorge Civera, and Alfons Juan, · 2020
Cited alongside, same era.
“Multilingual denoising pre-training for neural machine translation,”
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer, · 2020
Cited alongside, same era.
“Uwspeech: Speech to speech translation for unwritten languages,”
Chen Zhang, Xu Tan, Yi Ren, Tao Qin, Kejun Zhang, and Tie-Yan Liu, · 2021
Later among the works it cites.
“Textless speech-to-speech translation on real data,”
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Juan Pino, Jiatao Gu, and Wei-Ning Hsu, · 2021
Later among the works it cites.
Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino, Yossi Adi, Jiatao Gu, Wei-Ning Hsu, and Ann Lee, · 2022
Closest in time.
“Simple and effective unsupervised speech translation,”
Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang, Wei-Ning Hsu, Michael Auli, and Juan Pino, · 2022
Closest in time.
“Cvss corpus and massively multilingual speech-to-speech translation,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, et al., · 2021
Cited alongside, same era.
“VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,”
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, et al., · 2021
Cited alongside, same era.
“Translatotron 2: Robust direct speech-to-speech translation,”
Ye Jia, Michelle Tadmor Ramanovich, Tal Remez, and Roi Pomerantz, · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“Transformer-based direct speech-to-speech translation with transcoder,”
Takatomo Kano, Sakriani Sakti, and Satoshi Nakamura, · 2021
Cited alongside, same era.
“SpeechMatrix: A large-scale mined corpus of multilingual speech-to-speech translations,”
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswani, et al.,
Cited in the paper.
Ye Jia, Michelle Tadmor Ramanovich, Quan Wang, and Heiga Zen, · 2022
Closest in time.
“Leveraging pseudo-labeled data to improve direct speech-to-speech translation,”
Qianqian Dong, Fengpeng Yue, Tom Ko, Mingxuan Wang, Qibing Bai, and Yu Zhang, · 2022
Closest in time.
“Leveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translation,”
Ye Jia, Yifan Ding, Ankur Bapna, Colin Cherry, Yu Zhang, Alexis Conneau, and Nobuyuki Morioka, · 2022
Closest in time.
“mslam: Massively multilingual joint pre-training for speech and text,”
Ankur Bapna, Colin Cherry, Yu Zhang, Ye Jia, Melvin Johnson, Yong Cheng, Simran Khanuja, Jason Riesa, and Alexis Conneau, · 2022
Closest in time.
Ziqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu, Lirong Dai, Jinyu Li, and Furu Wei, · 2022
Closest in time.