Fetching the paper…
Reading the bibliography…
Direct speech-to-speech translation (S2ST) translates speech from one language into another using a single model.
Direct speech-to-speech translation with a sequence-to-sequence model
Ye Jia, Ron J. Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu · 1951
Earlier work this paper cites.
Error bounds for convolutional codes and an asymptotically optimum decoding algorithm
A. Viterbi · 1967
Earlier work this paper cites.
A tutorial on hidden markov models and selected applications in speech recognition
L.R. Rabiner · 1989
Earlier work this paper cites.
Janus-iii: speech-to-speech translation in multiple languages
A. Lavie, A. Waibel, L. Levin, M. Finke, D. Gates, M. Gavalda, T. Zeppenfeld, and Puming Zhan · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
The atr multilingual speech-to-speech translation system
S. Nakamura, K. Markov, H. Nakaiwa, G. Kikui, H. Kawai, T. Jitsuhiro, J.-S. Zhang, H. Yamamoto, E. Sumita, and S. Yamamoto · 2005
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Christophe Veaux, Junichi Yamagishi, and Kirsten MacDonald · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and koray kavukcuoglu · 2017
Earlier work this paper cites.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor O. K. Li, and Richard Socher · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post · 2018
Earlier work this paper cites.
Generalized end-to-end loss for speaker verification
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno · 2018
Earlier work this paper cites.
End-to-end non-autoregressive neural machine translation with connectionist temporal classification
Jindřich Libovický and Jindřich Helcl · 2018
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Earlier work this paper cites.
Speech-to-speech translation between untranscribed unknown languages
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura · 2019
Earlier work this paper cites.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Earlier work this paper cites.
Conformer: Convolution-augmented Transformer for Speech Recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Cited alongside, same era.
Covost 2: A massively multilingual speech-to-text translation corpus, 2020
Changhan Wang, Anne Wu, and Juan Pino · 2020
Cited alongside, same era.
Understanding knowledge distillation in non-autoregressive machine translation
Chunting Zhou, Jiatao Gu, and Graham Neubig · 2020
Cited alongside, same era.
Latent-variable non-autoregressive neural machine translation with deterministic inference using a delta posterior
Raphael Shu, Jason Lee, Hideki Nakayama, and Kyunghyun Cho · 2020
Cited alongside, same era.
Non-autoregressive machine translation with latent alignments
Chitwan Saharia, William Chan, Saurabh Saxena, and Mohammad Norouzi · 2020
Cited alongside, same era.
Unity: Two-pass direct speech-to-speech translation with discrete units, 2022
Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang, Ann Lee, Shinji Watanabe, and Juan Pino · 2022
Later among the works it cites.
Enhanced direct speech-to-speech translation using self-supervised pre-training and data augmentation
Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino, Yossi Adi, Jiatao Gu, Wei-Ning Hsu, and Ann Lee · 2022
Later among the works it cites.
Directed acyclic transformer for non-autoregressive machine translation
Fei Huang, Hao Zhou, Yang Liu, Hang Li, and Minlie Huang · 2022
Later among the works it cites.
Viterbi decoding of directed acyclic transformer for non-autoregressive machine translation
Chenze Shao, Zhengrui Ma, and Yang Feng · 2022
Later among the works it cites.
Stemm: Self-learning with speech-text manifold mixup for speech translation
Qingkai Fang, Rong Ye, Lei Li, Yang Feng, and Mingxuan Wang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chenze Shao, Jinchao Zhang, Yang Feng, Fandong Meng, and Jie Zhou · 2020
Cited alongside, same era.
Non-autoregressive neural text-to-speech
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao · 2020
Cited alongside, same era.
Flow-tts: A non-autoregressive network for text to speech based on flow
Chenfeng Miao, Shuang Liang, Minchuan Chen, Jun Ma, Shaojun Wang, and Jing Xiao · 2020
Cited alongside, same era.
Glow-tts: A generative flow for text-to-speech via monotonic alignment search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon · 2020
Cited alongside, same era.
Glancing transformer for non-autoregressive neural machine translation
Lihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang, Lin Qiu, Weinan Zhang, Yong Yu, and Lei Li · 2021
Cited alongside, same era.
R-drop: Regularized dropout for neural networks
xiaobo liang, Lijun Wu, Juntao Li, Yue Wang, Qi Meng, Tao Qin, Wei Chen, Min Zhang, and Tie-Yan Liu · 2021
Cited alongside, same era.
Uwspeech: Speech to speech translation for unwritten languages
Chen Zhang, Xu Tan, Yi Ren, Tao Qin, Kejun Zhang, and Tie-Yan Liu · 2021
Cited alongside, same era.
Textless speech-to-speech translation on real data
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino, Jiatao Gu, and Wei-Ning Hsu · 2022
Later among the works it cites.
Speech-to-speech translation for A real-world unwritten language
Peng-Jen Chen, Kevin Tran, Yilin Yang, Jingfei Du, Justine Kao, Yu-An Chung, Paden Tomasello, Paul-Ambroise Duquenne, Holger Schwenk, Hongyu Gong, Hirofumi Inaguma, Sravya Popuri, Changhan Wang, Juan Miguel Pino, Wei-Ning Hsu, and Ann Lee · 2022
Later among the works it cites.
Leveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translation
Ye Jia, Yifan Ding, Ankur Bapna, Colin Cherry, Yu Zhang, Alexis Conneau, and Nobu Morioka · 2022
Later among the works it cites.
Leveraging pseudo-labeled data to improve direct speech-to-speech translation
Qianqian Dong, Fengpeng Yue, Tom Ko, Mingxuan Wang, Qibing Bai, and Yu Zhang · 2022
Later among the works it cites.
Improving speech-to-speech translation through unlabeled text
Xuan-Phi Nguyen, Sravya Popuri, Changhan Wang, Yun Tang, Ilia Kulikov, and Hongyu Gong · 2022
Later among the works it cites.
One reference is not enough: Diverse distillation with reference selection for non-autoregressive translation
Chenze Shao, Xuanfu Wu, and Yang Feng · 2022
Later among the works it cites.
latent-glat: Glancing at latent variables for parallel text generation
Yu Bao, Hao Zhou, Shujian Huang, Dongqi Wang, Lihua Qian, Xinyu Dai, Jiajun Chen, and Lei Li · 2022
Later among the works it cites.
Non-monotonic latent alignments for CTC-based non-autoregressive machine translation
Chenze Shao and Yang Feng · 2022
Later among the works it cites.
Naturalspeech: End-to-end text to speech synthesis with human-level quality
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, Frank K. Soong, Tao Qin, Sheng Zhao, and Tie-Yan Liu · 2022
Later among the works it cites.
Fastdiff: A fast conditional diffusion model for high-quality speech synthesis
Rongjie Huang, Max W. Y. Lam, Jun Wang, Dan Su, Dong Yu, Yi Ren, and Zhou Zhao · 2022
Later among the works it cites.
CMOT: Cross-modal mixup via optimal transport for speech translation
Yan Zhou, Qingkai Fang, and Yang Feng · 2023
Closest in time.
Joint pre-training with speech and bilingual text for direct speech to speech translation
Kun Wei, Long Zhou, Ziqiang Zhang, Liping Chen, Shujie Liu, Lei He, Jinyu Li, and Furu Wei · 2023
Closest in time.
Non-autoregressive streaming transformer for simultaneous translation
Zhengrui Ma, Shaolei Zhang, Shoutao Guo, Chenze Shao, Min Zhang, and Yang Feng · 2023
Closest in time.
Non-autoregressive machine translation with probabilistic context-free grammar
Shangtong Gui, Chenze Shao, Zhengrui Ma, Xishan Zhang, Yunji Chen, and Yang Feng · 2023
Closest in time.