Fetching the paper…
Reading the bibliography…
We propose a cross-lingual neural codec language model, VALL-E X, for cross-lingual speech synthesis.
Janus-iii: Speech-to-speech translation in multiple languages
Alon Lavie, Alex Waibel, Lori Levin, Michael Finke, Donna Gates, Marsal Gavalda, Torsten Zeppenfeld, and Puming Zhan · 1997
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
The atr multilingual speech-to-speech translation system
Satoshi Nakamura, Konstantin Markov, Hiromi Nakaiwa, Gen-ichiro Kikui, Hisashi Kawai, Takatoshi Jitsuhiro, J-S Zhang, Hirofumi Yamamoto, Eiichiro Sumita, and Seiichi Yamamoto · 2006
Earlier work this paper cites.
Diffwave: A versatile diffusion model for audio synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2009
Earlier work this paper cites.
The emime bilingual database
Mirjam Wester · 2010
Earlier work this paper cites.
Verbmobil: foundations of speech-to-speech translation
Wolfgang Wahlster · 2013
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Direct speech-to-speech translation with a sequence-to-sequence model
Ye Jia, Ron J Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu · 2019
Earlier work this paper cites.
Neural speech synthesis with transformer network
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu · 2019
Earlier work this paper cites.
Unsupervised polyglot text-to-speech
Eliya Nachmani and Lior Wolf · 2019
Earlier work this paper cites.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Earlier work this paper cites.
Yu Zhang, Ron J Weiss, Heiga Zen, Yonghui Wu, Zhifeng Chen, RJ Skerry-Ryan, Ye Jia, Andrew Rosenberg, and Bhuvana Ramabhadran · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Code-switched speech synthesis using bilingual phonetic posteriorgram with only monolingual corpora
Yuewen Cao, Songxiang Liu, Xixin Wu, Shiyin Kang, Peng Liu, Zhiyong Wu, Xunying Liu, Dan Su, Dong Yu, and Helen Meng · 2020
Earlier work this paper cites.
Libri-light: A benchmark for asr with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux · 2020
Earlier work this paper cites.
Multi-lingual multi-speaker text-to-speech synthesis for voice cloning with online speaker enrollment
Zhaoyu Liu and Brian Mak · 2020
Cited alongside, same era.
Emotional speech synthesis with rich and granularized control
Se-Yun Um, Sangshin Oh, Kyungguen Byun, Inseon Jang, ChungHyun Ahn, and Hong-Goo Kang · 2020
Cited alongside, same era.
Towards universal text-to-speech
Jingzhou Yang and Lei He · 2020
Cited alongside, same era.
Multilingual speech synthesis and cross-language voice cloning, December 3 2020
Yu Zhang, Ron J Weiss, Byungha Chun, Yonghui Wu, Zhifeng Chen, Russell John Wyatt Skerry-Ryan, Ye Jia, Andrew M Rosenberg, and Bhuvana Ramabhadran · 2020
Cited alongside, same era.
Shengkui Zhao, Trung Hieu Nguyen, Hao Wang, and Bin Ma · 2020
Cited alongside, same era.
Speechmatrix: A large-scale mined corpus of multilingual speech-to-speech translations
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswani, Changhan Wang, Juan Pino, Benoît Sagot, and Holger Schwenk · 2022
Later among the works it cites.
Cross-lingual text-to-speech with flow-based voice conversion for improved pronunciation
Nikolaos Ellinas, Georgios Vamvoukakis, Konstantinos Markopoulos, Georgia Maniati, Panos Kakoulidis, June Sig Sung, Inchul Hwang, Spyros Raptis, Aimilios Chalamandaris, and Pirros Tsiakoulis · 2022
Later among the works it cites.
Transpeech: Speech-to-speech translation with bilateral perturbation
Rongjie Huang, Zhou Zhao, Jinglin Liu, Huadai Liu, Yi Ren, Lichao Zhang, and Jinzheng He · 2022
Later among the works it cites.
Audiogen: Textually guided audio generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, Wei-Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, et al · 2021
Cited alongside, same era.
W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Cited alongside, same era.
Translatotron 2: Robust direct speech-to-speech translation
Ye Jia, Michelle Tadmor Ramanovich, Tal Remez, and Roi Pomerantz · 2021
Cited alongside, same era.
Deep learning based assessment of synthetic speech naturalness
Gabriel Mittag and Sebastian Möller · 2021
Cited alongside, same era.
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2021
Cited alongside, same era.
An empirical study on l2 accents of cross-lingual text-to-speech systems via vowel space
Jihwan Lee, Jae-Sung Bae, Seongkyu Mun, Heejin Choi, Joun Yeop Lee, Hoon-Young Cho, and Chanwoo Kim · 2022
Later among the works it cites.
Textless direct speech-to-speech translation with discrete speech representation
Xinjian Li, Ye Jia, and Chung-Cheng Chiu · 2022
Later among the works it cites.
Normalization of code-switched text for speech synthesis
Sreeram Manghat, Sreeja Manghat, and Tanja Schultz · 2022
Later among the works it cites.
Naturalspeech: End-to-end text to speech synthesis with human-level quality
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, et al · 2022
Later among the works it cites.
Joint pre-training with speech and bilingual text for direct speech to speech translation
Kun Wei, Long Zhou, Ziqiang Zhang, Liping Chen, Shujie Liu, Lei He, Jinyu Li, and Furu Wei · 2022
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu · 2022
Later among the works it cites.
Cross-lingual text-to-speech using multi-task learning and speaker classifier joint training
Jingzhou Yang and Lei He · 2022
Later among the works it cites.
Gigast: A 10,000-hour pseudo speech translation corpus
Rong Ye, Chengqi Zhao, Tom Ko, Chutong Meng, Tao Wang, Mingxuan Wang, and Jun Cao · 2022
Later among the works it cites.
Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, Di Wu, and Zhendong Peng · 2022
Later among the works it cites.
Cross-lingual multi-speaker speech synthesis with limited bilingual training data
Zexin Cai, Yaogen Yang, and Ming Li · 2023
Closest in time.
Neural codec language models are zero-shot text to speech synthesizers
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Closest in time.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2074
Closest in time.