Fetching the paper…
Reading the bibliography…
To scale neural speech synthesis to various real-world languages, we present a multilingual end-to-end framework that maps byte inputs to spectrograms, thus allowing arbitrary input scripts.
Toward accurate dynamic time warping in linear time and space
Stan Salvador and Philip Chan · 2007
Earlier work this paper cites.
Multi-language multi-speaker acoustic modeling for LSTM-RNN based statistical parametric speech synthesis
Bo Li and Heiga Zen · 2016
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda B. Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2017
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz · 2017
Earlier work this paper cites.
An investigation of convolution attention based models for multilingual speech synthesis of indian languages
Pallavi Baljekar, Sai Krishna Rallabandi, and Alan W. Black · 2018
Earlier work this paper cites.
A unified phonological representation of south asian languages for multilingual text-to-speech
Isin Demirsahin, Martin Jansche, and Alexander Gutkin · 2018
Earlier work this paper cites.
Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ-Skerrv Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui Wu · 2018
Earlier work this paper cites.
End-to-end text-to-speech for low-resource languages by cross-lingual transfer learning
Yuan-Jui Chen, Tao Tu, Cheng-chieh Yeh, and Hung-yi Lee · 2019
Earlier work this paper cites.
Semi-supervised training for improving data efficiency in end-to-end speech synthesis
Yu-An Chung, Yuxuan Wang, Wei-Ning Hsu, Yu Zhang, and R. J. Skerry-Ryan · 2019
Earlier work this paper cites.
Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes
Bo Li, Yu Zhang, Tara N. Sainath, Yonghui Wu, and William Chan · 2019
Earlier work this paper cites.
Neural speech synthesis with transformer network
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu · 2019
Cited alongside, same era.
How language-neutral is multilingual BERT?
Jindrich Libovický, Rudolf Rosa, and Alexander Fraser · 2019
Cited alongside, same era.
Unsupervised polyglot text-to-speech
Eliya Nachmani and Lior Wolf · 2019
Cited alongside, same era.
Almost unsupervised text to speech and automatic speech recognition
Yi Ren, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Cited alongside, same era.
Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning
Yu Zhang, Ron J. Weiss, Heiga Zen, Yonghui Wu, Zhifeng Chen, R. J. Skerry-Ryan, Ye Jia, Andrew Rosenberg, and Bhuvana Ramabhadran · 2019
Cited alongside, same era.
Towards unsupervised speech recognition and synthesis with quantized speech representation learning
Alexander H. Liu, Tao Tu, Hung-yi Lee, and Lin-Shan Lee · 2020
Later among the works it cites.
Tone learning in low-resource bilingual TTS
Ruolan Liu, Xue Wen, Chunhui Lu, and Xiao Chen · 2020
Later among the works it cites.
One model, many languages: Meta-learning for multilingual text-to-speech
Tomás Nekvinda and Ondrej Dusek · 2020
Later among the works it cites.
Probing multilingual BERT for genetic and typological signals
Taraka Rama, Lisa Beinborn, and Steffen Eger · 2020
Later among the works it cites.
Phonological features for 0-shot multilingual speech synthesis
Marlene Staib, Tian Huey Teh, Alexandra Torresquintero, Devang S. Ram Mohan, Lorenzo Foglianti, Raphael Lenain, and Jiameng Gao · 2020
Later among the works it cites.
LRSpeech: Extremely low-resource speech synthesis and recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zexin Cai, Yaogen Yang, and Ming Li · 2020
Cited alongside, same era.
Multispeech: Multi-speaker text to speech with transformer
Mingjian Chen, Xu Tan, Yi Ren, Jin Xu, Hao Sun, Sheng Zhao, and Tao Qin · 2020
Cited alongside, same era.
Efficient neural speech synthesis for low-resource languages through multilingual modeling
Marcel de Korte, Jaebok Kim, and Esther Klabbers · 2020
Cited alongside, same era.
Open-source multi-speaker speech corpora for building gujarati, kannada, malayalam, marathi, tamil and telugu speech synthesis systems
Fei He, Shan-Hui Cathy Chu, Oddur Kjartansson, Clara Rivera, Anna Katanova, Alexander Gutkin, Isin Demirsahin, Cibu Johny, Martin Jansche, Supheakmungkol Sarin, and Knot Pipatsrisawat · 2020
Cited alongside, same era.
Jin Xu, Xu Tan, Yi Ren, Tao Qin, Jian Li, Sheng Zhao, and Tie-Yan Liu · 2020
Later among the works it cites.
Towards universal text-to-speech
Jingzhou Yang and Lei He · 2020
Later among the works it cites.
Investigation of learning abilities on linguistic features in sequence-to-sequence text-to-speech synthesis
Yusuke Yasuda, Xin Wang, and Junichi Yamagishi · 2020
Later among the works it cites.
Unsupervised learning for sequence-to-sequence text-to-speech for low-resource languages
Haitong Zhang and Yue Lin · 2020
Later among the works it cites.