Fetching the paper…
Reading the bibliography…
In this paper, we present AISHELL-3, a large-scale and high-fidelity multi-speaker Mandarin speech corpus which could be used to train multi-speaker Text-to-Speech (TTS) systems.
“Carnegie-mellon university pronouncing dictionary for american english,”
CMU Pronouncing Dictionary, · 1998
Earlier work this paper cites.
Text-to-speech synthesis
Paul Taylor, · 2009
Earlier work this paper cites.
“Improved prediction of japanese word accent sandhi using crf,”
Nobuaki Minematsu, Shumpei Kobayashi, Shinya Shimizu, and Keikichi Hirose, · 2012
Earlier work this paper cites.
“WaveNet: A generative model for raw audio,”
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Earlier work this paper cites.
“Superseded-CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,”
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al., · 2016
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Earlier work this paper cites.
“Deep voice 2: Multi-speaker neural text-to-speech,”
Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou, · 2017
Earlier work this paper cites.
“Zoneout: Regularizing rnns by randomly preserving hidden activations,”
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Aaron Courville, and Chris Pal, · 2017
Earlier work this paper cites.
“Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Earlier work this paper cites.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, Fei Ren, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, Yonghui Wu, et al., · 2018
Cited alongside, same era.
“Neural speech synthesis with transformer network,”
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu, · 2019
Cited alongside, same era.
“Waveglow: A flow-based generative network for speech synthesis,”
Ryan Prenger, Rafael Valle, and Bryan Catanzaro, · 2019
Cited alongside, same era.
“Melgan: Generative adversarial networks for conditional waveform synthesis,”
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville, · 2019
Cited alongside, same era.
“A mandarin prosodic boundary prediction model based on multi-task learning.,”
Huashan Pan, Xiulin Li, and Zhiqiang Huang, · 2019
Cited alongside, same era.
“Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,”
Rafael Valle, Jason Li, Ryan Prenger, and Bryan Catanzaro, · 2020
Closest in time.
“Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis,”
Guangzhi Sun, Yu Zhang, Ron J Weiss, Yuan Cao, Heiga Zen, and Yonghui Wu, · 2020
Closest in time.
“Generating diverse and natural text-to-speech samples using a quantized fine-grained vae and autoregressive prosody prior,”
Guangzhi Sun, Yu Zhang, Ron J Weiss, Yuan Cao, Heiga Zen, Andrew Rosenberg, Bhuvana Ramabhadran, and Yonghui Wu, · 2020
Closest in time.
“Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,”
Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang, Xin Wang, Nanxin Chen, and Junichi Yamagishi, · 2020
Closest in time.
“From Speaker Verification to Multispeaker Speech Synthesis, Deep Transfer with Feedback Constraint,”
Zexin Cai, Chuxiong Zhang, and Ming Li, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Polyphone Disambiguation for Mandarin Chinese Using Conditional Neural Network with Multi-Level Embedding Features,”
Zexin Cai, Yaogen Yang, Chuxiong Zhang, Xiaoyi Qin, and Ming Li, · 2019
Cited alongside, same era.
“A prosodic mandarin text-to-speech system based on tacotron,”
Chuxiong Zhang, Sheng Zhang, and Haibing Zhong, · 2019
Cited alongside, same era.
“Forward–backward decoding sequence for regularizing end-to-end tts,”
Yibin Zheng, Jianhua Tao, Zhengqi Wen, and Jiangyan Yi, · 2019
Cited alongside, same era.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ-Skerry Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A. Saurous,
Cited in the paper.
“A feature-enriched neural model for joint chinese word segmentation and part-of-speech tagging,”
Xinchi Chen, Xipeng Qiu, and Xuanjing Huang,
Cited in the paper.
“ https://openslr.org/62/
aidatatang_200zh,
Cited in the paper.
“ https://openslr.com/68/
Magic Data Technology Co., Ltd.,
Cited in the paper.
“Mandarin prosody boundary prediction based on sequence-to-sequence model,”
Yajing Yan, Jiaolong Jiang, and Hongwu Yang, · 2020
Closest in time.
“On-the-fly data loader and utterance-level aggregation for speaker and language recognition,”
Weicheng Cai, Jinkun Chen, Jun Zhang, and Ming Li, · 2020
Closest in time.
“Location-relative attention mechanisms for robust long-form speech synthesis,”
Eric Battenberg, RJ Skerry-Ryan, Soroosh Mariooryad, Daisy Stanton, David Kao, Matt Shannon, and Tom Bagby, · 2020
Closest in time.