Fetching the paper…
Reading the bibliography…
In Mandarin text-to-speech (TTS) system, the front-end text processing module significantly influences the intelligibility and naturalness of synthesized speech.
“A compression-based algorithm for chinese word segmentation,”
William J Teahan, Yingying Wen, Rodger McNab, and Ian H Witten, · 2000
Earlier work this paper cites.
“An rnn-based algorithm to detect prosodic phrase for chinese tts,”
Zhiwei Ying and Xiaohua Shi, · 2001
Earlier work this paper cites.
“Conditional random fields: Probabilistic models for segmenting and labeling sequence data,”
John Lafferty, Andrew McCallum, and Fernando CN Pereira, · 2001
Earlier work this paper cites.
“Grapheme-to-phoneme conversion for chinese text-to-speech,”
Jun Xu, Guohong Fu, and Haizhou Li, · 2004
Earlier work this paper cites.
“Inequality maximum entropy classifier with character features for polyphone disambiguation in mandarin tts systems,”
Xinnian Mao, Yuan Dong, Jinyu Han, Dezhi Huang, and Haila Wang, · 2007
Earlier work this paper cites.
“Disambiguating effectively chinese polyphonic ambiguity based on unify approach,”
Feng-Long Huang, · 2008
Earlier work this paper cites.
“Automatic prosody prediction and detection with conditional random field (crf) models,”
Yao Qian, Zhizheng Wu, Xuezhe Ma, and Frank Soong, · 2010
Earlier work this paper cites.
“Distributed representations of words and phrases and their compositionality,”
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean, · 2013
Cited alongside, same era.
“Impact of automatic feature extraction in deep learning architecture,”
Fatma Shaheen, Brijesh Verma, and Md Asafuddoula, · 2016
Cited alongside, same era.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Transformer: A novel neural network architecture for language understanding,”
Jakob Uszkoreit, · 2017
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Later among the works it cites.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Fei Ren, Ye Jia, and Rif A Saurous, · 2018
Later among the works it cites.
“Implementing prosodic phrasing in chinese end-to-end speech synthesis,”
Yanfeng Lu, Minghui Dong, and Ying Chen, · 2019
Closest in time.
“A mandarin prosodic boundary prediction model based on multi-task learning,”
Huashan Pan, Xiulin Li, and Zhiqiang Huang, · 2019
Closest in time.
“Effective use of variational embedding capacity in expressive end-to-end speech synthesis,”
Eric Battenberg, Soroosh Mariooryad, Daisy Stanton, RJ Skerry-Ryan, Matt Shannon, David Kao, and Tom Bagby, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, and Koray Kavukcuoglu, · 2018
Cited alongside, same era.
“Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,”
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron J Weiss, Rob Clark, and Rif A Saurous, · 2018
Cited alongside, same era.
Closest in time.
“A hybrid text normalization system using multi-head self-attention for mandarin,”
Junhui Zhang, Junjie Pan, Xiang Yin, Shichao Liu, Yang Zhang, Yuxuan Wang, and Zejun Ma, · 2020
Closest in time.