Fetching the paper…
Reading the bibliography…
This paper proposes a new approach to duration modelling for statistical parametric speech synthesis in which a recurrent statistical model is trained to output a phone transition probability at each timestep (acoustic frame).
“The mean, median, mode inequality and skewness for a class of densities,”
H. L. MacGillivray, · 1981
Earlier work this paper cites.
“Review of text-to-speech conversion for English,”
Dennis H. Klatt, · 1987
Earlier work this paper cites.
“Syllable-level duration determination,”
W. Nick Campbell, · 1989
Earlier work this paper cites.
“A statistical model of duration control for speech synthesis,”
K. Huber, · 1990
Earlier work this paper cites.
“Mel-cepstral distance measure for objective speech quality assessment,”
Robert F. Kubichek, · 1993
Earlier work this paper cites.
“The mean, median, and mode of unimodal distributions: a characterization,”
Sanjib Basu and Anirban DasGupta, · 1997
Earlier work this paper cites.
“Speech parameter generation algorithms for HMM-based speech synthesis,”
Keiichi Tokuda, Takayoshi Yoshimura, Takashi Masuko, Takao Kobayashi, and Tadashi Kitamura, · 2000
Earlier work this paper cites.
“Joint prosody prediction and unit selection for concatenative speech synthesis,”
Ivan Bulyko and Mari Ostendorf, · 2001
Earlier work this paper cites.
“Hidden semi-markov model based speech synthesis,”
Heiga Zen, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, and Tadashi Kitamura, · 2004
Earlier work this paper cites.
“Sub-phonetic modeling for capturing pronunciation variations for conversational speech synthesis,”
Kishore Prahallad, Alan W. Black, and Ravishankhar Mosur, · 2006
Cited alongside, same era.
“The HMM-based speech synthesis system (HTS) version 2.0,”
Heiga Zen, Takashi Nose, Junichi Yamagishi, Shinji Sako, Takashi Masuko, Alan Black, and Keiichi Tokuda, · 2007
Cited alongside, same era.
“Statistical parametric speech synthesis,”
Heiga Zen, Keiichi Tokuda, and Alan W. Black, · 2009
Cited alongside, same era.
“An introduction to statistical parametric speech synthesis,”
Simon King, · 2011
Cited alongside, same era.
“Statistical parametric speech synthesis using deep neural networks,”
Heiga Zen, Andrew Senior, and Mike Schuster, · 2013
Cited alongside, same era.
“Measuring the perceptual effects of modelling assumptions in speech synthesis using stimuli constructed from repeated natural speech,”
“A study of speaker adaptation for DNN-based speech synthesis,”
Zhizheng Wu, Pawel Swietojanski, Christophe Veaux, Steve Renals, and Simon King, · 2015
Later among the works it cites.
“Sentence-level control vectors for deep neural network speech synthesis,”
Oliver Watts, Zhizheng Wu, and Simon King, · 2015
Later among the works it cites.
“The NST–GlottHMM entry to the Blizzard Challenge 2015,”
Oliver Watts, Srikanth Ronanki, Zhizheng Wu, Tuomo Raitio, and Antti Suni, · 2015
Later among the works it cites.
“Wavenet: A generative model for raw audio,”
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Closest in time.
“Fast, compact, and high quality LSTM-RNN based statistical parametric speech synthesizers for mobile devices,”
Heiga Zen, Yannis Agiomyrgiannakis, Niels Egberts, Fergus Henderson, and Przemysław Szczepaniak, · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gustav Eje Henter, Thomas Merritt, Matt Shannon, Catherine Mayo, and Simon King, · 2014
Cited alongside, same era.
“Prosody contour prediction with long short-term memory, bi-directional, deep recurrent neural networks,”
Raul Fernandez, Asaf Rendel, Bhuvana Ramabhadran, and Ron Hoory, · 2014
Cited alongside, same era.
“Deep neural networks employing multi-task learning and stacked bottleneck features for speech synthesis,”
Zhizheng Wu, Cassia Valentini-Botinhao, Oliver Watts, and Simon King, · 2015
Cited alongside, same era.
“Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis,”
Heiga Zen and Haşim Sak, · 2015
Cited alongside, same era.
Gustav Eje Henter, Srikanth Ronanki, Oliver Watts, Mirjam Wester, Zhizheng Wu, and Simon King, · 2016
Closest in time.
“A template-based approach for speech synthesis intonation generation using LSTMs,”
Srikanth Ronanki, Gustav Eje Henter, Zhizheng Wu, and Simon King, · 2016
Closest in time.
“The Blizzard Challenge 2016,”
Simon King and Vasilis Karaiskos, · 2016
Closest in time.
“Investigating gated recurrent networks for speech synthesis,”
Zhizheng Wu and Simon King, · 2016
Closest in time.