Y. Fan, Y. Qian, F.-L. Xie, and F. K. Soong, “TTS synthesis with bidirectional LSTM based recurrent neural networks,” in Proc. Interspeech , 2014, pp. 1964–1968
1968
Earlier work this paper cites.
A. P. Dempster, N. M. Laird, and D. B. Rubin, “Maximum likelihood from incomplete data via the EM algorithm,” J. Roy. Stat. Soc. B , vol. 39, no. 1, pp. 1–38, 1977
1977
Earlier work this paper cites.
S. Holm, “A simple sequentially rejective multiple test procedure,” Scand. J. Stat. , vol. 6, no. 2, pp. 65–70, 1979
1979
Earlier work this paper cites.
D. H. Klatt, “Review of text-to-speech conversion for English,” The Journal of the Acoustical Society of America , vol. 82, no. 3, pp. 737–793, 1987
1987
Earlier work this paper cites.
L. R. Rabiner, “A tutorial on hidden Markov models and selected applications in speech recognition,” Proc. IEEE , vol. 77, no. 2, pp. 257–286, 1989
1989
Earlier work this paper cites.
P. Dayan, G. E. Hinton, R. M. Neal, and R. S. Zemel, “The Helmholtz machine,” Neural Comput. , vol. 7, no. 5, pp. 889–904, 1995
1995
Earlier work this paper cites.
J. R. Quinlan, “Improved use of continuous attributes in C4.5,” J. Artif. Intel. Res. , vol. 4, pp. 77–90, 1996
1996
Earlier work this paper cites.
M. J. F. Gales, “Cluster adaptive training of hidden Markov models,” IEEE T. Speech Audi. P. , vol. 8, no. 4, pp. 417–428, 2000
2000
Earlier work this paper cites.
K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, “Speech parameter generation algorithms for HMM-based speech synthesis,” in Proc. ICASSP , 2000, pp. 1315–1318
2000
Earlier work this paper cites.
K. Fujinaga, M. Nakai, H. Shimodaira, and S. Sagayama, “Multiple-regression hidden Markov model,” in Proc. ICASSP , 2001, pp. 513–516
2001
Earlier work this paper cites.
T. Masuko, T. Kobayashi, and K. Miyanaga, “A style control technique for HMM-based speech synthesis,” in Proc. Interspeech , 2004, pp. 1437–1439
2004
Earlier work this paper cites.
J. Yamagishi, K. Onishi, T. Masuko, and T. Kobayashi, “Acoustic modeling of speaking styles and emotional expressions in HMM-based speech synthesis,” IEICE T. Inf. Syst. , vol. 88, no. 3, pp. 502–509, 2005
2005
Earlier work this paper cites.
T. Yoshimura, K. Tokuda, T. Masuko, T. Kobayashi, and T. Kitamura, “Incorporating a mixed excitation model and postfilter into HMM-based text-to-speech synthesis,” Syst. Comput. Jpn. , vol. 36, no. 12, pp. 43–50, 2005
2005
Earlier work this paper cites.
C. M. Bishop, Pattern Recognition and Machine Learning , 1st ed. New York, NY: Springer, 2006
2006
Earlier work this paper cites.
H. Kawahara, “STRAIGHT, exploitation of the other aspect of VOCODER: Perceptually isomorphic decomposition of speech sounds,” Acoust. Sci. Technol. , vol. 27, no. 6, pp. 349–353, 2006
2006
Earlier work this paper cites.
T. Nose, Y. Kato, and T. Kobayashi, “Style estimation of speech based on multiple regression hidden semi-Markov model,” in Proc. Interspeech , 2007, pp. 2285–2288
2007
Earlier work this paper cites.
H. Zen, T. Nose, J. Yamagishi, S. Sako, T. Masuko, A. W. Black, and K. Tokuda, “The HMM-based speech synthesis system (HTS) version 2.0,” in Proc. SSW , 2007, pp. 294–299
2007
Earlier work this paper cites.
A. Camacho and J. G. Harris, “A sawtooth waveform inspired pitch estimator for speech and music,” The Journal of the Acoustical Society of America , vol. 124, no. 3, pp. 1638–1652, 2008
2008
Earlier work this paper cites.
L. van der Maaten and G. Hinton, “Visualizing data using t-SNE,” J. Mach. Learn. Res. , vol. 9, no. Nov, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval . Cambridge University Press, 2008
2008
Earlier work this paper cites.
H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,” Speech Commun. , vol. 51, no. 11, pp. 1039–1064, 2009
2009
Earlier work this paper cites.
D. Povey, L. Burget, M. Agarwal, P. Akyazi, K. Feng, A. Ghoshal, O. Glembek, N. K. Goel, M. Karafiát, A. Rastrow, R. C. Rose, P. Schwarz, and S. Thomas, “Subspace Gaussian mixture models for speech recognition,” in Proc. ICASSP , 2010, pp. 4330–4333
2010
Earlier work this paper cites.
R. Barra-Chicote, J. Yamagishi, S. King, J. M. Montero, and J. Macias-Guarasa, “Analysis of statistical parametric and unit selection speech synthesis systems applied to emotional speech,” Speech Commun. , vol. 52, no. 5, pp. 394–404, 2010
2010
Earlier work this paper cites.
D. Erro, E. Navas, I. Herndez, and I. Saratxaga, “Emotion conversion based on prosodic unit selection,” IEEE T. Audio Speech , vol. 18, no. 5, pp. 974–983, 2010
2010
Earlier work this paper cites.
K. Oura, S. Sako, and K. Tokuda, “Japanese text-to-speech synthesis system: Open JTalk,” in Proc. ASJ Spring , 2010, pp. 343–344
2010
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proc. AISTATS , 2010, pp. 249–256
2010
Earlier work this paper cites.
S. King, “An introduction to statistical parametric speech synthesis,” Sadhana , vol. 36, no. 5, pp. 837–852, 2011
2011
Earlier work this paper cites.
L. Chen, M. J. F. Gales, V. Wan, J. Latorre, and M. Akamine, “Exploring rich expressive information from audiobook data using cluster adaptive training,” in Proc. Interspeech , 2012, pp. 959–962
2012
Earlier work this paper cites.
O. Abdel-Hamid and H. Jiang, “Fast speaker adaptation of hybrid NN/HMM model for speech recognition based on discriminative learning of speaker code,” in Proc. ICASSP , 2013, pp. 7942–7946
2013
Earlier work this paper cites.
Z.-H. Ling, K. Richmond, and J. Yamagishi, “Articulatory control of HMM-based parametric speech synthesis using feature-space-switched multiple regression,” IEEE T. Audio Speech , vol. 21, no. 1, pp. 207–219, 2013
2013
Earlier work this paper cites.
T. Nose and T. Kobayashi, “An intuitive style control technique in HMM-based expressive speech synthesis using subjective style intensity and multiple-regression global variance model,” Speech Commun. , vol. 55, no. 2, pp. 347–357, 2013
2013
Earlier work this paper cites.
Y. Bengio, N. Léonard, and A. Courville, “Estimating or propagating gradients through stochastic neurons for conditional computation,” arXiv preprint arXiv:1308.3432 , 2013
Original
2013
Earlier work this paper cites.
G. E. Henter, T. Merritt, M. Shannon, C. Mayo, and S. King, “Measuring the perceptual effects of modelling assumptions in speech synthesis using stimuli constructed from repeated natural speech,” in Proc. Interspeech , 2014, pp. 1504–1508
2014
Earlier work this paper cites.
S. Xue, O. Abdel-Hamid, H. Jiang, L.-R. Dai, and Q. Liu, “Fast adaptation of deep neural network based on discriminant codes for speech recognition,” IEEE/ACM T. Audio Speech , vol. 22, no. 12, pp. 1713–1725, 2014
2014
Earlier work this paper cites.