Fetching the paper…
Reading the bibliography…
Nearly all Statistical Parametric Speech Synthesizers today use Mel Cepstral coefficients as the vocal tract parameterization of the speech signal.
S. Stevens, J. Volkmann, and E. Newman, “A scale for the measurement of the psychological magnitude pitch.”
1937
Earlier work this paper cites.
D. W. Robinson and R. S. Dadson, “A re-determination of the equal-loudness relations for pure tones,”
1956
Earlier work this paper cites.
F. Itakura, “Line spectrum representation of linear predictor coefficients of speech signals,”
1975
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams,
1988
Earlier work this paper cites.
K. Hornik, M. Stinchcombe, and H. White, “Multilayer feedforward networks are universal approximators,”
1989
Earlier work this paper cites.
E. Moulines and F. Charpentier, “Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones,”
1990
Earlier work this paper cites.
T. Dutoit and H. Leich, “MBR-PSOLA text-to-speech synthesis based on an MBE re-synthesis of the segments database,”
1993
Earlier work this paper cites.
R. F. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in
1993
Earlier work this paper cites.
K. Tokuda, T. Kobayashi, T. Masuko, and S. Imai, “Mel-generalized cepstral analysis-a unified approach to speech spectral estimation.” in
1994
Earlier work this paper cites.
K. Tokuda, T. Kobayashi, and S. Imai, “Speech parameter generation from HMM using dynamic features,” in
1995
Earlier work this paper cites.
T. Dutoit and B. Gosselin, “On the use of a hybrid harmonic/stochastic model for TTS synthesis-by-concatenation,”
1996
Cited alongside, same era.
Y. Stylianou, “Applying the harmonic plus noise model in concatenative speech synthesis,”
2001
Cited alongside, same era.
J. Kominek and A. W. Black, “The CMU arctic speech databases,” in
2004
Cited alongside, same era.
G. E. Hinton, S. Osindero, and Y.-W. Teh, “A fast learning algorithm for deep belief nets,”
2006
Cited alongside, same era.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,”
2006
Cited alongside, same era.
A. W. Black, “ClusterGen: a statistical parametric synthesizer using trajectory modeling.” in
2006
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol, “Extracting and composing robust features with denoising autoencoders,” in
2008
Later among the works it cites.
H. Zen, K. Tokuda, and A. Black, “Statistical parametric speech synthesis,”
2009
Later among the works it cites.
H. Larochelle, Y. Bengio, J. Louradour, and P. Lamblin, “Exploring strategies for training deep neural networks,”
2009
Later among the works it cites.
L. Deng, M. L. Seltzer, D. Yu, A. Acero, A.-R. Mohamed, and G. E. Hinton, “Binary coding of speech spectrograms using a deep auto-encoder.” in
2010
Later among the works it cites.
J. Bergstra, O. Breuleux, F. Bastien, P. Lamblin, R. Pascanu, G. Desjardins, J. Turian, D. Warde-Farley, and Y. Bengio, “Theano: a CPU and GPU math expression compiler,” in
2010
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
T. Tomoki and K. Tokuda, “A speech parameter generation algorithm considering global variance for HMM-based speech synthesis,”
2007
Cited alongside, same era.
Y. Bengio, P. Lamblin, D. Popovici, H. Larochelle
2007
Cited alongside, same era.
H. Kawahara, M. Morise, T. Takahashi, R. Nisimura, T. Irino, and H. Banno, “TANDEM-STRAIGHT: A temporally stable power spectral representation for periodic signals and applications to interference-free spectrum, f0, and aperiodicity estimation,” in
2008
Cited alongside, same era.
“Blizzard challenge,”
Cited in the paper.
Y. Bengio and O. Delalleau, “On the expressive power of deep architectures,” in
2011
Later among the works it cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath
2012
Later among the works it cites.
J. Gehring, Y. Miao, F. Metze, and A. Waibel, “Extracting deep bottleneck features using stacked auto-encoders,” in
2013
Later among the works it cites.
“Blizzard challenge 2014,”
2014
Closest in time.