Fetching the paper…
Reading the bibliography…
Thanks to improvements in machine learning techniques, including deep learning, speech synthesis is becoming a machine learning task.
“Vocal quality factors: Analysis, synthesis, and perception,”
D. G. Childers and C. K. Lee, · 1991
Earlier work this paper cites.
“Continuous probabilistic transform for voice conversion,”
Y. Stylianou, O. Cappé, and E. Moulines, · 1998
Earlier work this paper cites.
“Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based F0 extraction: Possible role of a repetitive structure in sounds,”
H. Kawahara, I. Masuda-Katsuse, and A. D. Cheveigne, · 1999
Earlier work this paper cites.
“Julius — an open source real-time large vocabulary recognition engine,”
A. Lee, T. Kawahara, and K. Shikano, · 2001
Earlier work this paper cites.
“Analysis and recognition of whispered speech,”
T. Ito, K. Takeda, and F. Itakura, · 2005
Earlier work this paper cites.
“Whispery speech recognition using adapted articulatory features,”
S.-C. Jou, T. Schultz, and A. Waibel, · 2005
Earlier work this paper cites.
“Voice conversion based on maximum likelihood estimation of spectral parameter trajectory,”
T. Toda, A. W. Black, and K. Tokuda, · 2007
Earlier work this paper cites.
“List of daily-use kanjis
Governments of Japan Agency for Cultural Affairs, · 2010
Earlier work this paper cites.
“Whispered speech prosody modeling for TTS synthesis,”
V. A. Petrushin, L. I. Tsirulnik, and V. Makarova, · 2010
Cited alongside, same era.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
G. Hinton, L. Deng, D. Yu, G. Dahl, A. r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, and B. Kingsbury, · 2012
Cited alongside, same era.
“Statistical voice conversion techniques for body-conducted unvoiced speech enhancement,”
T. Toda, M. Nakagiri, and K. Shikano, · 2012
Cited alongside, same era.
“Factorized context modeling for Text-to-Speech synthesis,”
H. Lu and S. King, · 2013
Cited alongside, same era.
“Multiple-average-voice-based speech synthesis,”
P. Lanchantin, Mark J.F. Gales, S. King, and J. Yamagishi, · 2014
Cited alongside, same era.
“Sampling-based speech parameter generation using moment-matching network,”
S. Takamichi, K. Tomoki, and H. Saruwatari, · 2017
Later among the works it cites.
“JSUT corpus: free large-scale japanese speech corpus for end-to-end speech synthesis,”
R. Sonobe, S. Takamichi, and H. Saruwatari, · 2017
Later among the works it cites.
“Statistical parametric speech synthesis incorporating generative adversarial networks,”
Y. Saito, S. Takamichi, and H. Saruwatari, · 2018
Later among the works it cites.
“Hands on voice conversion,”
T. Toda, · 2018
Later among the works it cites.
“Multi-speaker sequence-to-sequence speech synthesis for data augmentation in acoustic-to-word speech recognition,”
S. Ueno, M. Mimura, S. Sakai, and Tatsuya Kawahara, · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, · 2016
Cited alongside, same era.
“WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,”
M. Morise, F. Yokomori, and K. Ozawa, · 2016
Cited alongside, same era.
“D4C, a band-aperiodicity estimator for high-quality speech synthesis,”
M. Morise, · 2016
Cited alongside, same era.
“JSUT: Japanese speech corpus of Saruwatari Lab, the University of Tokyo corpus,”
Cited in the paper.
“Voice-actress corpus,”
y_benjo and MagnesiumRibbon,
Cited in the paper.
“REAPER: Robust Epoch And Pitch EstimatoR,”
D. Talkin,
Cited in the paper.
“Speech signal processing toolkit (SPTK),”
Cited in the paper.
“Emotional voice conversion using dual supervised adversarial networks with continuous wavelet transform F0 features,”
Z. Luo, J. Chen, T. Takiguchi, and Y. Ariki, · 2019
Closest in time.
“DNN-based speaker embedding using subjective inter-speaker similarity for multi-speaker modeling in speech synthesis,”
Y. Saito, S. Takamichi, and H. Saruwatari, · 2019
Closest in time.