Fetching the paper…
Reading the bibliography…
Recent speech technology research has seen a growing interest in using WaveNets as statistical vocoders, i.e., generating speech waveforms from acoustic features.
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,”
1992
Earlier work this paper cites.
D. Talkin, “A robust algorithm for pitch tracking (RAPT),”
1995
Earlier work this paper cites.
“Methods for Subjective Determination of Transmission Quality,” ITU-T SG12, Geneva, Switzerland, Recommendation P.800, Aug. 1996
1996
Earlier work this paper cites.
P. Alku, “Glottal inverse filtering analysis of human voice production – a review of estimation and parameterization methods of the glottal excitation and their applications. (invited article),”
2011
Earlier work this paper cites.
T. Raitio, A. Suni, J. Yamagishi, H. Pulakka, J. Nurminen, M. Vainio, and P. Alku, “HMM-based speech synthesis utilizing glottal inverse filtering,”
2011
Earlier work this paper cites.
M. Airaksinen, T. Raitio, B. Story, and P. Alku, “Quasi closed phase glottal inverse filtering analysis with weighted linear prediction,”
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
L. Juvela, B. Bollepalli, M. Airaksinen, and P. Alku, “High-pitched excitation generation for glottal vocoding in statistical parametric speech synthesis using a deep neural network,” in
2016
Earlier work this paper cites.
M. Airaksinen, B. Bollepalli, L. Juvela, Z. Wu, S. King, and P. Alku, “GlottDNN—a full-band glottal vocoder for statistical parametric speech synthesis,” in
2016
Cited alongside, same era.
2016
Cited alongside, same era.
S. O. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman, S. Sengupta, and M. Shoeybi, “Deep Voice: Real-time neural text-to-speech,” in
2017
Cited alongside, same era.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent WaveNet vocoder,” in
2017
Cited alongside, same era.
T. Hayashi, A. Tamamori, K. Kobayashi, K. Takeda, and T. Toda, “An investigation of multi-speaker training for WaveNet vocoder,” in
C. Valentini-Botinhao, “Noisy speech database for training speech enhancement algorithms and TTS models,” 2017. [Online]. Available:
2017
Later among the works it cites.
J. Shen, R. Pang, R. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” in
2018
Closest in time.
X. Wang, J. Lorenzo-Trueba, S. Takaki, L. Juvela, and J. Yamagishi, “A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis,” in
2018
Closest in time.
W. B. Kleijn, F. S. C. Lim, A. Luebs, J. Skoglund, F. Stimberg, Q. Wang, and T. C. Walters, “Wavenet based low rate speech coding,” in
2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
B. Bollepalli, L. Juvela, and P. Alku, “Generative adversarial network-based glottal waveform model for statistical parametric speech synthesis,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
M. Blaauw and J. Bonada, “A neural parametric singing synthesizer,” in
2017
Cited alongside, same era.
2018
Closest in time.
N. Adiga, V. Tsiaras, and Y. Stylianou, “On the use of WaveNet as a statistical vocoder,” in
2018
Closest in time.
CrowdFlower Inc., “Crowd-sourcing platform,” https://www.crowdflower.com/, accessed: 2018-03-22
2018
Closest in time.
“EF English proficiency index,” http://www.ef.com/epi/, accessed: 2018-03-22
2018
Closest in time.