Fetching the paper…
Reading the bibliography…
This paper proposes speaker-adaptive neural vocoders for parametric text-to-speech (TTS) systems.
S. Furui, “Speaker-independent isolated word recognition using dynamic features of speech spectrum,” IEEE Trans. Acoust., Speech Signal Process. , vol. 34, no. 1, pp. 52–59, 1986
1986
Earlier work this paper cites.
R. J. Williams and J. Peng, “An efficient gradient-based algorithm for on-line training of recurrent network trajectories,” Neural computat. , vol. 2, no. 4, pp. 490–501, 1990
1990
Earlier work this paper cites.
K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, “Speech parameter generation algorithms for HMM-based speech synthesis,” in Proc. ICASSP , 2000, pp. 1315–1318
2000
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proc. AISTATS , 2010, pp. 249–256
2010
Earlier work this paper cites.
H. Zen, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in Proc. ICASSP , 2013, pp. 7962–7966
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. Fan, Y. Qian, F. K. Soong, and L. He, “Multi-speaker modeling and speaker adaptation for DNN-based TTS synthesis,” in Proc. ICASSP , 2015, pp. 4475–4479
2015
Earlier work this paper cites.
A. Van Den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves et al. , “Conditional image generation with PixelCNN decoders,” in Proc. NIPS , 2016, pp. 4790–4798
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
L. Juvela, V. Tsiaras, B. Bollepalli, M. Airaksinen, J. Yamagishi, and P. Alku, “Speaker-independent raw waveform model for glottal excitation,” in Proc. INTERSPEECH , 2018, pp. 2012–2016
2016
Cited alongside, same era.
S. Pascual and A. Bonafonte, “Multi-output RNN-LSTM for multiple speaker speech synthesis and adaptation,” in Proc. EUSIPCO , 2016, pp. 2325–2329
2016
Cited alongside, same era.
2016
Cited alongside, same era.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent WaveNet vocoder,” in Proc. INTERSPEECH , 2017, pp. 1118–1122
2017
Cited alongside, same era.
T. Hayashi, A. Tamamori, K. Kobayashi, K. Takeda, and T. Toda, “An investigation of multi-speaker training for wavenet vocoder,” in Proc. ASRU , 2017, pp. 712–718
N. Adiga, V. Tsiaras, and Y. Stylianou, “On the use of WaveNet as a statistical vocoder,” in Proc. ICASSP , 2018, pp. 5674–5678
2018
Closest in time.
T. Yoshimura, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, “Mel-cepstrum-based quantization noise shaping applied to neural-network-based speech waveform synthesis,” IEEE/ACM Trans. Audio, Speech, and Lang. Process. , vol. 26, no. 7, pp. 1173–1180, 2018
2018
Closest in time.
2018
Closest in time.
K. Tachibana, T. Toda, Y. Shiga, and H. Kawai, “An investigation of noise shaping with perceptual weighting for WaveNet-based speech generation,” in Proc. ICASSP , 2018, pp. 5664–5668
2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Y.-J. Hu, C. Ding, L.-J. Liu, Z.-H. Ling, and L.-R. Dai, “The USTC system for blizzard challenge 2017,” in Proc. Blizzard Challenge Workshop , 2017
2017
Cited alongside, same era.
E. Song, F. K. Soong, and H.-G. Kang, “Effective spectral and excitation modeling techniques for LSTM-RNN-based speech synthesis systems,” IEEE/ACM Trans. Audio, Speech, and Lang. Process. , vol. 25, no. 11, pp. 2152–2161, 2017
2017
Cited alongside, same era.
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio, “SampleRNN: An unconditional end-to-end neural audio generation model,” in Proc. ICLR , 2017
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan et al. , “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” in Proc. ICASSP , 2018, pp. 4779–4783
2018
Cited alongside, same era.
X. Wang, J. Lorenzo-Trueba, S. Takaki, L. Juvela, and J. Yamagishi, “A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis,” in Proc. ICASSP , 2018, pp. 4804–4808
2018
Closest in time.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in Proc. ICML , 2018, pp. 2410–2419
2018
Closest in time.
E. Song, K. Byun, and H.-G. Kang, “ExcitNet vocoder: A neural excitation model for parametric speech synthesis systems,” in Proc. EUSIPCO , 2019, pp. 1179–1183
2019
Closest in time.
J.-M. Valin and J. Skoglund, “LPCnet: Improving neural speech synthesis through linear prediction,” in Proc. ICASSP , 2019, pp. 5891–5895
2019
Closest in time.
R. Prenger, R. Valle, and B. Catanzaro, “WaveGlow: A flow-based generative network for speech synthesis,” in Proc. ICASSP , 2019, pp. 3617–3621
2019
Closest in time.