Fetching the paper…
Reading the bibliography…
This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems.
M. R. Schroeder, B. S. Atal, and J. Hall, “Optimizing digital speech coders by exploiting masking properties of the human ear,” Journal of Acoust. Soc. of America , vol. 66, no. 6, pp. 1647–1652, 1979
1979
Earlier work this paper cites.
F. Soong and B. Juang, “Line spectrum pair (LSP) and speech data compression,” in Proc. ICASSP , 1984, pp. 37–40
1984
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proc. AISTATS , 2010, pp. 249–256
2010
Earlier work this paper cites.
H. Zen, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in Proc. ICASSP , 2013, pp. 7962–7966
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. P. Kingma, T. Salimans, R. Jozefowicz, X. Chen, I. Sutskever, and M. Welling, “Improved variational inference with inverse autoregressive flow,” in Proc. NIPS , 2016, pp. 4743–4751
2016
Earlier work this paper cites.
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” in Proc. NIPS , 2016, pp. 901–909
2016
Earlier work this paper cites.
A. Odena, V. Dumoulin, and C. Olah, “Deconvolution and checkerboard artifacts,” Distill , 2016. [Online]. Available: http://distill.pub/2016/deconv-checkerboard
2016
Earlier work this paper cites.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent WaveNet vocoder,” in Proc. INTERSPEECH , 2017, pp. 1118–1122
2017
Earlier work this paper cites.
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. Paul Smolley, “Least squares generative adversarial networks,” in Proc. ICCV , 2017, pp. 2794–2802
2017
Earlier work this paper cites.
B. Bollepalli, L. Juvela, and P. Alku, “Generative adversarial network-based glottal waveform model for statistical parametric speech synthesis,” in Proc. INTERSPEECH , 2017, pp. 3394–3398
2017
Earlier work this paper cites.
S. Pascual, A. Bonafonte, and J. Serrà, “SEGAN: Speech enhancement generative adversarial network,” in Proc. INTERSPEECH , 2017, pp. 3642–3646
2017
Cited alongside, same era.
E. Song, F. K. Soong, and H.-G. Kang, “Effective spectral and excitation modeling techniques for LSTM-RNN-based speech synthesis systems,” IEEE/ACM Trans. Audio, Speech, and Lang. Process. , vol. 25, no. 11, pp. 2152–2161, 2017
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in Proc. INTERSPEECH , 2017, pp. 4006–4010
2017
Cited alongside, same era.
2018
Cited alongside, same era.
S. Ö. Arık, H. Jun, and G. Diamos, “Fast spectrogram inversion using multi-head convolutional neural networks,” IEEE Signal Procees. Letters , vol. 26, no. 1, pp. 94–98, 2019
2019
Later among the works it cites.
T. Okamoto, T. Toda, Y. Shiga, and H. Kawai, “Real-time neural text-to-speech with sequence-to-sequence acoustic model and WaveGlow or single Gaussian WaveRNN vocoders,” in Proc. INTERSPEECH , 2019, pp. 1308–1312
2019
Later among the works it cites.
Q. Tian, X. Wan, and S. Liu, “Generative adversarial network based speaker adaptation for high fidelity WaveNet vocoder,” in Proc. SSW , 2019, pp. 19–23
2019
Later among the works it cites.
N. Li, S. Liu, Y. Liu, S. Zhao, M. Liu, and M. T. Zhou, “Neural speech synthesis with Transformer network,” in Proc. AAAI , 2019, pp. 6706–6713
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. C. Cobo, F. Stimberg et al. , “Parallel WaveNet: Fast high-fidelity speech synthesis,” in Proc. ICML , 2018, pp. 3915–3923
2018
Cited alongside, same era.
K. Tachibana, T. Toda, Y. Shiga, and H. Kawai, “An investigation of noise shaping with perceptual weighting for WaveNet-based speech generation,” in Proc. ICASSP , 2018, pp. 5664–5668
2018
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions,” in Proc. ICASSP , 2018, pp. 4779–4783
2018
Cited alongside, same era.
E. Song, K. Byun, and H.-G. Kang, “ExcitNet vocoder: A neural excitation model for parametric speech synthesis systems,” in Proc. EUSIPCO , 2019, pp. 1–5
2019
Cited alongside, same era.
W. Ping, K. Peng, and J. Chen, “ClariNet: Parallel wave generation in end-to-end text-to-speech,” in Proc. ICLR , 2019
2019
Cited alongside, same era.
R. Yamamoto, E. Song, and J.-M. Kim, “Probability density distillation with generative adversarial networks for high-quality parallel waveform generation,” in Proc. INTERSPEECH , 2019, pp. 699–703
2019
Cited alongside, same era.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, “MelGAN: Generative adversarial networks for conditional waveform synthesis,” in Proc. NeurIPS , 2019, pp. 14 881–14 892
2019
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
——, “Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in Proc. ICASSP , 2020, pp. 6199–6203
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Yang, J. Lee, Y. Kim, H.-Y. Cho, and I. Kim, “VocGAN: A high-fidelity real-time vocoder with a hierarchically-nested adversarial network,” in Proc. INTERSPEECH , 2020, pp. 200–204
2020
Later among the works it cites.
E. Song, M.-J. Hwang, R. Yamamoto, J.-S. Kim, O. Kwon, and J.-M. Kim, “Neural text-to-speech with a modeling-by-generation excitation vocoder,” in Proc. INTERSPEECH , 2020, pp. 3570–3574
2020
Later among the works it cites.