Fetching the paper…
Reading the bibliography…
This paper proposes voicing-aware conditional discriminators for Parallel WaveGAN-based waveform synthesis systems.
“Statistical parametric speech synthesis using deep neural networks,”
H Zen, A Senior, and M Schuster, · 2013
Earlier work this paper cites.
“Generative adversarial nets,”
I Goodfellow, J Pouget-Abadie, M Mirza, B Xu, D Warde-Farley, S Ozair, A Courville, and Y Bengio, · 2014
Earlier work this paper cites.
“WaveNet: A generative model for raw audio,”
A van den Oord, S Dieleman, H Zen, K Simonyan, O Vinyals, A Graves, N Kalchbrenner, A Senior, and K Kavukcuoglu, · 2016
Earlier work this paper cites.
“Deconvolution and checkerboard artifacts,”
A Odena, V Dumoulin, and C Olah, · 2016
Earlier work this paper cites.
“Speaker-dependent WaveNet vocoder,”
A Tamamori, T Hayashi, K Kobayashi, K Takeda, and T Toda, · 2017
Earlier work this paper cites.
“An investigation of multi-speaker training for WaveNet vocoder,”
T Hayashi, A Tamamori, K Kobayashi, K Takeda, and T Toda, · 2017
Earlier work this paper cites.
“Least squares generative adversarial networks,”
X Mao, Q Li, H Xie, R. Y Lau, Z Wang, and S Paul Smolley, · 2017
Earlier work this paper cites.
“Effective spectral and excitation modeling techniques for LSTM-RNN-based speech synthesis systems,”
E Song, F. K Soong, and H.-G Kang, · 2017
Earlier work this paper cites.
“A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis,”
X Wang, J Lorenzo-Trueba, S Takaki, L Juvela, and J Yamagishi, · 2018
Earlier work this paper cites.
“Efficient neural audio synthesis,”
N Kalchbrenner, E Elsen, K Simonyan, S Noury, N Casagrande, E Lockhart, F Stimberg, A. v. d Oord, S Dieleman, and K Kavukcuoglu, · 2018
Earlier work this paper cites.
“Parallel WaveNet: Fast high-fidelity speech synthesis,”
A van den Oord, Y Li, I Babuschkin, K Simonyan, O Vinyals, K Kavukcuoglu, G van den Driessche, E Lockhart, L. C Cobo, F Stimberg, et al., · 2018
Cited alongside, same era.
“cGANs with projection discriminator,”
T Miyato and M Koyama, · 2018
Cited alongside, same era.
“ExcitNet vocoder: A neural excitation model for parametric speech synthesis systems,”
E Song, K Byun, and H.-G Kang, · 2019
Cited alongside, same era.
“ClariNet: Parallel wave generation in end-to-end text-to-speech,”
W Ping, K Peng, and J Chen, · 2019
Cited alongside, same era.
“WaveGlow: A flow-based generative network for speech synthesis,”
R Prenger, R Valle, and B Catanzaro, · 2019
Cited alongside, same era.
“FloWaveNet : A generative flow for raw audio,”
S Kim, S Lee, J Song, J Kim, and S Yoon, · 2019
Cited alongside, same era.
“Neural speech synthesis with Transformer network,”
N Li, S Liu, Y Liu, S Zhao, M Liu, and M. T Zhou, · 2019
Later among the works it cites.
“Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,”
R Yamamoto, E Song, and J.-M Kim, · 2020
Closest in time.
“VocGAN: A high-fidelity real-time vocoder with a hierarchically-nested adversarial network,”
J Yang, J Lee, Y Kim, H.-Y Cho, and I Kim, · 2020
Closest in time.
“High fidelity speech synthesis with adversarial networks,”
M Bińkowski, J Donahue, S Dieleman, A Clark, E Elsen, N Casagrande, L. C Cobo, and K Simonyan, · 2020
Closest in time.
“On the variance of the adaptive learning rate and beyond,”
L Liu, H Jiang, P He, W Chen, X Liu, J Gao, and J Han, · 2020
Closest in time.
“ESPnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Probability density distillation with generative adversarial networks for high-quality parallel waveform generation,”
R Yamamoto, E Song, and J.-M Kim, · 2019
Cited alongside, same era.
“MelGAN: Generative adversarial networks for conditional waveform synthesis,”
K Kumar, R Kumar, T de Boissiere, L Gestin, W. Z Teoh, J Sotelo, A de Brébisson, Y Bengio, and A. C Courville, · 2019
Cited alongside, same era.
“Fast spectrogram inversion using multi-head convolutional neural networks,”
S. Ö Arık, H Jun, and G Diamos, · 2019
Cited alongside, same era.
“Investigation of enhanced Tacotron text-to-speech synthesis systems with self-attention for pitch accent language,”
Y Yasuda, X Wang, S Takaki, and J Yamagishi, · 2019
Cited alongside, same era.
T Hayashi, R Yamamoto, K Inoue, T Yoshimura, S Watanabe, T Toda, K Takeda, Y Zhang, and X Tan, · 2020
Closest in time.
“Improved Parallel WaveGAN vocoder with perceptually weighted spectrogram loss,”
E Song, R Yamamoto, M.-J Hwang, J.-S Kim, O Kwon, and J.-M Kim, · 2021
Closest in time.
“FastSpeech 2: Fast and high-quality end-to-end text-to-speech,”
Y Ren, C Hu, T Qin, S Zhao, Z Zhao, and T.-Y Liu, · 2021
Closest in time.
“FastPitch: Parallel text-to-speech with pitch prediction,”
A Łańcucki, · 2021
Closest in time.