Fetching the paper…
Reading the bibliography…
We propose Parallel WaveGAN, a distillation-free, fast, and small-footprint waveform generation method using a generative adversarial network.
I. Daubechies, “The wavelet transform, time-frequency localization and signal analysis,” IEEE trans. on information theory , vol. 36, no. 5, pp. 961–1005, 1990
1990
Earlier work this paper cites.
H. Zen, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in Proc. ICASSP , 2013, pp. 7962–7966
2013
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Proc. NIPS , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. P. Kingma, T. Salimans, R. Jozefowicz, X. Chen, I. Sutskever, and M. Welling, “Improved variational inference with inverse autoregressive flow,” in Proc. NIPS , 2016, pp. 4743–4751
2016
Earlier work this paper cites.
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” in Proc. NIPS , 2016, pp. 901–909
2016
Earlier work this paper cites.
A. Odena, V. Dumoulin, and C. Olah, “Deconvolution and checkerboard artifacts,” Distill , 2016. [Online]. Available: http://distill.pub/2016/deconv-checkerboard
2016
Earlier work this paper cites.
E. Song, F. K. Soong, and H.-G. Kang, “Effective spectral and excitation modeling techniques for LSTM-RNN-based speech synthesis systems,” IEEE/ACM Trans. Audio, Speech, and Lang. Process. , vol. 25, no. 11, pp. 2152–2161, 2017
2017
Earlier work this paper cites.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent WaveNet vocoder,” in Proc. INTERSPEECH , 2017, pp. 1118–1122
2017
Earlier work this paper cites.
T. Hayashi, A. Tamamori, K. Kobayashi, K. Takeda, and T. Toda, “An investigation of multi-speaker training for WaveNet vocoder,” in Proc. ASRU , 2017, pp. 712–718
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NIPS , 2017, pp. 5998–6008
2017
Cited alongside, same era.
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. Paul Smolley, “Least squares generative adversarial networks,” in Proc. ICCV , 2017, pp. 2794–2802
2017
Cited alongside, same era.
B. Bollepalli, L. Juvela, and P. Alku, “Generative adversarial network-based glottal waveform model for statistical parametric speech synthesis,” in Proc. INTERSPEECH , 2017, pp. 3394–3398
2017
Cited alongside, same era.
S. Pascual, A. Bonafonte, and J. Serrà, “SEGAN: Speech enhancement generative adversarial network,” in Proc. INTERSPEECH , 2017, pp. 3642–3646
2017
Cited alongside, same era.
R. Yamamoto, E. Song, and J.-M. Kim, “Probability density distillation with generative adversarial networks for high-quality parallel waveform generation,” in Proc. INTERSPEECH , 2019, pp. 699–703
2019
Closest in time.
C. Donahue, J. McAuley, and M. Puckette, “Adversarial audio synthesis,” in Proc. ICLR , 2019
2019
Closest in time.
N. Li, S. Liu, Y. Liu, S. Zhao, M. Liu, and M. T. Zhou, “Neural speech synthesis with Transformer network,” in Proc. AAAI , 2019, pp. 6706–6713
2019
Closest in time.
2019
Closest in time.
L. Juvela, B. Bollepalli, J. Yamagishi, and P. Alku, “GELP: GAN-excited linear prediction for speech synthesis from mel-spectrogram,” in Proc. INTERSPEECH , 2019, pp. 694–698
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
X. Wang, J. Lorenzo-Trueba, S. Takaki, L. Juvela, and J. Yamagishi, “A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis,” in Proc. ICASSP , 2018, pp. 4804–4808
2018
Cited alongside, same era.
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. C. Cobo, F. Stimberg et al. , “Parallel WaveNet: Fast high-fidelity speech synthesis,” in Proc. ICML , 2018, pp. 3915–3923
2018
Cited alongside, same era.
2018
Cited alongside, same era.
E. Song, K. Byun, and H.-G. Kang, “Excitnet vocoder: A neural excitation model for parametric speech synthesis systems,” in Proc. EUSIPCO , 2019, pp. 1179–1183
2019
Cited alongside, same era.
W. Ping, K. Peng, and J. Chen, “ClariNet: Parallel wave generation in end-to-end text-to-speech,” in Proc. ICLR , 2019
2019
Cited alongside, same era.
2019
Closest in time.
Q. Tian, X. Wan, and S. Liu, “Generative adversarial network based speaker adaptation for high fidelity WaveNet vocoder,” in Proc. SSW , 2019, pp. 19–23
2019
Closest in time.
S. Ö. Arık, H. Jun, and G. Diamos, “Fast spectrogram inversion using multi-head convolutional neural networks,” IEEE Signal Procees. Letters , vol. 26, no. 1, pp. 94–98, 2019
2019
Closest in time.
X. Wang, S. Takaki, and J. Yamagishi, “Neural source-filter-based waveform model for statistical parametric speech synthesis,” in Proc. ICASSP , 2019, pp. 5916–5920
2019
Closest in time.
2019
Closest in time.
Y. Yasuda, X. Wang, S. Takaki, and J. Yamagishi, “Investigation of enhanced Tacotron text-to-speech synthesis systems with self-attention for pitch accent language,” in Proc. ICASSP , 2019, pp. 6905–6909
2019
Closest in time.