Fetching the paper…
Reading the bibliography…
Recently, neural vocoders have been widely used in speech synthesis tasks, including text-to-speech and voice conversion.
D. Griffin and J. Lim, “Signal estimation from modified short-time fourier transform,”
1984
Earlier work this paper cites.
C. Recommendation, “Pulse code modulation (pcm) of voice frequencies,” in
1988
Earlier work this paper cites.
J. Kominek and A. W. Black, “The cmu arctic speech databases,” in
2004
Earlier work this paper cites.
H. Kawahara, “Straight, exploitation of the other aspect of vocoder: Perceptually isomorphic decomposition of speech sounds,”
2006
Earlier work this paper cites.
H. Banno, H. Hata, M. Morise, T. Takahashi, T. Irino, and H. Kawahara, “Implementation of realtime straight speech manipulation system: Report on its first implementation,”
2007
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “World: a vocoder-based high-quality speech synthesis system for real-time applications,”
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
K. Ito, “The LJ speech dataset,”
2017
Cited alongside, same era.
e. a. Christophe Veaux, Junichi Yamagishi, “CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” 2017
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Later among the works it cites.
2019
Closest in time.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “AutoVC: Zero-shot voice style transfer with only autoencoder loss,” in
2019
Closest in time.
2019
Closest in time.
J.-M. Valin and J. Skoglund, “Lpcnet: Improving neural speech synthesis through linear prediction,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo, “Stargan-vc: Non-parallel many-to-many voice conversion using star generative adversarial networks,” in
2018
Cited alongside, same era.
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, “Fftnet: A real-time speaker-dependent neural vocoder,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Closest in time.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in
2019
Closest in time.
2019
Closest in time.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in
2020
Closest in time.