Fetching the paper…
Reading the bibliography…
Several of the latest GAN-based vocoders show remarkable achievements, outperforming autoregressive and flow-based competitors in both qualitative and quantitative measures while synthesizing orders of magnitude faster.
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,”
1993
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,”
2014
Earlier work this paper cites.
2016
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,”
2017
Earlier work this paper cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,”
2018
Earlier work this paper cites.
A. Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. Driessche, E. Lockhart, L. Cobo, F. Stimberg, N. Casagrande, D. Grewe, S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis, “Parallel wavenet: fast high-fidelity speech synthesis,”
2018
Earlier work this paper cites.
C. Lee, A. Toffy, G. Jung, and W. Han, “Conditional wavegan,”
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
2018
Earlier work this paper cites.
W. Ping, K. Peng, and J. Chen, “Clarinet: parallel wave generation in end-to-end text-to-speech,”
2019
Earlier work this paper cites.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: a flow-based generative network for speech synthesis,”
2019
Earlier work this paper cites.
S. Kim, S. Lee, J. Song, J. Kim, and S. Yoon, “Flowavenet : a generative flow for raw audio,”
2019
Earlier work this paper cites.
C. Donahue, J. McAuley, and M. Puckette, “Adversarial audio synthesis,”
2019
Cited alongside, same era.
K. Kumar, R. Kumar, T. Boissiere, L. Gestin, W. Teoh, J. Sotelo, A. Brebisson, Y. Bengio, and A. Courville, “Melgan: generative adversarial networks for conditional waveform synthesis,”
2019
Cited alongside, same era.
R. Yamamoto, E. Song, and J. Kim, “Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,”
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J. Yang, J. Lee, Y. Kim, H. Cho, and J. Kim, “Vocgan: a high-fidelity real-time vocoder with a hierarchically-nested adversarial network,”
2020
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “Diffwave: a versatile diffusion model for audio synthesis,”
2021
Closest in time.
2021
Closest in time.
J. Kong, “Official repository of hifi-gan,”
2021
Closest in time.
W. Teoh, “Official repository of melgan,”
2021
Closest in time.
T. Hayashi, “Unofficial repository of parallel wavegan,”
2021
Closest in time.
Rishikesh, “Unofficial repository of vocgan,”
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J. Pons, S. Pascual, G. Cengarle, and J. Serra, “Upsampling artifacts in neural audio synthesis,”
2020
Cited alongside, same era.
R. Child, “Very deep vaes generalize autoregressive models and can outperform them on images,”
2020
Cited alongside, same era.
2020
Cited alongside, same era.
N. Chen, Y. Zhang, H. zen, R. Weiss, M. Norouzi, and W. Chan, “Wavegrad: estimating gradients for waveform generation,”
2021
Cited alongside, same era.
2021
Closest in time.
K. Ito, “The lj speech dataset,”
2021
Closest in time.
G. Karch, “The tacotron2 repository,”
2021
Closest in time.
“Tacotron2 pytorch checkpoint (fp32),”
2021
Closest in time.
“Amazon mechanical turks,”
2021
Closest in time.