Fetching the paper…
Reading the bibliography…
Generative adversarial network (GAN)-based vocoders have been intensively studied because they can synthesize high-fidelity audio waveforms faster than real-time.
“Mel-cepstral distance measure for objective speech quality assessment,”
R. Kubichek, · 1993
Earlier work this paper cites.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,”
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, · 2001
Earlier work this paper cites.
“Generative adversarial nets,”
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, · 2014
Earlier work this paper cites.
“WaveNet: A generative model for raw audio,”
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Earlier work this paper cites.
“f-GAN: Training generative neural samplers using variational divergence minimization,”
S. Nowozin, B. Cseke, and R. Tomioka, · 2016
Earlier work this paper cites.
“Least squares generative adversarial networks,”
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. Paul Smolley, · 2017
Earlier work this paper cites.
J. H. Lim and J. C. Ye, · 2017
Earlier work this paper cites.
“Generative adversarial network-based glottal waveform model for statistical parametric speech synthesis,”
B. Bollepalli, L. Juvela, and P. Alku, · 2017
Earlier work this paper cites.
“SEGAN: Speech enhancement generative adversarial network,”
S. Pascual, A. Bonafonte, and J. Serrà, · 2017
Earlier work this paper cites.
M. Arjovsky, S. Chintala, and L. Bottou, · 2017
Earlier work this paper cites.
“The LJ Speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
K. Ito and L. Johnson, · 2017
Earlier work this paper cites.
“Natural TTS synthesis by conditioning WaveNet on MEL spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, · 2018
Earlier work this paper cites.
“Efficient neural audio synthesis,”
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, · 2018
Earlier work this paper cites.
“Parallel WaveNet: Fast high-fidelity speech synthesis,”
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. Cobo, F. Stimberg, N. Casagrande, D. Grewe, S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis, · 2018
Earlier work this paper cites.
“WaveGlow: A flow-based generative network for speech synthesis,”
R. Prenger, R. Valle, and B. Catanzaro, · 2018
Cited alongside, same era.
“MelGAN: Generative adversarial networks for conditional waveform synthesis,”
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, · 2019
Cited alongside, same era.
“ClariNet: Parallel wave generation in end-to-end text-to-speech,”
W. Ping, K. Peng, and J. Chen, · 2019
Cited alongside, same era.
“FloWaveNet : A generative flow for raw audio,”
S. Kim, S.-G. Lee, J. Song, J. Kim, and S. Yoon, · 2019
Cited alongside, same era.
“A style-based generator architecture for generative adversarial networks,”
T. Karras, S. Laine, and T. Aila, · 2019
Cited alongside, same era.
“LibriTTS: A corpus derived from LibriSpeech for text-to-speech,”
“FastSpeech 2: Fast and high-quality end-to-end text to speech,”
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2021
Later among the works it cites.
“WaveGrad: Estimating gradients for waveform generation,”
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, · 2021
Later among the works it cites.
“DiffWave: A versatile diffusion model for audio synthesis,”
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, · 2021
Later among the works it cites.
“PriorGrad: Improving conditional denoising diffusion models with data-dependent adaptive prior,”
S.-g. Lee, H. Kim, C. Shin, X. Tan, C. Liu, Q. Meng, T. Qin, W. Chen, S. Yoon, and T.-Y. Liu, · 2022
Later among the works it cites.
“SpecGrad: Diffusion probabilistic model based neural vocoder with adaptive noise spectral shaping,”
Y. Koizumi, H. Zen, K. Yatabe, N. Chen, and M. Bacchiani, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, · 2019
Cited alongside, same era.
“CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit version 0.92,” https://doi.org/10.7488/ds/2645 , 2019
J. Yamagishi, C. Veaux, and K. MacDonald, · 2019
Cited alongside, same era.
“Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms,”
K. Kilgou, M. Zuluaga, D. Roblek, and M. Shari, · 2019
Cited alongside, same era.
“Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,”
R. Yamamoto, E. Song, and J.-M. Kim, · 2020
Cited alongside, same era.
“HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,”
J. Kong, J. Kim, and J. Bae, · 2020
Cited alongside, same era.
“WaveFlow: A compact flow-based model for raw audio,”
W. Ping, K. Peng, K. Zhao, and Z. Song, · 2020
Cited alongside, same era.
“NanoFlow: Scalable normalizing flows with sublinear parameter complexity,”
S.-g. Lee, S. Kim, and S. Yoon, · 2020
Cited alongside, same era.
Z. Xiao, K. Kreis, and A. Vahdat, · 2022
Later among the works it cites.
“StyleGAN-XL: Scaling StyleGAN to large diverse datasets,”
A. Sauer, K. Schwarz, and A. Geiger, · 2022
Later among the works it cites.
“Chunked autoregressive GAN for conditional waveform synthesis,”
M. Morrison, R. Kumar, K. Kumar, P. Seetharaman, A. Courville, and Y. Bengio, · 2022
Later among the works it cites.
“Vocbench: A neural vocoder benchmark for speech synthesis,”
E. A. AlBadawy, A. Gibiansky, Q. He, J. Wu, M.-C. Chang, and S. Lyu, · 2022
Later among the works it cites.
“Hierarchical diffusion models for singing voice neural vocoder,”
N. Takahashi, M. K. Singh, and Y. Mitsufuji, · 2023
Closest in time.
“BigVGAN: A universal neural vocoder with large-scale training,”
S.-g. Lee, W. Ping, B. Ginsburg, B. Catanzaro, and S. Yoon, · 2023
Closest in time.
“SAN: Inducing metrizability of gan with discriminative normalized linear layer,”
Y. Takida, M. Imaizumi, T. Shibuya, C.-H. Lai, T. Uesaka, N. Murata, and Y. Mitsufuji, · 2023
Closest in time.
“Investigating range-equalizing bias in mean opinion score ratings of synthesized speech,”
E. Cooper and J. Yamagishi, · 2023
Closest in time.