Fetching the paper…
Reading the bibliography…
Neural vocoders based on the generative adversarial neural network (GAN) have been widely used due to their fast inference speed and lightweight networks while generating high-quality speech waveforms.
Communication in the presence of noise
Shannon, C. E. 1949 · 1949
Earlier work this paper cites.
A digital signal processing approach to interpolation
Schafer, R. W.; and Rabiner, L. R. 1973 · 1973
Earlier work this paper cites.
Mel-cepstral distance measure for objective speech quality assessment
Kubichek, R. 1993 · 1993
Earlier work this paper cites.
Fundamentals of speech recognition
Rabiner, L.; and Juang, B.-H. 1993 · 1993
Earlier work this paper cites.
Near-perfect-reconstruction pseudo-QMF banks
Nguyen, T. 1994 · 1994
Earlier work this paper cites.
Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs
Rix, A. W.; Beerends, J. G.; Hollier, M. P.; and Hekstra, A. P. 2001 · 2001
Earlier work this paper cites.
Generative Adversarial Nets
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Coupled Generative Adversarial Networks
Liu, M.-Y.; and Tuzel, O. 2016 · 2016
Earlier work this paper cites.
WORLD: a vocoder-based high-quality speech synthesis system for real-time applications
Masanori, M.; Yokomori, F.; and Ozawa, K. 2016 · 2016
Earlier work this paper cites.
Improved Techniques for Training GANs
Salimans, T.; Goodfellow, I.; Zaremba, W.; Cheung, V.; Radford, A.; and Chen, X. 2016 · 2016
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio
van den Oord, A.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A. W.; and Kavukcuoglu, K. 2016 · 2016
Earlier work this paper cites.
Neural Photo Editing with Introspective Adversarial Networks
Brock, A.; Lim, T.; Ritchie, J. M.; and Weston, N. 2017 · 2017
Earlier work this paper cites.
The LJ Speech Dataset
Ito, K.; and Johnson, L. 2017 · 2017
Earlier work this paper cites.
Least squares generative adversarial networks
Mao, X.; Li, Q.; Xie, H.; Lau, R. Y.; Wang, Z.; and Paul Smolley, S. 2017 · 2017
Earlier work this paper cites.
Subband WaveNet with overlapped single-sideband filterbanks
Okamoto, T.; Tachibana, K.; Toda, T.; Shiga, Y.; and Kawai, H. 2017 · 2017
Cited alongside, same era.
Tacotron: Towards End-to-End Speech Synthesis
Wang, Y.; Skerry-Ryan, R. J.; Stanton, D.; Wu, Y.; Weiss, R. J.; Jaitly, N.; Yang, Z.; Xiao, Y.; Chen, Z.; Bengio, S.; Le, Q. V.; Agiomyrgiannakis, Y.; Clark, R.; and Saurous, R. A. 2017 · 2017
Cited alongside, same era.
Wave-U-Net: A Multi-Scale Neural Network for End-to-End Audio Source Separation
Stoller, D.; Ewert, S.; and Dixon, S. 2018 · 2018
Cited alongside, same era.
Parallel WaveNet: Fast High-Fidelity Speech Synthesis
van den Oord, A.; Li, Y.; Babuschkin, I.; Simonyan, K.; Vinyals, O.; Kavukcuoglu, K.; van den Driessche, G.; Lockhart, E.; Cobo, L.; Stimberg, F.; Casagrande, N.; Grewe, D.; Noury, S.; Dieleman, S.; Elsen, E.; Kalchbrenner, N.; Zen, H.; Graves, A.; King, H.; Walters, T.; Belov, D.; and Hassabis, D. 2018 · 2018
Cited alongside, same era.
Adversarial Audio Synthesis
Donahue, C.; McAuley, J.; and Puckette, M. 2019 · 2019
Cited alongside, same era.
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Kong, J.; Kim, J.; and Bae, J. 2020 · 2020
Later among the works it cites.
Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Yamamoto, R.; Song, E.; and Kim, J.-M. 2020 · 2020
Later among the works it cites.
VocGAN: A High-Fidelity Real-Time Vocoder with a Hierarchically-Nested Adversarial Network
Yang, J.; Lee, J.; Kim, Y.; Cho, H.-Y.; and Kim, I. 2020 · 2020
Later among the works it cites.
UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation
Jang, W.; Lim, D.; Yoon, J.; Kim, B.; and Kim, J. 2021 · 2021
Later among the works it cites.
Alias-Free Generative Adversarial Networks
Karras, T.; Aittala, M.; Laine, S.; Härkönen, E.; Hellsten, J.; Lehtinen, J.; and Aila, T. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
FloWaveNet : A Generative Flow for Raw Audio
Kim, S.; Lee, S.-G.; Song, J.; Kim, J.; and Yoon, S. 2019 · 2019
Cited alongside, same era.
MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
Kumar, K.; Kumar, R.; de Boissiere, T.; Gestin, L.; Teoh, W. Z.; Sotelo, J.; de Brébisson, A.; Bengio, Y.; and Courville, A. C. 2019 · 2019
Cited alongside, same era.
Neural Speech Synthesis with Transformer Network
Li, N.; Liu, S.; Liu, Y.; Zhao, S.; and Liu, M. 2019 · 2019
Cited alongside, same era.
Towards Achieving Robust Universal Neural Vocoding
Lorenzo-Trueba, J.; Drugman, T.; Latorre, J.; Merritt, T.; Putrycz, B.; Barra-Chicote, R.; Moinet, A.; and Aggarwal, V. 2019 · 2019
Cited alongside, same era.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Cited alongside, same era.
WaveGlow: A flow-based generative network for speech synthesis
Prenger, R.; Valle, R.; and Catanzaro, B. 2019 · 2019
Cited alongside, same era.
FastSpeech: Fast, Robust and Controllable Text to Speech
Ren, Y.; Ruan, Y.; Tan, X.; Qin, T.; Zhao, S.; Zhao, Z.; and Liu, T. 2019 · 2019
Cited alongside, same era.
Fre-GAN: Adversarial Frequency-Consistent Audio Synthesis
Kim, J.-H.; Lee, S.-H.; Lee, J.-H.; and Lee, S.-W. 2021 · 2021
Later among the works it cites.
StyleMelGAN: An efficient high-fidelity adversarial vocoder with temporal adaptive normalization
Mustafa, A.; Pia, N.; and Fuchs, G. 2021 · 2021
Later among the works it cites.
Upsampling artifacts in neural audio synthesis
Pons, J.; Pascual, S.; Cengarle, G.; and Serrà, J. 2021 · 2021
Later among the works it cites.
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Ren, Y.; Hu, C.; Tan, X.; Qin, T.; Zhao, S.; Zhao, Z.; and Liu, T. 2021 · 2021
Later among the works it cites.
Multi-band MelGAN: Faster Waveform Generation for High-Quality Text-to-Speech
Yang, G.; Yang, S.; Liu, K.; Fang, P.; Chen, W.; and Xie, L. 2021 · 2021
Later among the works it cites.
NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates
Han, S.; and Lee, J. 2022 · 2022
Closest in time.
Chunked Autoregressive GAN for Conditional Waveform Synthesis
Morrison, M.; Kumar, R.; Kumar, K.; Seetharaman, P.; Courville, A.; and Bengio, Y. 2022 · 2022
Closest in time.
Daft-Exprt: Cross-Speaker Prosody Transfer on Any Text for Expressive Speech Synthesis
Zaïdi, J.; Seuté, H.; van Niekerk, B.; and Carbonneau, M.-A. 2022 · 2022
Closest in time.
DurIAN: Duration Informed Attention Network for Speech Synthesis
Yu, C.; Lu, H.; Hu, N.; Yu, M.; Weng, C.; Xu, K.; Liu, P.; Tuo, D.; Kang, S.; Lei, G.; Su, D.; and Yu, D. 2020 · 2031
Closest in time.