Fetching the paper…
Reading the bibliography…
We propose Universal MelGAN, a vocoder that synthesizes high-fidelity speech in multiple domains.
“Signal estimation from modified Short-Time Fourier Transform,”
D. Griffin and J. Lim, · 1984
Earlier work this paper cites.
“WaveNet vocoder with limited training data for voice conversion.,”
L.-J. Liu, Z.-H. Ling, Y. Jiang, M. Zhou, and L.-R. Dai, · 1987
Earlier work this paper cites.
“Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based f0 extraction: Possible role of a repetitive structure in sounds,”
H. Kawahara, I. Masuda-Katsuse, and A. De Cheveigne, · 1999
Earlier work this paper cites.
“The Blizzard challenge 2013,” Blizzard Challenge Workshop, 2013
S. King and V. Karaiskos, · 2013
Earlier work this paper cites.
“Generative adversarial nets,”
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, · 2014
Earlier work this paper cites.
“WORLD: A vocoder-based high-quality speech synthesis system for real-time applications,”
M. Morise, F. Yokomori, and K. Ozawa, · 2016
Earlier work this paper cites.
“WaveNet: A generative model for raw audio,”
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Earlier work this paper cites.
“Conditional image generation with PixelCnn decoders,”
A. Van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves, et al., · 2016
Earlier work this paper cites.
“The LJ speech dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017
K. Ito et al., · 2017
Earlier work this paper cites.
“Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,”
J. Shen, R. Pang, R. J Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, et al., · 2018
Cited alongside, same era.
“Towards achieving robust universal neural vocoding,”
J. Lorenzo-Trueba, T. Drugman, J. Latorre, T. Merritt, B. Putrycz, R. Barra-Chicote, A. Moinet, and V. Aggarwal, · 2018
Cited alongside, same era.
“Efficient neural audio synthesis,”
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, · 2018
Cited alongside, same era.
“Direct speech-to-speech translation with a sequence-to-sequence model,”
Y. Jia, R. J. Weiss, F. Biadsy, W. Macherey, M. Johnson, Z. Chen, and Y. Wu, · 2019
Cited alongside, same era.
“Towards robust neural vocoding for speech generation: A survey,”
“DurIAN: Duration informed attention network for multimodal synthesis,”
C. Yu, H. Lu, N. Hu, M. Yu, C. Weng, K. Xu, P. Liu, D. Tuo, S. Kang, G. Lei, et al., · 2019
Later among the works it cites.
“FastSpeech 2: Fast and high-quality end-to-end text-to-speech,”
Y. Ren, C. Hu, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2020
Closest in time.
“FastPitch: Parallel text-to-speech with pitch prediction,”
A. Lańcucki, · 2020
Closest in time.
D. Lim, W. Jang, G. O, H. Park, B. Kim, and J. Yoon, · 2020
Closest in time.
“Multi-band MelGAN: Faster waveform generation for high-quality text-to-speech,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Hsu, C. Wang, A. T. Liu, and H. Lee, · 2019
Cited alongside, same era.
“MelGAN: Generative adversarial networks for conditional waveform synthesis,”
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, · 2019
Cited alongside, same era.
“WaveGlow: A flow-based generative network for speech synthesis,”
R. Prenger, R. Valle, and B. Catanzaro, · 2019
Cited alongside, same era.
“LibriTTS: A corpus derived from LibriSpeech for text-to-speech,”
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, · 2019
Cited alongside, same era.
“CSS10: A collection of single speaker speech datasets for 10 languages,”
K. Park and T. Mulc, · 2019
Cited alongside, same era.
G. Yang, S. Yang, K. Liu, P. Fang, W. Chen, and L. Xie, · 2020
Closest in time.
“Speaker independence of neural vocoders and their effect on parametric resynthesis speech enhancement,”
S. Maiti and M. I Mandel, · 2020
Closest in time.
D. Paul, Y. Pantazis, and Y. Stylianou, · 2020
Closest in time.
“Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,”
R. Yamamoto, E. Song, and J.-M. Kim, · 2020
Closest in time.
J. Su, Z. Jin, and A. Finkelstein, · 2020
Closest in time.