Fetching the paper…
Reading the bibliography…
Most GAN(Generative Adversarial Network)-based approaches towards high-fidelity waveform generation heavily rely on discriminators to improve their performance.
D. Griffin and J. Lim, “Signal estimation from modified short-time fourier transform,” IEEE Transactions on acoustics, speech, and signal processing , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
2010
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention . Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “World: a vocoder-based high-quality speech synthesis system for real-time applications,” IEICE TRANSACTIONS on Information and Systems , vol. 99, no. 7, pp. 1877–1884, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
M. Morise et al. , “Harvest: A high-performance fundamental frequency estimator from speech signals.” in INTERSPEECH , 2017, pp. 2321–2325
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Cited alongside, same era.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 3617–3621
2019
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
C. Le Moine and N. Obin, “Att-HACK: An Expressive Speech Database with Social Attitudes,” in Speech Prosody , Tokyo, Japan, May 2020. [Online]. Available: https://hal.archives-ouvertes.fr/hal-02508362
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Sodimana, K. Pipatsrisawat, L. Ha, M. Jansche, O. Kjartansson, P. D. Silva, and S. Sarin, “A Step-by-Step Process for Building TTS Voices Using Open Source Data and Framework for Bangla, Javanese, Khmer, Nepali, Sinhala, and Sundanese,” in Proc. The 6th Intl. Workshop on Spoken Language Technologies for Under-Resourced Languages (SLTU) , Gurugram, India, Aug. 2018, pp. 66–70. [Online]. Available: http://dx.doi.org/10.21437/SLTU.2018-14
2018
Cited alongside, same era.
2019
Cited alongside, same era.
2021
Closest in time.