Fetching the paper…
Reading the bibliography…
This paper describes the Microsoft end-to-end neural text to speech (TTS) system: DelightfulTTS for Blizzard Challenge 2021.
A. J. Hunt and A. W. Black, “Unit selection in a concatenative speech synthesis system using a large speech database,” in 1996 IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings , vol. 1. IEEE, 1996, pp. 373–376
1996
Earlier work this paper cites.
C. Benoît, M. Grice, and V. Hazan, “The sus test: A method for the assessment of text-to-speech synthesis intelligibility using semantically unpredictable sentences,” Speech communication , vol. 18, no. 4, pp. 381–392, 1996
1996
Earlier work this paper cites.
K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, “Speech parameter generation algorithms for hmm-based speech synthesis,” in 2000 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No. 00CH37100) , vol. 3. IEEE, 2000, pp. 1315–1318
2000
Earlier work this paper cites.
K. Bali, P. P. Talukdar, N. S. Krishna, and A. Ramakrishnan, “Tools for the development of a hindi speech synthesis system,” in Fifth ISCA Workshop on Speech Synthesis , 2004
2004
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing , vol. 13, no. 4, pp. 600–612, 2004
2004
Earlier work this paper cites.
A. W. Black and K. Tokuda, “The blizzard challenge-2005: Evaluating corpus-based speech synthesis on common datasets,” in Ninth European Conference on Speech Communication and Technology , 2005
2005
Earlier work this paper cites.
H. Ze, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in 2013 ieee international conference on acoustics, speech and signal processing . IEEE, 2013, pp. 7962–7966
2013
Earlier work this paper cites.
Y. Fan, Y. Qian, F.-L. Xie, and F. K. Soong, “Tts synthesis with bidirectional lstm based recurrent neural networks,” in Fifteenth annual conference of the international speech communication association , 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
V. Aubanel, M. L. G. Lecumberri, and M. Cooke, “The sharvard corpus: A phonemically-balanced spanish sentence resource for audiology,” International journal of audiology , vol. 53, no. 9, pp. 633–638, 2014
2014
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Cited alongside, same era.
D. Stanton, Y. Wang, and R. Skerry-Ryan, “Predicting expressive speaking style from text in end-to-end speech synthesis,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 595–602
2018
2020
Later among the works it cites.
N. Li, Y. Liu, Y. Wu, S. Liu, S. Zhao, and M. Liu, “Robutrans: A robust transformer-based text-to-speech model,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 05, 2020, pp. 8228–8235
2020
Later among the works it cites.
2020
Later among the works it cites.
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, A. Rosenberg, B. Ramabhadran, and Y. Wu, “Generating diverse and natural text-to-speech samples using a quantized fine-grained vae and autoregressive prosody prior,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6699–6703
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in International Conference on Machine Learning . PMLR, 2018, pp. 5180–5189
2018
Cited alongside, same era.
2019
Cited alongside, same era.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural speech synthesis with transformer network,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 6706–6713
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Later among the works it cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Later among the works it cites.
Microsoft, “Azure neural tts upgraded with hifinet, achieving higher audio fidelity and faster synthesis speed,” p. 0, Nov 2020, hiFiNet. [Online]. Available: https://techcommunity.microsoft.com/t5/azure-ai/azure-neural-tts-upgraded-with-hifinet-achieving-higher-audio/ba-p/1847860
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Closest in time.
2021
Closest in time.
I. Elias, H. Zen, J. Shen, Y. Zhang, Y. Jia, R. J. Weiss, and Y. Wu, “Parallel tacotron: Non-autoregressive and controllable tts,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5709–5713
2021
Closest in time.