Fetching the paper…
Reading the bibliography…
Several recent studies have tested the use of transformer language model representations to infer prosodic features for text-to-speech synthesis (TTS).
Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,” Psychometrika , vol. 12, no. 2, pp. 153–157, 1947
1947
Earlier work this paper cites.
J. Cohen, “A coefficient of agreement for nominal scales,” Educational and Psychological Measurement , vol. 20, no. 1, pp. 37–46, 1960
1960
Earlier work this paper cites.
M. Rooth, “A theory of focus interpretation,” Natural Language Semantics , vol. 1, no. 1, pp. 75–116, 1992
1992
Earlier work this paper cites.
J. Hirschberg, “Pitch accent in context: Predicting intonational prominence from text,” Artificial Intelligence , vol. 63, no. 1–2, pp. 305–340, 1993
1993
Earlier work this paper cites.
A. Nenkova, J. Brenier, A. Kothari, S. Calhoun, L. Whitton, D. Beaver, and D. Jurafsky, “To memorize or to predict: Prominence labeling in conversational speech,” in Proc. of the Conf. of the North Am. Chapter of the Assoc. for Computational Linguistics , Rochester, New York, 2007
2007
Earlier work this paper cites.
L. Badino, J. S. Andersson, J. Yamagishi, and R. A. Clark, “Identification of contrast and its emphatic realization in HMM based speech synthesis,” in Proc. of Interspeech , Brighton, UK, 2009
2009
Earlier work this paper cites.
N. Jillings, D. Moffat, B. De Man, and J. D. Reiss, “Web Audio Evaluation Tool: A browser-based listening test environment,” in Proc. of the Sound and Music Computing Conf. , Maynooth, Ireland, 2015
2015
Earlier work this paper cites.
J. Reeve, “Chapterize,” https://github.com/JonathanReeve/chapterize , 2016
2016
Earlier work this paper cites.
A. Suni, J. Simko, D. Aalto, and M. Vainio, “Hierarchical representation and estimation of prosody using continuous wavelet transform,” Computer Speech and Language , vol. 45, pp. 123–136, 2017
2017
Earlier work this paper cites.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal Forced Aligner: Trainable text-speech alignment using Kaldi,” in Proc. of Interspeech , Stockholm, Sweden, 2017
2017
Earlier work this paper cites.
D. Stanton, Y. Wang, and R. Skerry-Ryan, “Predicting expressive speaking style from text in end-to-end speech synthesis,” in Proc. of IEEE Spoken Language Tech. Workshop (SLT) , Athens, Greece, 2018
2018
Cited alongside, same era.
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in Int. Conf. on Machine Learning . PMLR, 2018
2018
Cited alongside, same era.
A. Talman, A. Suni, H. Celikkanat, S. Kakouros, J. Tiedemann, and M. Vainio, “Predicting prosodic prominence from text with pre-trained contextualized word representations,” in Proc. of Nordic Conference on Computational Linguistics (NoDaLiDa) , Turku, Finland, 2019
2019
Cited alongside, same era.
T. Hayashi, S. Watanabe, T. Toda, K. Takeda, S. Toshniwal, and K. Livescu, “Pre-trained text embeddings for enhanced text-to-speech synthesis,” in Proc. of Interspeech , Graz, Austria, 2019
M. Honnibal, I. Montani, S. Van Landeghem, and A. Boyd, “spaCy: Industrial-strength Natural Language Processing in Python,” https://spacy.io/ , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in Proc. of ICASSP , Barcelona, Spain (virtual conference), 2020
2020
Later among the works it cites.
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,” in Proc. of Interspeech , Shanghai, China (virtual conference), 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. of ACL , Florence, Italy, 2019
2019
Cited alongside, same era.
T. Kenter, M. Sharma, and R. Clark, “Improving the prosody of RNN-based english text-to-speech synthesis by incorporating a BERT model,” in Proc. of Interspeech , Shanghai, China (virtual conference), 2020
2020
Cited alongside, same era.
Y. Xiao, L. He, H. Ming, and F. K. Soong, “Improving prosody with linguistic and BERT derived features in multi-speaker based Mandarin Chinese neural TTS,” in Proc. of IEEE ICASSP , Barcelona, Spain (virtual conference), 2020
2020
Cited alongside, same era.
T. Raitio, R. Rasipuram, and D. Castellani, “Controllable neural text-to-speech synthesis using intuitive prosodic features,” in Proc. of Interspeech , Shanghai, China (virtual conference), 2020
2020
Cited alongside, same era.
A. Suni, S. Kakouros, M. Vainio, and J. Šimko, “Prosodic prominence and boundaries in sequence-to-sequence speech synthesis,” in Proc. of ISCA Int. Conf. on Speech Prosody , Tokyo, Japan, 2020
2020
Cited alongside, same era.
Y. Zou, S. Liu, X. Yin, H. Lin, C. Wang, H. Zhang, and Z. Ma, “Fine-grained prosody modeling in neural speech synthesis using ToBI representation,” in Proc. of Interspeech , Brno, Czech Republik, 2021
2021
Later among the works it cites.
Z. Hodari, A. Moinet, S. Karlapati, J. Lorenzo-Trueba, T. Merritt, A. Joly, A. Abbas, P. Karanasou, and T. Drugman, “CAMP: a two-stage approach to modelling prosody in context,” in Proc. of IEEE ICASSP , Toronto, Canada, 2021
2021
Later among the works it cites.
S. Latif, I. Kim, I. Calapodescu, and L. Besacier, “Controlling prosody in end-to-end TTS: A case study on contrastive focus generation,” in Proc. of the Conf. on Computational Natural Language Learning (CoNLL) , Punta Cana, Dominican Republic, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.