Fetching the paper…
Reading the bibliography…
Recent advances in Text-to-Speech (TTS) have improved quality and naturalness to near-human capabilities when considering isolated sentences.
A. Aubin, A. Cervone, O. Watts, and S. King, “Improving speech synthesis with discourse relations,” in
1945
Earlier work this paper cites.
D. Cer, M. Diab, E. Agirre, I. Lopez-Gazpio, and L. Specia, “SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation,” in
2001
Earlier work this paper cites.
I. Recommendation, “Method for the subjective assessment of intermediate sound quality (mushra),”
2001
Earlier work this paper cites.
M. Wagner and D. G. Watson, “Experimental and theoretical advances in prosody: A review,”
2010
Earlier work this paper cites.
W. Coster and D. Kauchak, “Learning to simplify sentences using Wikipedia,” in
2011
Earlier work this paper cites.
M. Schuster and K. Nakajima, “Japanese and korean voice search,” in
2012
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in
2013
Earlier work this paper cites.
R. Dall, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, “Redefining the linguistic context feature set for hmm and dnn tts through position and parsing,” in
2016
Earlier work this paper cites.
Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. V. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan
2018
Cited alongside, same era.
K. Akuzawa, Y. Iwasawa, and Y. Matsuo, “Expressive speech synthesis via modeling expressions with variational autoencoder,” in
2018
Cited alongside, same era.
A. Köhn, T. Baumann, and O. Dörfler, “An empirical analysis of the correlation of syntax and prosody,” in
2018
Cited alongside, same era.
S. Stehwien, N. T. Vu, and A. Schweitzer, “Effects of word embeddings on neural network-based pitch accent detection,” in
2018
Cited alongside, same era.
D. Stanton, Y. Wang, and R. Skerry-Ryan, “Predicting expressive speaking style from text in end-to-end speech synthesis,” in
2018
Cited alongside, same era.
W.-N. Hsu, Y. Zhang, R. Weiss, H. Zen, Y. Wu, Y. Cao, and Y. Wang, “Hierarchical generative modeling for controllable speech synthesis,” in
2019
Closest in time.
H. Guo, F. K. Soong, L. He, and L. Xie, “Exploiting syntactic features in a parsed tree to improve end-to-end tts,” in
2019
Closest in time.
I. Tenney, P. Xia, B. Chen, A. Wang, A. Poliak, R. T. McCoy, N. Kim, B. V. Durme, S. R. Bowman, D. Das, and E. Pavlick, “What do you learn from context? probing for sentence structure in contextualized word representations,” in
2019
Closest in time.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Closest in time.
N. Prateek, M. Lajszczak, R. Barra-Chicote, T. Drugman, J. Lorenzo-Trueba, T. Merritt, S. Ronanki, and T. Wood, “In other news: A bi-style text-to-speech model for synthesizing newscaster voice with limited data,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Shen, Z. Lin, C. wei Huang, and A. Courville, “Neural language modeling by jointly learning syntax and lexicon,” in
2018
Cited alongside, same era.
Y. Shen, Z. Lin, A. P. Jacob, A. Sordoni, A. Courville, and Y. Bengio, “Straight to the tree: Constituency parsing with neural syntactic distance,” in
2018
Cited alongside, same era.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in
2019
Cited alongside, same era.
J. Lorenzo-Trueba, T. Drugman, J. Latorre, T. Merritt, B. Putrycz, R. Barra-Chicote, A. Moinet, and V. Aggarwal, “Towards achieving robust universal neural vocoding,” in
2019
Cited alongside, same era.
A. Govender and S. King, “Using pupillometry to measure the cognitive load of synthetic speech,”
Cited in the paper.
2019
Closest in time.
J. Latorre, J. Lachowicz, J. Lorenzo-Trueba, T. Merritt, T. Drugman, S. Ronanki, and V. Klimkov, “Effect of data reduction on sequence-to-sequence neural tts,” in
2019
Closest in time.
Z. Hodari, O. Watts, and S. King, “Using generative modelling to produce varied intonation for speech synthesis,” in
2019
Closest in time.
R. Clark, H. Silen, T. Kenter, and R. Leith, “Evaluating long-form text-to-speech: Comparing the ratings of sentences and paragraphs,” in
2019
Closest in time.