Fetching the paper…
Reading the bibliography…
Many speech synthesis datasets, especially those derived from audiobooks, naturally comprise sequences of utterances.
“Communication and prosody: Functional aspects of prosody,”
Julia Hirschberg, · 2002
Earlier work this paper cites.
“Turn-taking cues in task-oriented dialogue,”
Agustín Gravano and Julia Hirschberg, · 2011
Earlier work this paper cites.
“Prosody in context: a review,”
Jennifer Cole, · 2015
Earlier work this paper cites.
“Paragraph-based prosodic cues for speech synthesis applications,”
Mireia Farrús, Catherine Lai, and Johanna D Moore, · 2016
Earlier work this paper cites.
“Phrase break prediction for long-form reading tts: Exploiting text structure information.,”
Viacheslav Klimkov, Adam Nadolski, Alexis Moinet, Bartosz Putrycz, Roberto Barra-Chicote, Thomas Merritt, and Thomas Drugman, · 2017
Earlier work this paper cites.
“Neural machine translation with extended context,”
Jörg Tiedemann and Yves Scherrer, · 2017
Earlier work this paper cites.
“Discourse-based objectives for fast unsupervised sentence representation learning,”
Yacine Jernite, Samuel R Bowman, and David Sontag, · 2017
Cited alongside, same era.
“The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017
Keith Ito and Linda Johnson, · 2017
Cited alongside, same era.
“Paragraph prosodic patterns to enhance text-to-speech naturalness,”
Àlex Peiró-Lilja and Mireia Farrús, · 2018
Cited alongside, same era.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Fei Ren, Ye Jia, and Rif A Saurous, · 2018
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Cited alongside, same era.
“Evaluating long-form text-to-speech: Comparing the ratings of sentences and paragraphs,”
Rob Clark, Hanna Silen, Tom Kenter, and Ralph Leith, · 2019
Later among the works it cites.
Shubhi Tyagi, Marco Nicolis, Jonas Rohnke, Thomas Drugman, and Jaime Lorenzo-Trueba, · 2019
Later among the works it cites.
“Waveglow: A flow-based generative network for speech synthesis,”
ICASSP, · 2019
Later among the works it cites.
“Prosodic prominence and boundaries in sequence-to-sequence speech synthesis,”
Antti Suni, Sofoklis Kakouros, Martti Vainio, and Juraj Šimko, · 2020
Closest in time.
“Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,”
Rafael Valle, Jason Li, Ryan Prenger, and Bryan Catanzaro, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…