Fetching the paper…
Reading the bibliography…
Prosodic phrasing is crucial to the naturalness and intelligibility of end-to-end Text-to-Speech (TTS).
“Silent and non-silent pauses in three speech styles,”
Danielle Duez, · 1982
Earlier work this paper cites.
“An rnn-based prosodic information synthesizer for mandarin text-to-speech,”
Sin-Horng Chen, Shaw-Hwa Hwang, and Yih-Ru Wang, · 1998
Earlier work this paper cites.
“Prosodic phrasing is central to language comprehension,”
Lyn Frazier, Katy Carlson, and Charles Clifton, · 2006
Earlier work this paper cites.
“Pause prediction from lexical and syntax information,”
Venkatesh Keri, Sathish Chandra Pammi, and Kishore Prahallad, · 2007
Earlier work this paper cites.
“Iemocap: Interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan, · 2008
Earlier work this paper cites.
“Automatic prosody prediction and detection with conditional random field (crf) models,”
Yao Qian, Zhizheng Wu, Xuezhe Ma, and Frank Soong, · 2010
Earlier work this paper cites.
“Learning speaker-specific phrase breaks for text-to-speech systems,”
Kishore Prahallad, E Veera Raghavendra, and Alan W Black, · 2010
Earlier work this paper cites.
“Unsupervised continuous-valued word features for phrase-break prediction without a part-of-speech tagger.,”
Oliver Watts, Junichi Yamagishi, and Simon King, · 2011
Earlier work this paper cites.
“An investigation of recurrent neural network architectures using word embeddings for phrase break prediction.,”
Anandaswarup Vadapalli and Suryakanth V Gangashetty, · 2016
Earlier work this paper cites.
“Speaker specific phrase break modeling with conditional random fields for text-to-speech,”
Johannes A Louw and Avashlin Moodley, · 2016
Cited alongside, same era.
“Deep voice 3: Scaling text-to-speech with convolutional sequence learning,”
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller, · 2017
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Cited alongside, same era.
“Phrase break prediction for long-form reading tts: Exploiting text structure information,”
Viacheslav Klimkov, Adam Nadolski, Alexis Moinet, Bartosz Putrycz, Roberto Barra-Chicote, Tom Merritt, and Thomas Drugman, · 2018
Cited alongside, same era.
“Mongolian text-to-speech system based on deep neural network,”
Rui Liu, Feilong Bao, Guanglai Gao, and Yonghe Wang, · 2018
“Roberta: A robustly optimized bert pretraining approach,”
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov, · 2019
Later among the works it cites.
“Exploiting morphological and phonological features to improve prosodic phrasing for mongolian speech synthesis,”
Rui Liu, Berrak Sisman, Feilong Bao, Jichen Yang, Guanglai Gao, and Haizhou Li, · 2020
Later among the works it cites.
“Comparative analyses of bert, roberta, distilbert, and xlnet for text-based emotion recognition,”
Acheampong Francisca Adoma, Nunoo-Mensah Henry, and Wenyu Chen, · 2020
Later among the works it cites.
“Modelling representations in speech normalization of prosodic cues,”
Chen Si, Caicai Zhang, Puiyin Lau, Yike Yang, and Bei Li, · 2022
Later among the works it cites.
“Emotional voice conversion: Theory, databases and esd,”
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Improving mongolian phrase break prediction by using syllable and morphological embeddings with bilstm model.,”
Rui Liu, Feilong Bao, Guanglai Gao, Hui Zhang, and Yonghe Wang, · 2018
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Cited alongside, same era.
“Neural speech synthesis with transformer network,”
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu, · 2019
Cited alongside, same era.
“Prosodic structure prediction using deep self-attention neural network,”
Yao Du, Zhiyong Wu, Shiyin Kang, Dan Su, Dong Yu, and Helen Meng, · 2019
Cited alongside, same era.
“Simple matching coefficient,”
Cited in the paper.
“Adversarial multi-task learning for mandarin prosodic boundary prediction with multi-modal embeddings,”
Jiangyan Yi, Jianhua Tao, Ruibo Fu, Tao Wang, Chu Yuan Zhang, and Chenglong Wang, · 2023
Closest in time.
“Duration-aware pause insertion using pre-trained language model for multi-speaker text-to-speech,”
Dong Yang, Tomoki Koriyama, Yuki Saito, Takaaki Saeki, Detai Xin, and Hiroshi Saruwatari, · 2023
Closest in time.
“An investigation of speaker independent phrase break models in end-to-end tts systems,”
Anandaswarup Vadapalli, · 2023
Closest in time.
“Dailytalk: Spoken dialogue dataset for conversational text-to-speech,”
Keon Lee, Kyumin Park, and Daeyoung Kim, · 2023
Closest in time.