Fetching the paper…
Reading the bibliography…
The expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data.
“Expressive speech synthesis in mary tts using audiobook data and emotionml.,”
Marcela Charfuelan and Ingmar Steiner, · 2013
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“Neural discrete representation learning,”
Aaron Van Den Oord, Oriol Vinyals, et al., · 2017
Earlier work this paper cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Earlier work this paper cites.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ-Skerry Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A Saurous, · 2018
Earlier work this paper cites.
“Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,”
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron Weiss, Rob Clark, and Rif A Saurous, · 2018
Earlier work this paper cites.
“Predicting expressive speaking style from text in end-to-end speech synthesis,”
Daisy Stanton, Yuxuan Wang, and RJ Skerry-Ryan, · 2018
Earlier work this paper cites.
“Synpaflex-corpus: An expressive french audiobooks corpus dedicated to expressive speech synthesis,”
Aghilas Sini, Damien Lolive, Gaëlle Vidal, Marie Tahon, and Élisabeth Delais-Roussarie, · 2018
Earlier work this paper cites.
“Neural speech synthesis with transformer network,”
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu, · 2019
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova, · 2019
Cited alongside, same era.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2020
Cited alongside, same era.
“A simple framework for contrastive learning of visual representations,”
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton, · 2020
Cited alongside, same era.
“Improved prosody from learned f0 codebook representations for vq-vae speech waveform reconstruction,”
Yi Zhao, Haoyu Li, Cheng-I Lai, Jennifer Williams, Erica Cooper, and Junichi Yamagishi, · 2020
Cited alongside, same era.
“Supporting clustering with contrastive learning,”
Dejiao Zhang, Feng Nan, Xiaokai Wei, Shang-Wen Li, Henghui Zhu, Kathleen Mckeown, Ramesh Nallapati, Andrew O Arnold, and Bing Xiang, · 2021
Later among the works it cites.
“Learning disentangled phone and speaker representations in a semi-supervised vq-vae paradigm,”
Jennifer Williams, Yi Zhao, Erica Cooper, and Junichi Yamagishi, · 2021
Later among the works it cites.
“A character-level span-based model for mandarin prosodic structure prediction,”
Xueyuan Chen, Changhe Song, Yixuan Zhou, Zhiyong Wu, Changbin Chen, Zhongqin Wu, and Helen Meng, · 2022
Later among the works it cites.
“Hilvoice: Human-in-the-loop style selection for elder-facing speech synthesis,”
Xueyuan Chen, Qiaochu Huang, Xixin Wu, Zhiyong Wu, and Helen Meng, · 2022
Later among the works it cites.
“Towards expressive speaking style modelling with hierarchical context information for mandarin speech synthesis,”
Shun Lei, Yixuan Zhou, Liyang Chen, Zhiyong Wu, Shiyin Kang, and Helen Meng, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sven Buechel, Susanna Rücker, and Udo Hahn, · 2020
Cited alongside, same era.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Cited alongside, same era.
“Improving prosody modelling with cross-utterance bert embeddings for end-to-end speech synthesis,”
Guanghui Xu, Wei Song, Zhengchen Zhang, Chao Zhang, Xiaodong He, and Bowen Zhou, · 2021
Cited alongside, same era.
“Discourse-level prosody modeling with a variational autoencoder for non-autoregressive expressive speech synthesis,”
Ning-Qian Wu, Zhao-Ci Liu, and Zhen-Hua Ling, · 2022
Later among the works it cites.
“Unsupervised multi-scale expressive speaking style modeling with hierarchical context information for audiobook speech synthesis,”
Xueyuan Chen, Shun Lei, Zhiyong Wu, Dong Xu, Weifeng Zhao, and Helen Meng, · 2022
Later among the works it cites.
“Self-supervised context-aware style representation for expressive speech synthesis,”
Yihan Wu, Xi Wang, Shaofei Zhang, Lei He, Ruihua Song, and Jian-Yun Nie, · 2022
Later among the works it cites.