Fetching the paper…
Reading the bibliography…
The majority of current Text-to-Speech (TTS) datasets, which are collections of individual utterances, contain few conversational aspects.
“The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017
Keith Ito and Linda Johnson, · 2017
Earlier work this paper cites.
“DailyDialog: A manually labelled multi-turn dialogue dataset,”
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu, · 2017
Earlier work this paper cites.
“Montreal Forced Aligner: trainable text-speech alignment using Kaldi,”
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Earlier work this paper cites.
“LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Cited alongside, same era.
“MELD: A multimodal multi-party dataset for emotion recognition in conversations,”
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea, · 2019
Cited alongside, same era.
“Conversational end-to-end tts for voice agents,”
Haohan Guo, Shaofei Zhang, Frank K. Soong, Lei He, and Lei Xie, · 2021
Cited alongside, same era.
“Adaspeech 3: Adaptive text to speech for spontaneous style,”
Yuzi Yan, Xu Tan, Bohan Li, Guangyan Zhang, Tao Qin, Sheng Zhao, Yuan Shen, Wei-Qiang Zhang, and Tie-Yan Liu, · 2021
Cited alongside, same era.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2021
Later among the works it cites.
“One tts alignment to rule them all,”
Rohan Badlani, Adrian Lancucki, Kevin J Shih, Rafael Valle, Wei Ping, and Bryan Catanzaro, · 2021
Later among the works it cites.
“STYLER: Style Factor Modeling with Rapidity and Robustness via Speech Decomposition for Expressive and Controllable Neural Text to Speech,”
Keon Lee, Kyumin Park, and Daeyoung Kim, · 2021
Later among the works it cites.
“Paratts: Learning linguistic and prosodic cross-sentence information in paragraph-based tts,”
Liumeng Xue, Frank K Soong, Shaofei Zhang, and Lei Xie, · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…