Fetching the paper…
Reading the bibliography…
Conversational Text-to-Speech (TTS) aims to synthesis an utterance with the right linguistic and affective prosody in a conversational context.
“Dialog speech acts and prosody: Considerations for tts,”
Ann K Syrdal and Yeon-Jun Kim, · 2008
Earlier work this paper cites.
“Statistical parametric speech synthesis,”
Heiga Zen, Keiichi Tokuda, and Alan W Black, · 2009
Earlier work this paper cites.
“Conversational spontaneous speech synthesis using average voice model,”
Tomoki Koriyama, Takashi Nose, and Takao Kobayashi, · 2010
Earlier work this paper cites.
“On the use of extended context for hmm-based spontaneous conversational speech synthesis,”
Tomoki Koriyama, Takashi Nose, and Takao Kobayashi, · 2011
Earlier work this paper cites.
“Statistical parametric speech synthesis using deep neural networks,”
Heiga Ze, Andrew Senior, and Mike Schuster, · 2013
Earlier work this paper cites.
“Mean opinion score (mos) revisited: methods and applications, limitations and alternatives,”
Robert C Streijl, Stefan Winkler, and David S Hands, · 2016
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Earlier work this paper cites.
“Deep voice 2: Multi-speaker neural text-to-speech,”
Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Dailydialog: A manually labelled multi-turn dialogue dataset,”
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu, · 2017
Earlier work this paper cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Earlier work this paper cites.
“Deep voice 3: Scaling text-to-speech with convolutional sequence learning,”
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller, · 2018
Cited alongside, same era.
“Fastspeech: Fast, robust and controllable text to speech,”
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Cited alongside, same era.
“Joint multiple intent detection and slot labeling for goal-oriented dialog,”
Rashmi Gangadharaiah and Balakrishnan Narayanaswamy, · 2019
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“Multi-domain task-completion dialog challenge,”
Sungjin Lee, Hannes Schulz, Adam Atkinson, Jianfeng Gao, Kaheer Suleman, Layla El Asri, Mahmoud Adada, Minlie Huang, Shikhar Sharma, Wendy Tay, and Xiujun Li, · 2019
Cited alongside, same era.
“TOD-BERT: Pre-trained natural language understanding for task-oriented dialogue,”
Chien-Sheng Wu, Steven C.H. Hoi, Richard Socher, and Caiming Xiong, · 2020
Later among the works it cites.
“Conversational end-to-end tts for voice agents,”
Haohan Guo, Shaofei Zhang, Frank K Soong, Lei He, and Lei Xie, · 2021
Later among the works it cites.
“Controllable context-aware conversational speech synthesis,”
Jian Cong, Shan Yang, Na Hu, Guangzhi Li, Lei Xie, and Dan Su, · 2021
Later among the works it cites.
“Duplex conversation: Towards human-like interaction in spoken dialogue systems,”
Ting-En Lin, Yuchuan Wu, Fei Huang, Luo Si, Jian Sun, and Yongbin Li, · 2022
Closest in time.
Kentaro Mitsui, Tianyu Zhao, Kei Sawada, Yukiya Hono, Yoshihiko Nankaku, and Keiichi Tokuda, · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset,” 2019
Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan, · 2019
Cited alongside, same era.
“Conversational ai: Dialogue systems, conversational agents, and chatbots,”
Michael McTear, · 2020
Cited alongside, same era.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2020
Cited alongside, same era.
“Modeling long context for task-oriented dialogue state generation,”
Jun Quan and Deyi Xiong, · 2020
Cited alongside, same era.
“Ma-dst: Multi-attention-based scalable dialog state tracking,”
Adarsh Kumar, Peter Ku, Anuj Goyal, Angeliki Metallinou, and Dilek Hakkani-Tur, · 2020
Cited alongside, same era.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Cited alongside, same era.
“Hierarchical and multi-view dependency modelling network for conversational emotion recognition,”
Yu-Ping Ruan, Shu-Kai Zheng, Taihao Li, Fen Wang, and Guanxiong Pei, · 2022
Closest in time.
“Information-enhanced hierarchical self-attention network for multiturn dialog generation,”
Jiamin Wang, Xiao Sun, Qian Chen, and Meng Wang, · 2022
Closest in time.
“One tts alignment to rule them all,”
Rohan Badlani, Adrian Lańcucki, Kevin J Shih, Rafael Valle, Wei Ping, and Bryan Catanzaro, · 2022
Closest in time.
“Inferring speaking styles from multi-modal conversational context by multi-scale relational graph convolutional networks,”
Jingbei Li, Yi Meng, Xixin Wu, Zhiyong Wu, Jia Jia, Helen Meng, Tian Qiao, Yuping Wang, and Yuxuan Wang, · 2022
Closest in time.
“Dailytalk: Spoken dialogue dataset for conversational text-to-speech,”
Keon Lee, Kyumin Park, and Daeyoung Kim, · 2022
Closest in time.