Fetching the paper…
Reading the bibliography…
Conversational speech synthesis (CSS) incorporates historical dialogue as supplementary information with the aim of generating speech that has dialogue-appropriate prosody.
“Learning a similarity metric discriminatively, with application to face verification,”
Sumit Chopra, Raia Hadsell, and Yann LeCun, · 2005
Earlier work this paper cites.
“Dynamic time warping,”
Meinard Müller, · 2007
Earlier work this paper cites.
“Visualizing data using t-sne.,”
Laurens Van der Maaten and Geoffrey Hinton, · 2008
Earlier work this paper cites.
“Speaker–listener neural coupling underlies successful communication,”
Greg J Stephens, Lauren J Silbert, and Uri Hasson, · 2010
Earlier work this paper cites.
“Enriching text-to-speech synthesis using automatic dialog act tags,”
Vivek Kumar Rangarajan Sridhar, Ann K. Syrdal, Alistair Conkie, and Srinivas Bangalore, · 2011
Earlier work this paper cites.
“Facenet: A unified embedding for face recognition and clustering,”
Florian Schroff, Dmitry Kalenichenko, and James Philbin, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ-Skerrv Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui Wu, · 2018
Earlier work this paper cites.
“Neural dynamics of semantic composition,”
Bingjiang Lyu, Hun S Choi, William D Marslen-Wilson, Alex Clarke, Billi Randall, and Lorraine K Tyler, · 2019
Cited alongside, same era.
“BERT: pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2021
Cited alongside, same era.
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
Jaehyeon Kim, Jungil Kong, and Juhee Son, · 2021
Cited alongside, same era.
“Comparing acoustic and textual representations of previous linguistic context for improving text-to-speech,”
Pilar Oplustil Gallegos, Johannah O’Mahony, and Simon King, · 2021
Cited alongside, same era.
Yuto Nishimura, Yuki Saito, Shinnosuke Takamichi, Kentaro Tachibana, and Hiroshi Saruwatari, · 2022
Later among the works it cites.
“Fctalker: Fine and coarse grained context modeling for expressive conversational speech synthesis,”
Yifan Hu, Rui Liu, Guanglai Gao, and Haizhou Li, · 2022
Later among the works it cites.
“Enhancing speaking styles in conversational text-to-speech synthesis with graph-based multi-modal context modeling,”
Jingbei Li, Yi Meng, Chenyi Li, Zhiyong Wu, Helen Meng, Chao Weng, and Dan Su, · 2022
Later among the works it cites.
“Inferring speaking styles from multi-modal conversational context by multi-scale relational graph convolutional networks,”
Jingbei Li, Yi Meng, Xixin Wu, Zhiyong Wu, Jia Jia, Helen Meng, Qiao Tian, Yuping Wang, and Yuxuan Wang, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haohan Guo, Shaofei Zhang, Frank K Soong, Lei He, and Lei Xie, · 2021
Cited alongside, same era.
Shun Lei, Yixuan Zhou, Liyang Chen, Jiankun Hu, Zhiyong Wu, Shiyin Kang, and Helen Meng, · 2022
Cited alongside, same era.
“Dailytalk: Spoken dialogue dataset for conversational text-to-speech,”
Keon Lee, Kyumin Park, and Daeyoung Kim, · 2022
Cited alongside, same era.
“Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech,”
Yi Ren, Ming Lei, Zhiying Huang, Shiliang Zhang, Qian Chen, Zhijie Yan, and Zhou Zhao, · 2022
Later among the works it cites.
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian, · 2023
Closest in time.
“M2-ctts: End-to-end multi-scale multi-modal conversational text-to-speech synthesis,”
Jinlong Xue, Yayue Deng, Fengping Wang, Ya Li, Yingming Gao, Jianhua Tao, Jianqing Sun, and Jiaen Liang, · 2023
Closest in time.