Fetching the paper…
Reading the bibliography…
We propose an end-to-end empathetic dialogue speech synthesis (DSS) model that considers both the linguistic and prosodic contexts of dialogue history.
K. K. Liu and R. W. Picard, “Embedded empathy in continuous interactive health assessment,” in Proc. CHI Workshop on HCI Challenges in Health Assessment , Portland, U.S.A., Apr. 2005
2005
Earlier work this paper cites.
C. Regenbogen, D. A. Schneider, A. Finkelmeyer, N. Kohn, B. Derntl, T. Kellermann, R. E. Gur, F. Schneider, and U. Habel, “The differential contribution of facial expressions, prosody, and speech content to empathy,” Cognition & Emotion , vol. 26, no. 6, pp. 995–1014, 2012
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
N. Komiya, Professional Counselor’s Lesson on Listening Skills for Each Situation . Natsumesha CO., LTD, 2015 (in Japanese)
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,” IEICE Transactions on Information and Systems , vol. E99-D, no. 7, pp. 1877–1884, 2016
2016
Earlier work this paper cites.
M. Morise, “D4C, a band-aperiodicity estimator for high-quality speech synthesis,” Speech Communication , vol. 84, pp. 57–65, Nov. 2016
2016
Earlier work this paper cites.
H. Chen, X. Liu, D. Yin, and J. Tang, “A survey on dialogue systems: Recent advances and new frontiers,” ACM SIGKDD Explorations Newsletter , vol. 19, no. 2, pp. 25–35, 2017
2017
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R.-J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R.-A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in Proc. INTERSPEECH , Stockholm, Sweden, Aug. 2017, pp. 4006–4010
2017
Earlier work this paper cites.
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, “Char2Wav: End-to-end speech synthesis,” in Proc. ICLR Workshop , Toulon, France, May 2017
2017
Earlier work this paper cites.
M. H. Davis, Empathy: A Social Psychological Approach . Routledge, 2018
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” in Proc. ICASSP , Calgary, Canada, Apr. 2018, pp. 4779–4783
2018
Earlier work this paper cites.
Y. Chiba, T. Nose, T. Kase, M. Yamanaka, and A. Ito, “An analysis of the effect of emotional speech synthesis on non-task-oriented dialogue system,” in Proc. SIGDIAL , Melbourne, Australia, Jul. 2018, pp. 371–375
2018
Cited alongside, same era.
H. Rashkin, E. M. Smith, M. Li, and Y.-L. Boureau, “Towards empathetic open-domain conversation models: A new benchmark and dataset,” in Proc. ACL , Florence, Italy, Jul. 2019, pp. 5370–5381
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT , Minneapolis, U.S.A., Jun. 2019, pp. 4171–4186
2019
Cited alongside, same era.
W. Ping, K. Peng, and J. Chen, “ClariNet: Parallel wave generation in end-to-end text-to-speech,” in Proc. ICLR , New Orleans, U.S.A., May 2019
2019
Cited alongside, same era.
Y. Yamazaki, Y. Chiba, T. Nose, and A. Ito, “Neural spoken-response generation using prosodic and linguistic context for conversational systems,” in Proc. INTERSPEECH , Brno, Czech Republic, Sep. 2021, pp. 246–250
2021
Later among the works it cites.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech 2: Fast and high-quality end-to-end text to speech,” in Proc. ICLR , Vienna, Austria, May 2021
2021
Later among the works it cites.
R. J. Weiss, R. Skerry-Ryan, E. Battenberg, S. Mariooryad, and D. P. Kingma, “Wave-Tacotron: Spectrogram-free end-to-end text-to-speech synthesis,” in Proc. ICASSP , Montreal, Canada, Jun. 2021, pp. 5679–5683
2021
Later among the works it cites.
J. Cong, S. Yang, N. Hu, G. Li, L. Xie, and D. Su, “Controllable context-aware conversational speech synthesis,” in Proc. INTERSPEECH , Brno, Czech Republic, Sep. 2021, pp. 4658–4662
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Park and T. Mulc, “CSS10: A collection of single speaker speech datasets for 10 languages,” in Proc. INTERSPEECH , Graz, Austria, Sep. 2019, pp. 1566–1570
2019
Cited alongside, same era.
Q. Lui, H. Chen, Z. Ren, P. Ren, Z. Tu, and Z. Chen, “EmpDG: Multi-resolution interactive empathetic dialogue generation,” in Proc. COLING , Barcelona, Spain, Dec. 2020, pp. 4454–4466
2020
Cited alongside, same era.
S. Takamichi, R. Sonobe, K. Mitsui, Y. Saito, T. Koriyama, N. Tanji, and H. Saruwatari, “JSUT and JVS: Free Japanese voice corpora for accelerating speech synthesis research,” Acoustical Science and Technology , vol. 41, no. 5, pp. 761–768, Sep. 2020
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Proc. NeurIPS , Vancouver, Canada, Dec. 2020
2020
Cited alongside, same era.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. NeurIPS , Vancouver, Canada, Dec. 2020
2020
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common Voice: A massively-multilingual speech corpus,” in Proc. LREC , Marseille, France, May 2020, pp. 4218–4222
2020
Cited alongside, same era.
H. Guo, S. Zhang, F. K. Soong, L. He, and L. Xie, “Conversational end-to-end TTS for voice agents,” in Proc. SLT , Shenzhen, China, Jan. 2021, pp. 403–409
2021
Cited alongside, same era.
Y. Liu, W. Maier, W. Minker, and S. Ultes, “Empathetic dialogue generation with pre-trained RoBERTa-GPT2 and external knowledge,” in Proc. IWSDS , Singapore, Nov. 2021
2021
Later among the works it cites.
Y. Xie and P. Pu, “Empathetic dialog generation with fine-grained intents,” in Proc. CoNLL , Punta Cana, Dominican Republic, Nov. 2021, pp. 133–147
2021
Later among the works it cites.
Wataru-Nakata, “FastSpeech2-JSUT,” https://github.com/Wataru-Nakata/FastSpeech2-JSUT
2021
Later among the works it cites.
A. Łańcucki, “FastPitch: Parallel text-to-speech with pitch prediction,” in Proc. ICASSP , Montreal, Canada, Jun. 2021, pp. 6588–6592
2021
Later among the works it cites.
keonlee9420, “Expressive-FastSpeech2,” https://github.com/keonlee9420/Expressive-FastSpeech2/tree/conversational
2021
Later among the works it cites.
Y. Saito, Y. Nishimura, S. Takamichi, K. Tachibana, and H. Saruwatari, “STUDIES: Corpus of japanese empathetic dialogue speech towards user-friendly voice agent,” in Proc. INTERSPEECH (Accepted) , Incheon, South Korea, Sep. 2022
2022
Closest in time.
C. Du and K. Yu, “Phone-level prosody modelling with GMM-based MDN for diverse and controllable speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 190–201, Jan. 2022
2022
Closest in time.