Fetching the paper…
Reading the bibliography…
The cross-speaker emotion transfer task in text-to-speech (TTS) synthesis particularly aims to synthesize speech for a target speaker with the emotion transferred from reference speech recorded by another (source) speaker.
V. Ferrari and A. Zisserman, “Learning visual attributes,” in Conference on Advances in Neural Information Processing Systems , 2008
2008
Earlier work this paper cites.
V. D. M. Laurens and G. Hinton, “Visualizing data using t-SNE,” Journal of Machine Learning Research , vol. 9, no. 2605, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in NIPS , 2014
2014
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” CoRR , vol. abs/1312.6114, 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Z. Ling, S. Kang, H. Zen, A. Senior, M. Schuster, X. Qian, H. Meng, and L. Deng, “Deep learning for acoustic modeling in parametric speech generation: A systematic review of existing techniques and future trends,” IEEE Signal Processing Magazine , vol. 32, pp. 35–52, 2015
2015
Earlier work this paper cites.
Y. Ohtani, N. Yu, M. Morita, and M. Akamine, “Emotional transplant in statistical speech synthesis based on emotion additive model,” in Interspeech , 2015
2015
Earlier work this paper cites.
B. S. Akanksh, S. Vekkot, and S. Tripathi, “Interconversion of emotions in speech using td-psola,” in SIRS , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for real-time style transfer and super-resolution,” in European conference on computer vision . Springer, 2016, pp. 694–711
2016
Earlier work this paper cites.
S. Ö. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman, S. Sengupta, and M. Shoeybi, “Deep voice: Real-time neural text-to-speech,” in ICML , 2017
2017
Earlier work this paper cites.
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. C. Courville, and Y. Bengio, “Char2wav: End-to-end speech synthesis,” in ICLR , 2017
2017
Earlier work this paper cites.
W. Ping, K. Peng, A. Gibiansky, S. Ö. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, “Deep voice 3: 2000-speaker neural text-to-speech.” 2017
2017
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, and S. Bengio, “Tacotron: Towards end-to-end speech synthesis,” INTERSPEECH , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. Inoue, S. Hara, M. Abe, N. Hojo, and Y. Ijima, “An investigation to transplant emotional expressions in dnn-based tts synthesis,” in 2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) , 2017
2017
Earlier work this paper cites.
A. Gibiansky, S. Ö. Arik, G. Diamos, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, “Deep voice 2: Multi-speaker neural text-to-speech,” in NIPS , 2017
2017
Earlier work this paper cites.
J. Lee, K. Cho, and T. Hofmann, “Fully character-level neural machine translation without explicit segmentation,” Transactions of the Association for Computational Linguistics , vol. 5, pp. 365–378, 2017
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in Proc. ICASSP . IEEE, 2018, pp. 4779–4783
2018
Earlier work this paper cites.
Y. Wang, D. Stanton, Y. Zhang, R. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, F. Ren, Y. Jia, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in Proc. ICML, 2018, pp. 5180–5189
2018
Cited alongside, same era.
H. Li, Y. Kang, and Z. Wang, “Emphasis: An emotional phoneme-based acoustic model for speech synthesis system,” in INTERSPEECH , 2018
2018
Cited alongside, same era.
X. Wu, Y. Cao, M. Wang, S. Liu, S. Kang, Z. Wu, X. Liu, D. Su, D. Yu, and H. Meng, “Rapid style adaptation using residual error embedding for expressive speech synthesis,” in INTERSPEECH , 2018
2018
Cited alongside, same era.
R. Li, Z. Wu, Y. Huang, J. Jia, H. Meng, and L. Cai, “Emphatic speech generation with conditioned input layer and bidirectional lstms for expressive speech synthesis,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 5129–5133, 2018
2018
Cited alongside, same era.
S. Ma, D. McDuff, and Y. Song, “Neural TTS stylization with adversarial and collaborative games,” in ICLR , 2019
2019
Later among the works it cites.
S.-Y. Um, S. Oh, K. Byun, I. Jang, and H.-G. Kang, “Emotional speech synthesis with rich and granularized control,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020
2020
Later among the works it cites.
S. Karlapati, A. Moinet, A. Joly, V. Klimkov, D. Saez-Trigueros, and T. Drugman, “Copycat: Many-to-many fine-grained prosody transfer for neural text-to-speech,” in INTERSPEECH , 2020
2020
Later among the works it cites.
A. Kulkarni, V. Colotte, and D. Jouvet, “Transfer learning of the expressivity using flow metric learning in multispeaker text-to-speech synthesis,” in INTERSPEECH , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous, “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,” in international conference on machine learning . PMLR, 2018, pp. 4693–4702
2018
Cited alongside, same era.
Y. Jia, Y. Zhang, R. J. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. Lopez-Moreno, and Y. Wu, “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in NeurIPS , 2018
2018
Cited alongside, same era.
X. Wu, L. Sun, S. Kang, S. Liu, Z. Wu, X. Liu, and H. Meng, “Feature based adaptation for speaking style synthesis,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 5304–5308, 2018
2018
Cited alongside, same era.
F. Yang, S. Yang, P. Zhu, P. Yan, and L. Xie, “Improving mandarin end-to-end speech synthesis by self-attention and learnable gaussian bias,” 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , pp. 208–213, 2019
2019
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in NeurIPS , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y.-J. Zhang, S. Pan, L. He, and Z.-H. Ling, “Learning latent representations for style control and transfer in end-to-end speech synthesis,” ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6945–6949, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
R. Valle, J. Li, R. Prenger, and B. Catanzaro, “Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,” ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6189–6193, 2020
2020
Later among the works it cites.
A. Sorin, S. Shechtman, and R. Hoory, “Principal style components: Expressive style control and cross-speaker transfer in neural tts,” in INTERSPEECH , 2020
2020
Later among the works it cites.
S. Um, S. Oh, K. Byun, I. Jang, C. Ahn, and H.-G. Kang, “Emotional speech synthesis with rich and granularized control,” ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 7254–7258, 2020
2020
Later among the works it cites.
E. Battenberg, R. Skerry-Ryan, S. Mariooryad, D. Stanton, D. Kao, M. Shannon, and T. Bagby, “Location-relative attention mechanisms for robust long-form speech synthesis,” in ICASSP , 2020
2020
Later among the works it cites.
B. Desplanques, J. Thienpondt, and K. Demuynck, “Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,” in INTERSPEECH , 2020
2020
Later among the works it cites.
Q. Xie, X. Tian, G. Liu, K. Song, L. Xie, Z. Wu, H. Li, S. Shi, H. Li, F. Hong et al. , “The multi-speaker multi-style voice cloning challenge 2021,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 8613–8617
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
L. Xue, S. Pan, L. He, L. Xie, and F. Soong, “Cycle consistent network for end-to-end style transfer tts training.” Neural networks : the official journal of the International Neural Network Society , vol. 140, pp. 223–236, 2021
2021
Closest in time.
Z. Wang, X. Zhou, F. Yang, T. Li, H. Du, L. Xie, W. Gan, H. Chen, and H. Li, “Enriching Source Style Transfer in Recognition-Synthesis Based Non-Parallel Voice Conversion,” in Proc. Interspeech , 2021, pp. 831–835
2021
Closest in time.
A. Kulkarni, V. Colotte, and D. Jouvet, “Improving transfer of expressivity for end-to-end multispeaker text-to-speech synthesis,” in 29th European Signal Processing Conference (EUSIPCO 2021) . Dublin (virtuel), Ireland: European Association for Signal Processing (EURASIP), Aug. 2021. [Online]. Available: https://hal.archives-ouvertes.fr/hal-02978485
2021
Closest in time.
T. Li, S. Yang, L. Xue, and L. Xie, “Controllable emotion transfer for end-to-end speech synthesis,” 2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP) , pp. 1–5, 2021
2021
Closest in time.
X. Li, C. Song, J. Li, Z. Wu, J. Jia, and H. Meng, “Towards multi-scale style control for expressive speech synthesis,” in INTERSPEECH , 2021
2021
Closest in time.
S. Kannan, P. R. Raju, R. S. S. Madhav, and S. Tripathi, “Voice conversion using spectral mapping and td-psola,” 2021
2021
Closest in time.