Fetching the paper…
Reading the bibliography…
Style transfer TTS has shown impressive performance in recent years.
V. D. M. Laurens and G. Hinton, “Visualizing data using t-SNE,” Journal of Machine Learning Research , vol. 9, no. 2605, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al. , “Tacotron: Towards end-to-end speech synthesis,” Proc. Interspeech 2017 , pp. 4006–4010, 2017
2017
Earlier work this paper cites.
H. Li, Y. Kang, and Z. Wang, “Emphasis: An emotional phoneme-based acoustic model for speech synthesis system,” Proc. Interspeech 2018 , pp. 3077–3081, 2018
2018
Earlier work this paper cites.
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous, “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,” in international conference on machine learning . PMLR, 2018, pp. 4693–4702
2018
Earlier work this paper cites.
Y. Jia, Y. Zhang, R. J. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. L. Moreno et al. , “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems , 2018, pp. 4485–4495
2018
Earlier work this paper cites.
Y. Wang, D. Stanton, Y. Zhang, R. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, F. Ren, Y. Jia, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in Proc. ICML, 2018, pp. 5180–5189
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
X. Zhu, S. Yang, G. Yang, and L. Xie, “Controlling emotion strength with relative attribute for end-to-end speech synthesis,” 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , pp. 192–199, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y.-J. Zhang, S. Pan, L. He, and Z.-H. Ling, “Learning latent representations for style control and transfer in end-to-end speech synthesis,” ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6945–6949, 2019
2019
Earlier work this paper cites.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural speech synthesis with transformer network,” in Proceedings of the AAAI conference on artificial intelligence , vol. 33, no. 01, 2019, pp. 6706–6713
2019
Cited alongside, same era.
P.-F. Wu, Z. Ling, L. juan Liu, Y. Jiang, H.-C. Wu, and L.-R. Dai, “End-to-end emotional speech synthesis using style tokens and semi-supervised training,” 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) , pp. 623–627, 2019
2019
Cited alongside, same era.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
N. Tits, K. El Haddad, and T. Dutoit, “Exploring transfer learning for low resource emotional tts,” in Intelligent Systems and Applications: Proceedings of the 2019 Intelligent Systems Conference (IntelliSys) Volume 1 . Springer, 2020, pp. 52–60
2021
Later among the works it cites.
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in International Conference on Machine Learning . PMLR, 2021, pp. 5530–5540
2021
Later among the works it cites.
2021
Later among the works it cites.
T. Li, X. Wang, Q. Xie, Z. Wang, and L. Xie, “Cross-speaker emotion disentangling and transfer for end-to-end speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1448–1460, 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 022–17 033, 2020
2020
Cited alongside, same era.
2021
Cited alongside, same era.
Z. Shang, Z. Huang, H. Zhang, P. Zhang, and Y. Yan, “Incorporating cross-speaker style transfer for multi-language text-to-speech.” in Interspeech , 2021, pp. 1619–1623
2021
Cited alongside, same era.
A. Kulkarni, V. Colotte, and D. Jouvet, “Improving transfer of expressivity for end-to-end multispeaker text-to-speech synthesis,” 2021 29th European Signal Processing Conference (EUSIPCO) , pp. 31–35, 2021
2021
Cited alongside, same era.
T. Li, S. Yang, L. Xue, and L. Xie, “Controllable emotion transfer for end-to-end speech synthesis,” 2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP) , pp. 1–5, 2021
2021
Cited alongside, same era.
Z. Zhang, Y. Gu, X. Han, S. Chen, C. Xiao, Z. Sun, Y. Yao, F. Qi, J. Guan, P. Ke et al. , “Cpm-2: Large-scale cost-effective pre-trained language models,” AI Open , vol. 2, pp. 216–224, 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
T. Li, X. Wang, Q. Xie, Z. Wang, M. Jiang, and L. Xie, “Cross-speaker Emotion Transfer Based On Prosody Compensation for End-to-End Speech Synthesis,” in Proc. Interspeech 2022 , 2022, pp. 5498–5502
2022
Later among the works it cites.
P. H. Seo, A. Nagrani, A. Arnab, and C. Schmid, “End-to-end generative pretraining for multimodal video captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 17 959–17 968
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.