Fetching the paper…
Reading the bibliography…
Generating speech across different accents while preserving speaker identity is crucial for various real-world applications.
M. Zhang, Y. Zhou, Z. Wu, and H. Li, “Zero-shot multi-speaker accent tts with limited accent data,” in 2023 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) . IEEE, 2023, pp. 1931–1936
1936
Earlier work this paper cites.
D. L. Bolinger, “A theory of pitch accent in english,” Word , vol. 14, no. 2-3, pp. 109–149, 1958
1958
Earlier work this paper cites.
J. Terken, “Fundamental frequency and perceived prominence of accented syllables,” The Journal of the Acoustical Society of America , vol. 89, no. 4, pp. 1768–1776, 1991
1991
Earlier work this paper cites.
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in Proceedings of IEEE pacific rim conference on communications computers and signal processing , vol. 1. IEEE, 1993, pp. 125–128
1993
Earlier work this paper cites.
J. H. Hansen and L. M. Arslan, “Foreign accent classification using source generator based prosodic features,” in 1995 International Conference on Acoustics, Speech, and Signal Processing , vol. 1. IEEE, 1995, pp. 836–839
1995
Earlier work this paper cites.
C. I. Watson, J. Harrington, and Z. Evans, “An acoustic comparison between new zealand and australian english vowels,” Australian journal of linguistics , vol. 18, no. 2, pp. 185–207, 1998
1998
Earlier work this paper cites.
L. W. Kat and P. Fung, “Fast accent identification and accented speech recognition,” in 1999 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings. ICASSP99 (Cat. No. 99CH36258) , vol. 1. IEEE, 1999, pp. 221–224
1999
Earlier work this paper cites.
Q. Yan, S. Vaseghi, D. Rentzos, C.-H. Ho, and E. Turajlic, “Analysis of acoustic correlates of british, australian and american accents,” in 2003 IEEE Workshop on Automatic Speech Recognition and Understanding (IEEE Cat. No. 03EX721) . IEEE, 2003, pp. 345–350
2003
Earlier work this paper cites.
J. Vaissière and P. B. de Mareüil, “Identifying a language or an accent: from segments to prosody,” in Workshop MIDL 2004 , 2004, pp. 1–4
2004
Earlier work this paper cites.
T. Cho and J. M. McQueen, “Prosodic influences on consonant production in dutch: Effects of prosodic boundaries, phrasal accent and lexical stress,” Journal of Phonetics , vol. 33, no. 2, pp. 121–157, 2005
2005
Earlier work this paper cites.
J. Fletcher, E. Grabe, and P. Warren, “Intonational variation in four dialects of english: the high rising tune,” Prosodic typology: The phonology of intonation and phrasing , pp. 390–409, 2005
2005
Earlier work this paper cites.
P. B. d. Mareüil and B. Vieru-Dimulescu, “The contribution of prosody to the perception of foreign accent,” Phonetica , vol. 63, no. 4, pp. 247–267, 2006
2006
Earlier work this paper cites.
P. Angkititrakul and J. H. Hansen, “Advances in phone-based modeling for automatic accent classification,” IEEE transactions on audio, speech, and language processing , vol. 14, no. 2, pp. 634–646, 2006
2006
Earlier work this paper cites.
M. Müller, “Dynamic time warping,” Information retrieval for music and motion , pp. 69–84, 2007
2007
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.” Journal of machine learning research , vol. 9, no. 11, 2008
2008
Earlier work this paper cites.
H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,” speech communication , vol. 51, no. 11, pp. 1039–1064, 2009
2009
Earlier work this paper cites.
P. Gomez, “British and american english pronunciation differences,” 2009
2009
Earlier work this paper cites.
I. Cohen, Y. Huang, J. Chen, J. Benesty, J. Benesty, J. Chen, Y. Huang, and I. Cohen, “Pearson correlation coefficient,” Noise reduction in speech processing , pp. 1–4, 2009
2009
Earlier work this paper cites.
L. Loots and T. Niesler, “Automatic conversion between pronunciations of different english accents,” Speech Communication , vol. 53, no. 1, pp. 75–84, 2011
2011
Earlier work this paper cites.
F. Biadsy, J. B. Hirschberg, and D. P. Ellis, “Dialect and accent recognition using phonetic-segmentation supervectors,” 2011
2011
Earlier work this paper cites.
H. Ze, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in 2013 ieee international conference on acoustics, speech and signal processing . IEEE, 2013, pp. 7962–7966
2013
Earlier work this paper cites.
E. Reinisch and L. L. Holt, “Lexically guided phonetic retuning of foreign-accented speech and its generalization.” Journal of Experimental Psychology: Human Perception and Performance , vol. 40, no. 2, p. 539, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. Fan, Y. Qian, F. K. Soong, and L. He, “Multi-speaker modeling and speaker adaptation for dnn-based tts synthesis,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 4475–4479
2015
Earlier work this paper cites.
H. Ding, R. Hoffmann, and D. Hirst, “Prosodic transfer: A comparison study of f0 patterns in l2 english by chinese speakers,” in Speech Prosody , vol. 2016, 2016, pp. 756–760
2016
Cited alongside, same era.
R. C. Streijl, S. Winkler, and D. S. Hands, “Mean opinion score (mos) revisited: methods and applications, limitations and alternatives,” Multimedia Systems , vol. 22, no. 2, pp. 213–227, 2016
2016
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al. , “Tacotron: Towards end-to-end speech synthesis,” Interspeech 2017 , 2017
2017
Cited alongside, same era.
C. Veaux, J. Yamagishi, K. MacDonald et al. , “Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” University of Edinburgh. The Centre for Speech Technology Research (CSTR) , vol. 6, p. 15, 2017
2017
Cited alongside, same era.
2021
Later among the works it cites.
R. Liu, B. Sisman, G. Gao, and H. Li, “Expressive tts training with frame and style reconstruction loss,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 1806–1818, 2021
2021
Later among the works it cites.
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, “Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,” in International Conference on Machine Learning . PMLR, 2022, pp. 2709–2720
2022
Later among the works it cites.
T. Li, X. Wang, Q. Xie, Z. Wang, and L. Xie, “Cross-speaker emotion disentangling and transfer for end-to-end speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1448–1460, 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Cited alongside, same era.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. Lopez Moreno, Y. Wu et al. , “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” Advances in neural information processing systems , vol. 31, 2018
2018
Cited alongside, same era.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4879–4883
2018
Cited alongside, same era.
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in International Conference on Machine Learning . PMLR, 2018, pp. 5180–5189
2018
Cited alongside, same era.
G. Zhao, S. Sonsaat, A. Silpachai, I. Lucic, E. Chukharev-Hudilainen, J. Levis, and R. Gutierrez-Osuna, “L2-arctic: A non-native english speech corpus.” in Interspeech , 2018, pp. 2783–2787
2018
Cited alongside, same era.
X. Wang, S. Takaki, and J. Yamagishi, “Autoregressive neural f0 model for statistical parametric speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 8, pp. 1406–1419, 2018
2018
Cited alongside, same era.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural speech synthesis with transformer network,” in Proceedings of the AAAI conference on artificial intelligence , vol. 33, no. 01, 2019, pp. 6706–6713
2019
Cited alongside, same era.
Y.-J. Zhang, S. Pan, L. He, and Z.-H. Ling, “Learning latent representations for style control and transfer in end-to-end speech synthesis,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6945–6949
2019
Cited alongside, same era.
Y. Lei, S. Yang, X. Wang, and L. Xie, “Msemotts: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 853–864, 2022
2022
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
G. Tinchev, M. Czarnowska, K. Deja, K. Yanagisawa, and M. Cotescu, “Modelling low-resource accents without accent-specific tts frontend,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
T.-N. Nguyen, N.-Q. Pham, and A. Waibel, “Syntacc: Synthesizing multi-accent speech by weight factorization,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
M. Zhang, X. Zhou, Z. Wu, and H. Li, “Towards zero-shot multi-speaker multi-accent text-to-speech synthesis,” IEEE Signal Processing Letters , 2023
2023
Later among the works it cites.
L. Ma, Y. Zhang, X. Zhu, Y. Lei, Z. Ning, P. Zhu, and L. Xie, “Accent-vits: accent transfer for end-to-end tts,” in National Conference on Man-Machine Speech Communication . Springer, 2023, pp. 203–214
2023
Later among the works it cites.
C. Gong, X. Wang, E. Cooper, D. Wells, L. Wang, J. Dang, K. Richmond, and J. Yamagishi, “Zmm-tts: Zero-shot multilingual and multispeaker speech synthesis conditioned on self-supervised discrete speech representations,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024
2024
Closest in time.
W. Wang, Y. Song, and S. Jha, “Usat: A universal speaker-adaptive text-to-speech approach,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024
2024
Closest in time.
M. Nishihara, D. Wells, K. Richmond, and A. Pine, “Low-dimensional style token control for hyperarticulated speech synthesis,” in Proc. Interspeech , 2024
2024
Closest in time.
X. Chen, X. Wang, S. Zhang, L. He, Z. Wu, X. Wu, and H. Meng, “Stylespeech: Self-supervised style enhancing with vq-vae-based pre-training for expressive audiobook speech synthesis,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 12 316–12 320
2024
Closest in time.
Y. Li, C. Yu, G. Sun, W. Zu, Z. Tian, Y. Wen, W. Pan, C. Zhang, J. Wang, Y. Yang et al. , “Cross-utterance conditioned vae for speech generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 4263–4276, 2024
2024
Closest in time.
S. Dutta and S. Ganapathy, “Zero shot audio to audio emotion transfer with speaker disentanglement,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 10 371–10 375
2024
Closest in time.
H. Choi, J.-S. Bae, J. Y. Lee, S. Mun, J. Lee, H.-Y. Cho, and C. Kim, “Mels-tts: Multi-emotion multi-lingual multi-speaker text-to-speech system via disentangled style tokens,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 12 682–12 686
2024
Closest in time.
H. Tang, X. Zhang, N. Cheng, J. Xiao, and J. Wang, “Ed-tts: Multi-scale emotion modeling using cross-domain emotion diarization for emotional speech synthesis,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 12 146–12 150
2024
Closest in time.
X. Zhou, M. Zhang, Y. Zhou, Z. Wu, and H. Li, “Accented text-to-speech synthesis with limited data,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 1699–1711, 2024
2024
Closest in time.
2024
Closest in time.
J. Melechovsky, A. Mehrish, B. SISMAN, and D. Herremans, “Dart: Disentanglement of accent and speaker representation in multispeaker text-to-speech,” in Audio Imagination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Generation , 2024
2024
Closest in time.
2024
Closest in time.
R. Liu, B. Sisman, G. Gao, and H. Li, “Controllable accented text-to-speech synthesis with fine and coarse-grained intensity rendering,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024
2024
Closest in time.