Fetching the paper…
Reading the bibliography…
While recent Zero-Shot Text-to-Speech (ZS-TTS) models have achieved high naturalness and speaker similarity, they fall short in accent fidelity and control.
M. Zhang, Y. Zhou, Z. Wu, and H. Li, “Zero-Shot Multi-Speaker Accent TTS with Limited Accent Data,” in Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) , 2023, pp. 1931–1936
1936
Earlier work this paper cites.
P. J. Rousseeuw, “Silhouettes: A Graphical Aid to the Interpretation and Validation of Cluster Analysis,” Journal of Computational and Applied Mathematics , vol. 20, pp. 53–65, 1987
1987
Earlier work this paper cites.
L.-G. Rosina, English with an Accent: Language Ideology and Discrimination in the United States . Routledge, 1997
1997
Earlier work this paper cites.
C. X. Ling and C. Li, “Data Mining for Direct marketing: Problems and Solutions,” in Proceedings of the 4th International Conference on Knowledge Discovery and Data Mining , vol. 98, 1998, pp. 73–79
1998
Earlier work this paper cites.
D. Felps, H. Bortfeld, and R. Gutierrez-Osuna, “Foreign Accent Conversion in Computer Assisted Pronunciation Training,” Speech Communication , vol. 51, no. 10, pp. 920–932, 2009
2009
Earlier work this paper cites.
A. Gluszek and J. F. Dovidio, “The Way They Speak: A Social Psychological Perspective on the Stigma of Nonnative Accents in Communication,” Personality and Social Psychology Review , vol. 14, no. 2, pp. 214–237, 2010
2010
Earlier work this paper cites.
J. Yamagishi, C. Veaux, and K. MacDonald, “CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92),” https://datashare.ed.ac.uk/handle/10283/3443
2012
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio Augmentation for Speech Recognition,” in Proc. Interspeech , 2015, pp. 3586–3589
2015
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A Study on Data Augmentation of Reverberant Speech for Robust Speech Recognition,” in IEEE ICASSP , 2017, pp. 5220–5224
2017
Earlier work this paper cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, z. Chen, P. Nguyen, R. Pang, I. Lopez Moreno, and Y. Wu, “Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis,” in Advances in Neural Information Processing Systems , vol. 31. Curran Associates, Inc., 2018
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, “Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions,” in IEEE ICASSP , 2018, pp. 4779–4783
2018
Earlier work this paper cites.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized End-to-End Loss for Speaker Verification,” in IEEE ICASSP , 2018, pp. 4879–4883
2018
Earlier work this paper cites.
C. Agarwal and P. Chakraborty, “A Review of Tools and Techniques for Computer Aided Pronunciation Training (CAPT) in English,” Education and Information Technologies , vol. 24, no. 6, pp. 3731–3743, 2019
2019
Earlier work this paper cites.
D. Pal, C. Arpnikanondt, S. Funilkul, and V. Varadarajan, “User Experience with Smart Voice Assistants: The Accent Perspective,” in 10th International Conference on Computing, Communication and Networking Technologies (ICCCNT) . IEEE, 2019, pp. 1–6
2019
Cited alongside, same era.
G. Zhao, S. Ding, and R. Gutierrez-Osuna, “Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams,” in Proc. Interspeech , 2019, pp. 2843–2847
2019
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common Voice: A Massively-Multilingual Speech Corpus,” in Proceedings of the 12th Language Resources and Evaluation Conference . European Language Resources Association, 2020, pp. 4218–4222
2020
Cited alongside, same era.
J. J. Webber, O. Perrotin, and S. King, “Hider-Finder-Combiner: An Adversarial Architecture for General Speech Signal Modification,” in Proc. Interspeech , 2020, pp. 3206–3210
2020
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar, and W.-N. Hsu, “Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale,” in Advances in Neural Information Processing Systems , vol. 36. Curran Associates, Inc., 2023, pp. 14 005–14 034
2023
Later among the works it cites.
J. Zuluaga-Gomez, S. Ahmed, D. Visockas, and C. Subakan, “CommonAccent: Exploring Large Acoustic Pretrained Models for Accent Classification Based on Common Voice,” in Proc. Interspeech , 2023, pp. 5291–5295
2023
Later among the works it cites.
M. Zhang, X. Zhou, Z. Wu, and H. Li, “Towards Zero-Shot Multi-Speaker Multi-Accent Text-to-Speech Synthesis,” IEEE Signal Processing Letters , vol. 30, pp. 947–951, 2023
2023
Later among the works it cites.
Y. Koizumi, H. Zen, S. Karita, Y. Ding, K. Yatabe, N. Morioka, M. Bacchiani, Y. Zhang, W. Han, and A. Bapna, “LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,” in Proc. Interspeech , 2023, pp. 5496–5500
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
G. Spiteri Miggiani, “Exploring Applied Strategies for English-Language Dubbing,” Journal of Audiovisual Translation , vol. 4, no. 1, p. 137–156, 2021
2021
Cited alongside, same era.
X. Shi, F. Yu, Y. Lu, Y. Liang, Q. Feng, D. Wang, Y. Qian, and L. Xie, “The Accented English Speech Recognition Challenge 2020: Open Datasets, Tracks, Baselines, Results and Methods,” in IEEE ICASSP , 2021, pp. 6918–6922
2021
Cited alongside, same era.
J. Melechovsky, A. Mehrish, D. Herremans, and B. Sisman, “Learning Accent Representation with Multi-Level VAE Towards Controllable Speech Synthesis,” in IEEE Spoken Language Technology Workshop (SLT) , 2022, pp. 928–935
2022
Cited alongside, same era.
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, “YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for Everyone,” in Proceedings of the 39th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 162. PMLR, 2022, pp. 2709–2720
2022
Cited alongside, same era.
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y. Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli, “XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,” in Proc. Interspeech , 2022, pp. 2278–2282
2022
Cited alongside, same era.
2023
Cited alongside, same era.
K. Deja, G. Tinchev, M. Czarnowska, M. Cotescu, and J. Droppo, “Diffusion-based Accent Modelling in Speech Synthesis,” in Proc. Interspeech , 2023, pp. 5516–5520
2023
Cited alongside, same era.
R. Liu, H. Zuo, D. Hu, G. Gao, and H. Li, “Explicit Intensity Control for Accented Text-to-Speech,” in Proc. Interspeech , 2023, pp. 22–26
2023
Cited alongside, same era.
Later among the works it cites.
R. Sanabria, N. Bogoychev, N. Markl, A. Carmantini, O. Klejch, and P. Bell, “The Edinburgh International Accents of English Corpus: Towards the Democratization of English ASR,” in IEEE ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
X. Zhou, M. Zhang, Y. Zhou, Z. Wu, and H. Li, “Accented Text-to-Speech Synthesis With Limited Data,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 1699–1711, 2024
2024
Closest in time.
L. Ma, Y. Zhang, X. Zhu, Y. Lei, Z. Ning, P. Zhu, and L. Xie, “Accent-VITS: Accent Transfer for End-to-End TTS,” in Man-Machine Speech Communication . Springer Nature Singapore, 2024, pp. 203–214
2024
Closest in time.
2024
Closest in time.
R. Liu, B. Sisman, G. Gao, and H. Li, “Controllable Accented Text-to-Speech Synthesis With Fine and Coarse-Grained Intensity Rendering,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 2188–2201, 2024
2024
Closest in time.