Fetching the paper…
Reading the bibliography…
In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content.
Second language speech learning: Theory, findings and problems
James Flege, · 1995
Earlier work this paper cites.
“Machine transliteration,”
Kevin Knight and Jonathan Graehl, · 1998
Earlier work this paper cites.
“The cmu arctic speech databases,”
John Kominek and Alan W. Black, · 2004
Earlier work this paper cites.
“Machine transliteration survey,”
Sarvnaz Karimi, Falk Scholer, and Andrew Turpin, · 2011
Earlier work this paper cites.
“Segmental and prosodic approaches to accent management,”
Alison Behrman, · 2014
Earlier work this paper cites.
“Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” [sound], 2017
Christophe Veaux, Junichi Yamagishi, and Kirsten MacDonald, · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“L2-arctic: A non-native english speech corpus,”
Guanlong Zhao, Sinem Sonsaat, Alif Silpachai, Ivana Lucic, Evgeny Chukharev-Hudilainen, John Levis, and Ricardo Gutierrez-Osuna, · 2018
Earlier work this paper cites.
“Voice transformer network: Sequence-to-sequence voice conversion using transformer with text-to-speech pretraining,” 2019
Wen-Chin Huang, Tomoki Hayashi, Yi-Chiao Wu, Hirokazu Kameoka, and Tomoki Toda, · 2019
Earlier work this paper cites.
“End-to-end accent conversion without using native utterances,”
Songxiang Liu, Disong Wang, Yuewen Cao, Lifa Sun, Xixin Wu, Shiyin Kang, Zhiyong Wu, Xunying Liu, Dan Su, Dong Yu, and Helen Meng, · 2020
Earlier work this paper cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Earlier work this paper cites.
“Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,”
Edresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior, Eren Gölge, and Moacir Antonelli Ponti, · 2021
Earlier work this paper cites.
“Converting foreign accent speech without a reference,”
Guanlong Zhao, Shaojin Ding, and Ricardo Gutierrez-Osuna, · 2021
Earlier work this paper cites.
“Accent and speaker disentanglement in many-to-many voice conversion,”
Zhichao Wang, Wenshuo Ge, Xiong Wang, Shan Yang, Wendong Gan, Haitao Chen, Hai Li, Lei Xie, and Xiulin Li, · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“Accent Conversion using Pre-trained Model and Synthesized Data from Voice Conversion,”
Tuan Nam Nguyen, Ngoc-Quan Pham, and Alexander Waibel, · 2022
Cited alongside, same era.
“Training text-to-speech systems from synthetic data: A practical approach for accent transfer tasks,” 2022
Lev Finkelstein, Heiga Zen, Norman Casagrande, Chun an Chan, Ye Jia, Tom Kenter, Alexey Petelin, Jonathan Shen, Vincent Wan, Yu Zhang, Yonghui Wu, and Rob Clark, · 2022
Cited alongside, same era.
“Training language models to follow instructions with human feedback,”
“Speak foreign languages with your own voice: Cross-lingual neural codec language modeling,”
Zi-Hua Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, Yu Wu, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei, · 2023
Later among the works it cites.
“Gpt-4 technical report,”
OpenAI, · 2023
Later among the works it cites.
“The edinburgh international accents of english corpus: Towards the democratization of english asr,” 2023
Ramon Sanabria, Nikolay Bogoychev, Nina Markl, Andrea Carmantini, Ondrej Klejch, and Peter Bell, · 2023
Later among the works it cites.
“Voice-preserving zero-shot multiple accent conversion,”
Mumin Jin, Prashant Serai, Jilong Wu, Andros Tjandra, Vimal Manohar, and Qing He, · 2023
Later among the works it cites.
“Different models of transliteration - a comprehensive review,”
Mahima Yadav, Ishan Kumar, and Ayush Kumar, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe, · 2022
Cited alongside, same era.
“Zero-shot foreign accent conversion without a native reference,”
Waris Quamer, Anurag Das, John M. Levis, Evgeny Chukharev-Hudilainen, and Ricardo Gutierrez-Osuna, · 2022
Cited alongside, same era.
“Accentron: Foreign accent conversion to arbitrary non-native speakers using zero-shot learning,”
Shaojin Ding, Guanlong Zhao, and Ricardo Gutierrez-Osuna, · 2022
Cited alongside, same era.
“gpt-3.5-turbo-1106,” Online, 2022,
OpenAI, · 2022
Cited alongside, same era.
“Robust speech recognition via large-scale weak supervision,” 2022
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever, · 2022
Cited alongside, same era.
“Accent-vits:accent transfer for end-to-end tts,” 2023
Linhan Ma, Yongmao Zhang, Xinfa Zhu, Yi Lei, Ziqian Ning, Pengcheng Zhu, and Lei Xie, · 2023
Cited alongside, same era.
“Evaluating methods for ground-truth-free foreign accent conversion,” 2023
Wen-Chin Huang and Tomoki Toda, · 2023
Cited alongside, same era.
“Modelling low-resource accents without accent-specific tts frontend,”
Georgi Tinchev, Marta Czarnowska, Kamil Deja, Kayoko Yanagisawa, and Marius Cotescu, · 2023
Cited alongside, same era.
“gpt-4o-2024-05-13,” Online, 2023,
OpenAI, · 2023
Later among the works it cites.
“Libritts-r: A restored multi-speaker text-to-speech corpus,” 2023
Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding, Kohei Yatabe, Nobuyuki Morioka, Michiel Bacchiani, Yu Zhang, Wei Han, and Ankur Bapna, · 2023
Later among the works it cites.
“Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice,” 2023
Juan Zuluaga-Gomez, Sara Ahmed, Danielius Visockas, and Cem Subakan, · 2023
Later among the works it cites.
“Voiceshop: A unified speech-to-speech framework for identity-preserving zero-shot voice editing,” 2024
Philip Anastassiou, Zhenyu Tang, Kainan Peng, Dongya Jia, Jiaxin Li, Ming Tu, Yuping Wang, Yuxuan Wang, and Mingbo Ma, · 2024
Closest in time.
“Xtts: a massively multilingual zero-shot text-to-speech model,” 2024
Edresson Casanova, Kelly Davis, Eren Gölge, Görkem Göknar, Iulian Gulea, Logan Hart, Aya Aljafari, Joshua Meyer, Reuben Morais, Samuel Olayemi, and Julian Weber, · 2024
Closest in time.
“Globe: A high-quality english corpus with global accents for zero-shot speaker adaptive text-to-speech,” 2024
Wenbin Wang, Yang Song, and Sanjay Jha, · 2024
Closest in time.
“Transfer the linguistic representations from tts to accent conversion with non-parallel data,” 2024
Xi Chen, Jiakun Pei, Liumeng Xue, and Mingyang Zhang, · 2024
Closest in time.