Fetching the paper…
Reading the bibliography…
Voice conversion aims to modify the source speaker's voice to resemble the target speaker while preserving the original speech content.
Cross-language voice conversion. In International Conference on Acoustics, Speech, and Signal Processing . IEEE, 345–348
Masanobu Abe, Kiyohiro Shikano, and Hisao Kuwabara. 1990 · 1990
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
The EMIME bilingual database
Mirjam Wester. 2010 · 2010
Earlier work this paper cites.
Voice conversion using dynamic frequency warping with amplitude scaling, for parallel or nonparallel corpora
Elizabeth Godoy, Olivier Rosec, and Thierry Chonavel. 2011 · 2011
Earlier work this paper cites.
Combining source and system information for limited data speaker verification.. In Interspeech . 1836–1840
Rohan Kumar Das, S Abhiram, SR Mahadeva Prasanna, and AG Ramakrishnan. 2014 · 2014
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015 · 2015
Earlier work this paper cites.
REAPER: Robust epoch and pitch estimator,
D. Talkin. 2015 · 2015
Earlier work this paper cites.
Dual learning for machine translation
Di He, Yingce Xia, Tao Qin, Liwei Wang, Nenghai Yu, Tie-Yan Liu, and Wei-Ying Ma. 2016 · 2016
Earlier work this paper cites.
Phonetic posteriorgrams for many-to-one voice conversion without parallel data training. In 2016 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 1–6
Lifa Sun, Kun Li, Hao Wang, Shiyin Kang, and Helen Meng. 2016 · 2016
Earlier work this paper cites.
Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline. In 2017 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA) . IEEE, 1–5
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng. 2017 · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1125–1134
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. 2017 · 2017
Earlier work this paper cites.
Least squares generative adversarial networks. In Proceedings of the IEEE international conference on computer vision . 2794–2802
Xudong Mao, Qing Li, Haoran Xie, Raymond YK Lau, Zhen Wang, and Stephen Paul Smolley. 2017 · 2017
Earlier work this paper cites.
Montreal forced aligner: Trainable text-speech alignment using kaldi.. In Interspeech , Vol. 2017. 498–502
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Tacotron: Towards End-to-End Speech Synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Earlier work this paper cites.
Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks. In 2018 26th European Signal Processing Conference (EUSIPCO) . IEEE, 2100–2104
Takuhiro Kaneko and Hirokazu Kameoka. 2018 · 2018
Earlier work this paper cites.
Average Modeling Approach to Voice Conversion with Non-Parallel Data.. In Odyssey , Vol. 2018. 227–232
Xiaohai Tian, Junchao Wang, Haihua Xu, Eng Siong Chng, and Haizhou Li. 2018 · 2018
Earlier work this paper cites.
The M-AILABS Speech Dataset – caito
Munich Artificial Intelligence Laboratories GmbH. 2017 · 2019
Cited alongside, same era.
CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92)
Junichi Yamagishi, Christophe Veaux, Kirsten MacDonald, et al · 2019
Cited alongside, same era.
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019 · 2019
Cited alongside, same era.
Many-to-many cross-lingual voice conversion with a jointly trained speaker embedding network. In 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) . IEEE, 1282–1287
Yi Zhou, Xiaohai Tian, Rohan Kumar Das, and Haizhou Li. 2019a · 2019
Cited alongside, same era.
Cross-lingual voice conversion with bilingual phonetic posteriorgram and average modeling. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6790–6794
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2022 · 2022
Later among the works it cites.
Cross-lingual text-to-speech using multi-task learning and speaker classifier joint training
Lei He. 2022 · 2022
Later among the works it cites.
An Empirical Study on L2 Accents of Cross-lingual Text-to-Speech Systems via Vowel Space
Jihwan Lee, Jae-Sung Bae, Seongkyu Mun, Heejin Choi, Joun Yeop Lee, Hoon-Young Cho, and Chanwoo Kim. 2022 · 2022
Later among the works it cites.
Simple and effective unsupervised speech synthesis
Alexander H Liu, Cheng-I Jeff Lai, Wei-Ning Hsu, Michael Auli, Alexei Baevski, and James Glass. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yi Zhou, Xiaohai Tian, Haihua Xu, Rohan Kumar Das, and Haizhou Li. 2019b · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Cited alongside, same era.
Conformer: Convolution-augmented Transformer for Speech Recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al · 2020
Cited alongside, same era.
Glow-tts: A generative flow for text-to-speech via monotonic alignment search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon. 2020 · 2020
Cited alongside, same era.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020 · 2020
Cited alongside, same era.
Multi-Lingual Multi-Speaker Text-to-Speech Synthesis for Voice Cloning with Online Speaker Enrollment.. In Interspeech . 2932–2936
Zhaoyu Liu and Brian Mak. 2020 · 2020
Cited alongside, same era.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020 · 2020
Cited alongside, same era.
Neural analysis and synthesis: Reconstructing speech from self-supervised representations
Hyeong-Seok Choi, Juheon Lee, Wansoo Kim, Jie Lee, Hoon Heo, and Kyogu Lee. 2021 · 2021
Cited alongside, same era.
Junrui Ni, Liming Wang, Heting Gao, Kaizhi Qian, Yang Zhang, Shiyu Chang, and Mark Hasegawa-Johnson. 2022 · 2022
Later among the works it cites.
Contentvec: An improved self-supervised speech representation by disentangling speakers. In International Conference on Machine Learning . PMLR, 18003–18017
Kaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni, Cheng-I Lai, David Cox, Mark Hasegawa-Johnson, and Shiyu Chang. 2022 · 2022
Later among the works it cites.
Bag of tricks for unsupervised text-to-speech. In The Eleventh International Conference on Learning Representations
Yi Ren, Chen Zhang, and YAN Shuicheng. 2022 · 2022
Later among the works it cites.
A comparison of discrete and soft speech units for improved voice conversion. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6562–6566
Benjamin van Niekerk, Marc-André Carbonneau, Julian Zaïdi, Matthew Baas, Hugo Seuté, and Herman Kamper. 2022 · 2022
Later among the works it cites.
Yixuan Zhou, Changhe Song, Xiang Li, Luwen Zhang, Zhiyong Wu, Yanyao Bian, Dan Su, and Helen Meng. 2022 · 2022
Later among the works it cites.
Cross-lingual multi-speaker speech synthesis with limited bilingual training data
Zexin Cai, Yaogen Yang, and Ming Li. 2023 · 2023
Later among the works it cites.
Diff-HierVC: Diffusion-based Hierarchical Voice Conversion with Robust Pitch Generation and Masked Prior for Zero-shot Speaker Adaptation
Ha-Yeong Choi, Sang-Hoon Lee, and Seong-Whan Lee. 2023 · 2023
Later among the works it cites.
Using joint training speaker encoder with consistency loss to achieve cross-lingual voice conversion and expressive voice conversion. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 1–8
Houjian Guo, Chaoran Liu, Carlos Toshinori Ishi, and Hiroshi Ishiguro. 2023 · 2023
Later among the works it cites.
Mega-TTS 2: Zero-Shot Text-to-Speech with Arbitrary Length Speech Prompts
Ziyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He, Chen Zhang, Zhenhui Ye, Pengfei Wei, Chunfeng Wang, Xiang Yin, Zejun Ma, and Zhou Zhao. 2023 · 2023
Later among the works it cites.
Freevc: Towards high-quality text-free one-shot voice conversion. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1–5
Jingyi Li, Weiping Tu, and Li Xiao. 2023 · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision. In International Conference on Machine Learning . PMLR, 28492–28518
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023 · 2023
Later among the works it cites.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Ziqiang Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, Yu Wu, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Later among the works it cites.