Fetching the paper…
Reading the bibliography…
Recent language model (LM) advancements have showcased impressive zero-shot voice conversion (VC) performance.
The EMIME bilingual database
Mirjam Wester. 2010 · 2010
Earlier work this paper cites.
Phonetic posteriorgrams for many-to-one voice conversion without parallel data training
Lifa Sun, K. Li, Hao Wang, Shiyin Kang, and H. Meng. 2016 · 2016
Earlier work this paper cites.
One-shot voice conversion by separating speaker and content representations with instance normalization
Juchieh Chou and Hungyi Lee. 2019 · 2019
Earlier work this paper cites.
An Unsupervised Autoregressive Model for Speech Representation Learning
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass. 2019 · 2019
Earlier work this paper cites.
Streaming end-to-end speech recognition with joint ctc-attention based models
Niko Moritz, Takaaki Hori, and Jonathan Le Roux. 2019 · 2019
Earlier work this paper cites.
Autovc: Zero-shot voice style transfer with only autoencoder loss
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson. 2019 · 2019
Earlier work this paper cites.
ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck. 2020 · 2020
Earlier work this paper cites.
The NUS & NWPU system for Voice Conversion Challenge 2020
Xiaohai Tian, Zhichao Wang, Shan Yang, Xinyong Zhou, Hongqiang Du, Yi Zhou, Mingyang Zhang, Kun Zhou, Berrak Sisman, Lei Xie, and Haizhou Li. 2020 · 2020
Earlier work this paper cites.
Neural analysis and synthesis: Reconstructing speech from self-supervised representations
Hyeong-Seok Choi, Juheon Lee, Wansoo Kim, Jie Lee, Hoon Heo, and Kyogu Lee. 2021 · 2021
Earlier work this paper cites.
w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu. 2021 · 2021
Earlier work this paper cites.
Contrastive predictive coding supported factorized variational autoencoder for unsupervised learning of disentangled speech representations
Janek Ebbers, Michael Kuhlmann, Tobias Cord-Landwehr, and Reinhold Haeb-Umbach. 2021 · 2021
Earlier work this paper cites.
MediumVC: Any-to-any voice conversion using synthetic specific-speaker speeches as intermedium features
Yewei Gu, Zhenyu Zhang, Xiaowei Yi, and Xianfeng Zhao. 2021 · 2021
Earlier work this paper cites.
Didispeech: A large scale mandarin speech corpus
Tingwei Guo, Cheng Wen, Dongwei Jiang, Ne Luo, Ruixiong Zhang, Shuaijiang Zhao, Wubo Li, Cheng Gong, Wei Zou, Kun Han, and Xiangang Li. 2021 · 2021
Cited alongside, same era.
Fasts2s-vc: Streaming non-autoregressive sequence-to-sequence voice conversion
Hirokazu Kameoka, Kou Tanaka, and Takuhiro Kaneko. 2021 · 2021
Cited alongside, same era.
AISHELL-3: A multi-speaker mandarin tts corpus
Yao Shi, Hui Bu, Xin Xu, Shaoji Zhang, and Ming Li. 2021 · 2021
Cited alongside, same era.
Vqmivc: Vector quantization and mutual information-based unsupervised speech representation disentanglement for one-shot voice conversion
Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen, Xunying Liu, and Helen Meng. 2021 · 2021
Cited alongside, same era.
WeNet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit
Zhuoyuan Yao, Di Wu, Xiong Wang, Binbin Zhang, Fan Yu, Chao Yang, Zhendong Peng, Xiaoyu Chen, Lei Xie, and Xin Lei. 2021 · 2021
Cited alongside, same era.
Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al. 2022 · 2022
Later among the works it cites.
Audiolm: A language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour. 2023 · 2023
Later among the works it cites.
Speak, Read and Prompt: High-fidelity text-to-speech with minimal supervision
Eugene Kharitonov, Damien Vincent, Zalán Borsos, Raphaël Marinier, Sertan Girgin, Olivier Pietquin, Matt Sharifi, Marco Tagliasacchi, and Neil Zeghidour. 2023 · 2023
Later among the works it cites.
Fast-u2++: Fast and accurate end-to-end speech recognition in joint ctc/attention frames
Chengdong Liang, Xiao-Lei Zhang, BinBin Zhang, Di Wu, Shengqiang Li, Xingchen Song, Zhendong Peng, and Fuping Pan. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Retriever: Learning content-style representation as a token-level bipartite graph
Dacheng Yin, Xuanchi Ren, Chong Luo, Yuwang Wang, Zhiwei Xiong, and Wenjun Zeng. 2021 · 2021
Cited alongside, same era.
SoundStream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. 2021 · 2021
Cited alongside, same era.
Streaming non-autoregressive model for any-to-many voice conversion
Ziyi Chen, Haoran Miao, and Pengyuan Zhang. 2022 · 2022
Cited alongside, same era.
An investigation of streaming non-autoregressive sequence-to-sequence voice conversion
Tomoki Hayashi, Kazuhiro Kobayashi, and Tomoki Toda. 2022 · 2022
Cited alongside, same era.
Revisiting over-smoothness in text to speech
Yi Ren, Xu Tan, Tao Qin, Zhou Zhao, and Tie-Yan Liu. 2022 · 2022
Cited alongside, same era.
Streamable speech representation disentanglement and multi-level prosody modeling for live one-shot voice conversion
Haoquan Yang, Liqun Deng, Yu Ting Yeung, Nianzu Zheng, and Yong Xu. 2022 · 2022
Cited alongside, same era.
Add 2022: the first audio deep synthesis detection challenge
Jiangyan Yi, Ruibo Fu, Jianhua Tao, Shuai Nie, Haoxin Ma, Chenglong Wang, Tao Wang, Zhengkun Tian, Ye Bai, Cunhan Fan, Shan Liang, Shiming Wang, Shuai Zhang, Xinrui Yan, Le Xu, Zhengqi Wen, and Haizhou Li. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Audiodec: An open-source streaming high-fidelity neural audio codec
Yi-Chiao Wu, Israel D Gebru, Dejan Marković, and Alexander Richard. 2023 · 2023
Later among the works it cites.
Uniaudio: An audio foundation model toward universal audio generation
Dongchao Yang, Jinchuan Tian, Xu Tan, Rongjie Huang, Songxiang Liu, Xuankai Chang, Jiatong Shi, Sheng Zhao, Jiang Bian, Xixin Wu, et al. 2023 · 2023
Later among the works it cites.
Vec-tok speech: speech vectorization and tokenization for neural speech generation
Xinfa Zhu, Yuanjun Lv, Yi Lei, Tao Li, Wendi He, Hongbin Zhou, Heng Lu, and Lei Xie. 2023 · 2023
Later among the works it cites.
Dualvc 2: Dynamic masked convolution for unified streaming and non-streaming voice conversion
Ziqian Ning, Yuepeng Jiang, Pengcheng Zhu, Shuai Wang, Jixun Yao, Lei Xie, and Mengxiao Bi. 2024 · 2024
Closest in time.
NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Kai Shen, Zeqian Ju, Xu Tan, Eric Liu, Yichong Leng, Lei He, Tao Qin, sheng zhao, and Jiang Bian. 2024 · 2024
Closest in time.
Dualvc: Dual-mode voice conversion using intra-model knowledge distillation and hybrid predictive coding
Ziqian Ning, Yuepeng Jiang, Pengcheng Zhu, Jixun Yao, Shuai Wang, Lei Xie, and Mengxiao Bi. 2023 · 2067
Closest in time.
Alo-vc: Any-to-any low-latency one-shot voice conversion
Bohan Wang, Damien Ronssin, and Milos Cernak. 2023a · 2077
Closest in time.