Fetching the paper…
Reading the bibliography…
Singing Voice Conversion (SVC) is a technique that enables any singer to perform any song.
“Mel-cepstral distance measure for objective speech quality assessment,”
Robert Kubichek, · 1993
Earlier work this paper cites.
“Application of voice conversion for cross-language rap singing transformation,”
Oytun Türk, Osman Büyük, Ali Haznedaroglu, and Levent Mustafa Arslan, · 2009
Earlier work this paper cites.
“Statistical singing voice conversion with direct waveform modification based on the spectrum differential,”
Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, and Satoshi Nakamura, · 2014
Earlier work this paper cites.
“Statistical singing voice conversion based on direct waveform modification with global variance,”
Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, and Satoshi Nakamura, · 2015
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Earlier work this paper cites.
“Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,”
Lifa Sun, Kun Li, Hao Wang, Shiyin Kang, and Helen M. Meng, · 2016
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Introducing Parselmouth: A Python interface to Praat,”
Yannick Jadoul, Bill Thompson, and Bart de Boer, · 2018
Earlier work this paper cites.
“Unsupervised singing voice conversion,”
Eliya Nachmani and Lior Wolf, · 2019
Earlier work this paper cites.
“Singing voice conversion with non-parallel data,”
Xin Chen, Wei Chu, Jinxi Guo, and Ning Xu, · 2019
Earlier work this paper cites.
“Autovc: Zero-shot voice style transfer with only autoencoder loss,”
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson, · 2019
Earlier work this paper cites.
“Zero-shot singing voice conversion,”
Shahan Nercessian, · 2020
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Earlier work this paper cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang, · 2020
Cited alongside, same era.
“Pitchnet: Unsupervised singing voice conversion with pitch adversarial network,”
Chengqi Deng, Chengzhu Yu, Heng Lu, Chao Weng, and Dong Yu, · 2020
Cited alongside, same era.
“Ppg-based singing voice conversion with adversarial representation learning,”
Zhonghao Li, Benlai Tang, Xiang Yin, Yuan Wan, Ling Xu, Chen Shen, and Zejun Ma, · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit,”
Zhuoyuan Yao, Di Wu, Xiong Wang, Binbin Zhang, Fan Yu, Chao Yang, Zhendong Peng, Xiaoyu Chen, Lei Xie, and Xin Lei, · 2021
Cited alongside, same era.
“Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis,”
Yu Wang, Xinsheng Wang, Pengcheng Zhu, Jie Wu, Hanzhao Li, Heyang Xue, Yongmao Zhang, Lei Xie, and Mengxiao Bi, · 2022
Later among the works it cites.
“M4singer: A multi-style, multi-singer and musical score provided mandarin singing corpus,”
Lichao Zhang, Ruiqi Li, Shoutong Wang, Liqun Deng, Jinglin Liu, Yi Ren, Jinzheng He, Rongjie Huang, Jieming Zhu, Xiao Chen, and Zhou Zhao, · 2022
Later among the works it cites.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei, · 2022
Later among the works it cites.
“The singing voice conversion challenge 2023,”
Wen-Chin Huang, Lester Phillip Violeta, Songxiang Liu, Jiatong Shi, Yusuke Yasuda, and Tomoki Toda, · 2023
Closest in time.
“Self-supervised representations for singing voice conversion,”
Tejas Jayashankar, Jilong Wu, Leda Sari, David Kant, Vimal Manohar, and Qing He, · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Diffsvc: A diffusion probabilistic model for singing voice conversion,”
Songxiang Liu, Yuewen Cao, Dan Su, and Helen Meng, · 2021
Cited alongside, same era.
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
Jaehyeon Kim, Jungil Kong, and Juhee Son, · 2021
Cited alongside, same era.
“Diffwave: A versatile diffusion model for audio synthesis,”
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro, · 2021
Cited alongside, same era.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2021
Cited alongside, same era.
“High-resolution image synthesis with latent diffusion models,”
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer, · 2022
Cited alongside, same era.
“A comparative study of self-supervised speech representation based voice conversion,”
Wen-Chin Huang, Shu-Wen Yang, Tomoki Hayashi, and Tomoki Toda, · 2022
Cited alongside, same era.
“Robust one-shot singing voice conversion,”
Naoya Takahashi, Mayank Kumar Singh, and Yuki Mitsufuji, · 2022
Cited alongside, same era.
Closest in time.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever, · 2023
Closest in time.
“Pronunciation dictionary-free multilingual speech synthesis using learned phonetic representations,”
Chang Liu, Zhen-Hua Ling, and Ling-Hui Chen, · 2023
Closest in time.
“Vits-based singing voice conversion leveraging whisper and multi-scale f0 modeling,”
Ziqian Ning, Yuepeng Jiang, Zhichao Wang, Bin Zhang, and Lei Xie, · 2023
Closest in time.
“Mega-tts: Zero-shot text-to-speech at scale with intrinsic inductive bias,”
Ziyue Jiang, Yi Ren, Zhenhui Ye, Jinglin Liu, Chen Zhang, Qian Yang, Shengpeng Ji, Rongjie Huang, Chunfeng Wang, Xiang Yin, Zejun Ma, and Zhou Zhao, · 2023
Closest in time.
“Amphion: An open-source audio, music and speech generation toolkit,”
Xueyao Zhang, Liumeng Xue, Yuancheng Wang, Yicheng Gu, Xi Chen, Zihao Fang, Haopeng Chen, Lexiao Zou, Chaoren Wang, Jun Han, Kai Chen, Haizhou Li, and Zhizheng Wu, · 2023
Closest in time.
“Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,”
Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu, Chen Yang, Jiaqi Li, Peiyang Shi, Yuancheng Wang, Kai Chen, Pengyuan Zhang, and Zhizheng Wu, · 2024
Closest in time.
“Singfake: Singing voice deepfake detection,”
Yongyi Zang, You Zhang, Mojtaba Heydari, and Zhiyao Duan, · 2024
Closest in time.
“Neural concatenative singing voice conversion: rethinking concatenation-based approach for one-shot singing voice conversion,”
Binzhu Sha, Xu Li, Zhiyong Wu, Ying Shan, and Helen Meng, · 2024
Closest in time.
“Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,”
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian, · 2024
Closest in time.