Fetching the paper…
Reading the bibliography…
The diffusion-based Singing Voice Conversion (SVC) methods have achieved remarkable performances, producing natural audios with high similarity to the target timbre.
Perceptual evaluation of speech quality (PESQ), an objective method for end-to-end speech quality assessment of narrowband telephone networks and speech codecs
Intl. Telecommunications Union (ITU-T) · 2001
Earlier work this paper cites.
Fast and reliable f0 estimation method based on the period extraction of vocal fold vibration of singing voice and speech
Masanori Morise, Hideki Kawahara, and Haruhiro Katayose · 2009
Earlier work this paper cites.
Statistical singing voice conversion with direct waveform modification based on the spectrum differential
Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, and Satoshi Nakamura · 2014
Earlier work this paper cites.
A wavenet for speech denoising
Dario Rethage, Jordi Pons, and Xavier Serra · 2018
Earlier work this paper cites.
Unsupervised singing voice conversion
Eliya Nachmani and Lior Wolf · 2019
Earlier work this paper cites.
Pitchnet: Unsupervised singing voice conversion with pitch adversarial network
Chengqi Deng, Chengzhu Yu, Heng Lu, Chao Weng, and Dong Yu · 2020
Earlier work this paper cites.
Unsupervised cross-domain singing voice conversion
Adam Polyak, Lior Wolf, Yossi Adi, and Yaniv Taigman · 2020
Cited alongside, same era.
Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus
Rongjie Huang, Feiyang Chen, Yi Ren, Jinglin Liu, Chenye Cui, and Zhou Zhao · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2021
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine · 2022
Cited alongside, same era.
ContentVec: An improved self-supervised speech representation by disentangling speakers
Kaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni, Cheng-I Lai, David Cox, Mark Hasegawa-Johnson, and Shiyu Chang · 2022
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Later among the works it cites.
Consistency models
Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever · 2023
Later among the works it cites.
Softvc vits singing voice conversion
SVC-Develop-Team · 2023
Later among the works it cites.
Comospeech: One-step speech and singing voice synthesis via consistency model
Zhen Ye, Wei Xue, Xu Tan, Jie Chen, Qifeng Liu, and Yike Guo · 2023
Later among the works it cites.
Leveraging content-based features from multiple acoustic models for singing voice conversion
Xueyao Zhang, Yicheng Gu, Haopeng Chen, Zihao Fang, Lexiao Zou, Liumeng Xue, and Zhizheng Wu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lichao Zhang, Ruiqi Li, Shoutong Wang, Liqun Deng, Jinglin Liu, Yi Ren, Jinzheng He, Rongjie Huang, Jieming Zhu, Xiao Chen, et al · 2022
Cited alongside, same era.
Statistical singing voice conversion based on direct waveform modification and its parameter generation algorithms
Kazuhiro Kobayashi, Tomoki Toda, and Satoshi Nakamura
Cited in the paper.
Statistical singing voice conversion based on direct waveform modification with global variance
Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, and Satoshi Nakamura
Cited in the paper.
FastSVC: Fast cross-domain singing voice conversion with feature-wise linear modulation
Songxiang Liu, Yuewen Cao, Na Hu, Dan Su, and Helen Meng
Cited in the paper.
Diffsvc: A diffusion probabilistic model for singing voice conversion
Songxiang Liu, Yuewen Cao, Dan Su, and Helen Meng
Cited in the paper.