Fetching the paper…
Reading the bibliography…
Singing voice conversion (SVC) is one promising technique which can enrich the way of human-computer interaction by endowing a computer the ability to produce high-fidelity and expressive singing voice.
“Estimation of non-normalized statistical models by score matching,”
Aapo Hyvärinen, · 2005
Earlier work this paper cites.
“One-to-many and many-to-one voice conversion based on eigenvoices,”
Tomoki Toda, Yamato Ohtani, and Kiyohiro Shikano, · 2007
Earlier work this paper cites.
“Fast and reliable f0 estimation method based on the period extraction of vocal fold vibration of singing voice and speech,”
Masanori Morise, Hideki Kawahara, and Haruhiro Katayose, · 2009
Earlier work this paper cites.
“Speech signal processing toolkit (sptk),”
Keiichi Tokuda, Keiichiro Oura, and et al., · 2009
Earlier work this paper cites.
“Statistical singing voice conversion with direct waveform modification based on the spectrum differential,”
Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, and Satoshi Nakamura, · 2014
Earlier work this paper cites.
“Deep unsupervised learning using nonequilibrium thermodynamics,”
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli, · 2015
Earlier work this paper cites.
“Statistical singing voice conversion based on direct waveform modification with global variance,”
Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, and Satoshi Nakamura, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Earlier work this paper cites.
“Deconvolution and checkerboard artifacts,”
Augustus Odena, Vincent Dumoulin, and Chris Olah, · 2016
Earlier work this paper cites.
“Searching for activation functions,”
Prajit Ramachandran, Barret Zoph, and Quoc V Le, · 2017
Cited alongside, same era.
“Deep-fsmn for large vocabulary continuous speech recognition,”
Shiliang Zhang, Ming Lei, Zhijie Yan, and Lirong Dai, · 2018
Cited alongside, same era.
“Efficient neural audio synthesis,”
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu, · 2018
Cited alongside, same era.
“Unsupervised singing voice conversion,”
Eliya Nachmani and Lior Wolf, · 2019
Cited alongside, same era.
“Singan: Singing voice conversion with generative adversarial networks,”
Berrak Sisman, Karthika Vijayan, Minghui Dong, and Haizhou Li, · 2019
Cited alongside, same era.
“Generative modeling by estimating gradients of the data distribution,”
“Diffwave: A versatile diffusion model for audio synthesis,”
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro, · 2020
Later among the works it cites.
“Wavegrad: Estimating gradients for waveform generation,”
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan, · 2020
Later among the works it cites.
“Singing voice conversion with disentangled representations of singer and vocal technique using variational autoencoders,”
Yin-Jyun Luo, Chin-Cheng Hsu, Kat Agres, and Dorien Herremans, · 2020
Later among the works it cites.
“Vaw-gan for singing voice conversion with non-parallel training data,”
Junchen Lu, Kun Zhou, Berrak Sisman, and Haizhou Li, · 2020
Later among the works it cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang Song and Stefano Ermon, · 2019
Cited alongside, same era.
“Singing voice conversion with non-parallel data,”
Xin Chen, Wei Chu, Jinxi Guo, and Ning Xu, · 2019
Cited alongside, same era.
“Pitchnet: Unsupervised singing voice conversion with pitch adversarial network,”
Chengqi Deng, Chengzhu Yu, Heng Lu, Chao Weng, and Dong Yu, · 2020
Cited alongside, same era.
“Unsupervised cross-domain singing voice conversion,”
Adam Polyak, Lior Wolf, Yossi Adi, and Yaniv Taigman, · 2020
Cited alongside, same era.
“Denoising diffusion probabilistic models,”
Jonathan Ho, Ajay Jain, and Pieter Abbeel, · 2020
Cited alongside, same era.
Later among the works it cites.
“Fastsvc: Fast cross-domain singing voice conversion with feature-wise linear modulation,”
Songxiang Liu, Yuewen Cao, Na Hu, Dan Su, and Helen Meng, · 2021
Closest in time.
“Ppg-based singing voice conversion with adversarial representation learning,”
Zhonghao Li, Benlai Tang, Xiang Yin, Yuan Wan, Ling Xu, Chen Shen, and Zejun Ma, · 2021
Closest in time.
“Denoising diffusion implicit models,”
Jiaming Song, Chenlin Meng, and Stefano Ermon, · 2021
Closest in time.
“Score-based generative modeling through stochastic differential equations,”
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole, · 2021
Closest in time.