Fetching the paper…
Reading the bibliography…
Voice conversion is a common speech synthesis task which can be solved in different ways depending on a particular real-world scenario.
Statistics of Random Processes , volume 5 of Stochastic Modelling and Applied Probability
Robert S. Liptser and Albert N. Shiryaev · 1978
Earlier work this paper cites.
Reverse-time Diffusion Equation Models
Brian D.O. Anderson · 1982
Earlier work this paper cites.
Numerical Solution of Stochastic Differential Equations , volume 23 of Stochastic Modelling and Applied Probability
Peter E. Kloeden and Eckhard Platen · 1992
Earlier work this paper cites.
Exact Simulation of Diffusions
Alexandros Beskos and Gareth O. Roberts · 2005
Earlier work this paper cites.
Estimation of Non-Normalized Statistical Models by Score Matching
Aapo Hyvärinen · 2005
Earlier work this paper cites.
A Connection Between Score Matching and Denoising Autoencoders
Pascal Vincent · 2011
Earlier work this paper cites.
Librispeech: An ASR Corpus Based on Public Domain Audio Books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Unbiased Simulation of Stochastic Differential Equations
Pierre Henry-Labordère, Xiaolu Tan, and Nizar Touzi · 2017
Earlier work this paper cites.
Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger · 2017
Earlier work this paper cites.
Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis
Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, and Yonghui Wu · 2018
Earlier work this paper cites.
Non-Parallel Voice Conversion Using Variational Autoencoders Conditioned by Phonetic Posteriorgrams and D-Vectors
Yuki Saito, Yusuke Ijima, Kyosuke Nishida, and Shinnosuke Takamichi · 2018
Earlier work this paper cites.
One-Shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization
Ju-Chieh Chou and Hung-yi Lee · 2019
Earlier work this paper cites.
AutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson · 2019
Cited alongside, same era.
CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92), 2019
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald · 2019
Cited alongside, same era.
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Heiga Zen, Rob Clark, Ron J. Weiss, Viet Dang, Ye Jia, Yonghui Wu, Yu Zhang, and Zhifeng Chen · 2019
Cited alongside, same era.
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Attention-Based Speaker Embeddings for One-Shot Voice Conversion
Assem-VC: Realistic Voice Conversion by Assembling Modern Speech Synthesis Techniques, 2021
Kang-wook Kim, Seung-won Park, and Myun-chul Joe · 2021
Closest in time.
Variational Diffusion Models, 2021
Diederik P. Kingma, Tim Salimans, Ben Poole, and Jonathan Ho · 2021
Closest in time.
On Fast Sampling of Diffusion Probabilistic Models
Zhifeng Kong and Wei Ping · 2021
Closest in time.
DiffWave: A Versatile Diffusion Model for Audio Synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2021
Closest in time.
PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Driven Adaptive Prior, 2021
Sang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan, Chang Liu, Qi Meng, Tao Qin, Wei Chen, Sungroh Yoon, and Tie-Yan Liu · 2021
Closest in time.
FragmentVC: Any-To-Any Voice Conversion by End-To-End Extracting and Fusing Fine-Grained Voice Fragments with Attention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tatsuma Ishihara and Daisuke Saito · 2020
Cited alongside, same era.
VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics, 2020
Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo, and Shogo Seki · 2020
Cited alongside, same era.
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Cited alongside, same era.
Improved Zero-Shot Voice Conversion Using Explicit Conditioning Signals
Shahan Nercessian · 2020
Cited alongside, same era.
F0-Consistent Many-To-Many Non-Parallel Voice Conversion Via Conditional Autoencoder
Kaizhi Qian, Zeyu Jin, Mark Hasegawa-Johnson, and Gautham J. Mysore · 2020
Cited alongside, same era.
VQVC+: One-Shot Voice Conversion by Vector Quantization and U-Net Architecture
Da-Yi Wu, Yen-Hao Chen, and Hung-yi Lee · 2020
Cited alongside, same era.
Again-VC: A One-Shot Voice Conversion Using Activation Guidance and Adaptive Instance Normalization
Yen-Hao Chen, D. Wu, Tsung-Han Wu, and Hung yi Lee · 2021
Cited alongside, same era.
Yist Y. Lin, Chung-Ming Chien, Jheng-Hao Lin, Hung-yi Lee, and Lin-Shan Lee · 2021
Closest in time.
Many-to-Many Voice Conversion Based Feature Disentanglement Using Variational Autoencoder
Manh Luong and Viet Anh Tran · 2021
Closest in time.
Improved Denoising Diffusion Probabilistic Models
Alexander Quinn Nichol and Prafulla Dhariwal · 2021
Closest in time.
Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov · 2021
Closest in time.
Noise Estimation for Generative Diffusion Models, 2021
Robin San-Roman, Eliya Nachmani, and Lior Wolf · 2021
Closest in time.
VQMIVC: Vector Quantization and Mutual Information-Based Unsupervised Speech Representation Disentanglement for One-Shot Voice Conversion
Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen, Xunying Liu, and Helen Meng · 2021
Closest in time.
Learning to Efficiently Sample from Diffusion Probabilistic Models, 2021
Daniel Watson, Jonathan Ho, Mohammad Norouzi, and William Chan · 2021
Closest in time.