Fetching the paper…
Reading the bibliography…
Voice conversion has gained increasing popularity in many applications of speech synthesis.
Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based f0 extraction: Possible role of a repetitive structure in sounds
H. Kawahara, I. Masuda-Katsuse, and A. de Cheveigné · 1999
Earlier work this paper cites.
On the impact of alignment on voice conversion performance
E. Helander, J. Schwarz, J. Nurminen, H. Silen, and M. Gabbouj · 2008
Earlier work this paper cites.
Supervisory data alignment for text-independent voice conversion
J. Tao, M. Zhang, J. Nurminen, J. Tian, and X. Wang · 2010
Earlier work this paper cites.
Speaking-aid systems using gmm-based voice conversion for electrolaryngeal speech
K. Nakamura, T. Toda, H. Saruwatari, and K. Shikano · 2012
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Auto-encoding variational Bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Accelerating t-SNE using tree-based algorithms
L. van der Maaten · 2014
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Voice conversion from non-parallel corpora using variational auto-encoder
C. Hsu, H. Hwang, Y. Wu, Y. Tsao, and H. Wang · 2016
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric
A. B. L. Larsen, S. K. Sønderby, H. Larochelle, and O. Winther · 2016
Earlier work this paper cites.
WORLD: A vocoder-based high-quality speech synthesis system for real-time applications
M. Morise, F. Yokomori, and K. Ozawa · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. P. Kingma · 2016
Earlier work this paper cites.
WaveNet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Analysis of the voice conversion challenge 2016 evaluation results
M. Wester, Z. Wu, and J. Yamagishi · 2016
Earlier work this paper cites.
A learned representation for artistic style
V. Dumoulin, J. Shlens, and M. Kudlur · 2017
Earlier work this paper cites.
Neural audio synthesis of musical notes with WaveNet autoencoders
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan · 2017
Earlier work this paper cites.
Voice conversion from unaligned corpora using variational autoencoding wasserstein generative adversarial networks
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
P. Isola, J. Zhu, T. Zhou, and A. A. Efros · 2017
Earlier work this paper cites.
Parallel-data-free voice conversion using cycle-consistent adversarial networks
T. Kaneko and H. Kameoka · 2017
Cited alongside, same era.
Sequence-to-sequence voice conversion with similarity metric learned using generative adversarial networks
T. Kaneko, H. Kameoka, K. Hiramatsu, and K. Kashino · 2017
Cited alongside, same era.
Least squares generative adversarial networks
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. P. Smolley · 2017
Cited alongside, same era.
Superseded - cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit, 2017
C. Veaux, J. Yamagishi, and K. MacDonald · 2017
Cited alongside, same era.
Stargan: Unified generative adversarial networks for multi-domain image-to-image translation
Y. Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo · 2018
Cited alongside, same era.
StarGAN-VC2: Rethinking conditional methods for stargan-based voice conversion
T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo · 2019
Later among the works it cites.
MelGAN: Generative adversarial networks for conditional waveform synthesis
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville · 2019
Later among the works it cites.
Few-shot unsupervised image-to-image translation
M.-Y. Liu, X. Huang, A. Mallya, T. Karras, T. Aila, J. Lehtinen, and J. Kautz · 2019
Later among the works it cites.
Waveglow: A flow-based generative network for speech synthesis
R. Prenger, R. Valle, and B. Catanzaro · 2019
Later among the works it cites.
AutoVC: Zero-shot voice style transfer with only autoencoder loss
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson · 2019
Later among the works it cites.
Faceforensics++: Learning to detect manipulated facial images
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
High-quality nonparallel voice conversion based on cycle-consistent adversarial network
F. Fang, J. Yamagishi, I. Echizen, and J. Lorenzo-Trueba · 2018
Cited alongside, same era.
Efficient neural audio synthesis
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu · 2018
Cited alongside, same era.
StarGAN-VC: Non-parallel many-to-many voice conversion using star generative adversarial networks
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo · 2018
Cited alongside, same era.
CycleGAN-VC: Non-parallel voice conversion using cycle-consistent adversarial networks
T. Kaneko and H. Kameoka · 2018
Cited alongside, same era.
Glow: Generative flow with invertible 1x1 convolutions
D. P. Kingma and P. Dhariwal · 2018
Cited alongside, same era.
WaveNet vocoder with limited training data for voice conversion
L.-J. Liu, Z.-H. Ling, Y. Jiang, M. Zhou, and L.-R. Dai · 2018
Cited alongside, same era.
The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, and Z. Ling · 2018
Cited alongside, same era.
A. Rossler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nießner · 2019
Later among the works it cites.
Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion
J. Serrà, S. Pascual, and C. Segura Perales · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in NLP
E. Strubell, A. Ganesh, and A. McCallum · 2019
Later among the works it cites.
High fidelity speech synthesis with adversarial networks
M. Bińkowski, J. Donahue, S. Dieleman, A. Clark, E. Elsen, N. Casagrande, L. C. Cobo, and K. Simonyan · 2020
Later among the works it cites.
DDSP: Differentiable digital signal processing
J. Engel, L. Hantrakul, C. Gu, and A. Roberts · 2020
Later among the works it cites.
Analyzing and improving the image quality of stylegan
T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila · 2020
Later among the works it cites.
Deepfake detection: Current challenges and next steps
S. Lyu · 2020
Later among the works it cites.
Deepsonar: Towards effective and robust detection of ai-synthesized fake voices
R. Wang, F. Juefei-Xu, Y. Huang, Q. Guo, X. Xie, L. Ma, and Y. Liu · 2020
Later among the works it cites.
One-shot voice conversion by vector quantization
D.-Y. Wu and H.-y. Lee · 2020
Later among the works it cites.
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
R. Yamamoto, E. Song, and J.-M. Kim · 2020
Later among the works it cites.
Voice conversion by cascading automatic speech recognition and text-to-speech synthesis with prosody transfer
J.-X. Zhang, L.-J. Liu, Y.-N. Chen, Y.-J. Hu, Y. Jiang, Z.-H. Ling, and L.-R. Dai · 2020
Later among the works it cites.
An overview of voice conversion and its challenges: From statistical modeling to deep learning
B. Sisman, J. Yamagishi, S. King, and H. Li · 2021
Closest in time.