Fetching the paper…
Reading the bibliography…
We propose a novel architecture and improved training objectives for non-parallel voice conversion.
D. Griffin and J. Lim, “Signal estimation from modified short time fourier transform,”
1984
Earlier work this paper cites.
F. Mamalet and C. Garcia, “Simplifying convnets for fast learning,”
2012
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2015
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “World: A vocoder-based high-quality speech synthesis system for real-time applications,”
2016
Earlier work this paper cites.
S. Mohammadi and A. Kain, “An overview of voice conversion systems,”
2017
Earlier work this paper cites.
J. Zhu, T. Park, P. Isola, and A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,”
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,”
2017
Earlier work this paper cites.
X. Mao, Q. Li, H. Xie, R. Lau, Z. Wang, and S. Smolley, “Least squares generative adversarial networks,”
2017
Earlier work this paper cites.
J. Lim and J. Ye, “Geometric gan,”
2017
Earlier work this paper cites.
P. Isola, J. Zhu, T. Zhou, and A. Efros, “Image-to-image translation with conditional adversarial networks,”
2017
Earlier work this paper cites.
Y. Taigman, A. Polyak, and L. Wolf, “Towards principled methods for training generative adversarial networks,”
2017
Earlier work this paper cites.
M. Arjovsky and L. Bottou, “Towards principled methods for training generative adversarial networks,”
2017
Earlier work this paper cites.
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo, “Acvae-vc: Non-parallel many-to-many voice conversion with auxiliary classifier variational autoencoder,”
2018
Cited alongside, same era.
Y. Choi, M. Choi, M. Kim, J. Ha, S. Kim, and J. Choo, “Stargan: Unified generative adversarial networks for multi-domain image-to-image translation,”
2018
Cited alongside, same era.
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo, “Stargan: Unified generative adversarial networks for multi-domain image-to-image translation,”
2018
Cited alongside, same era.
——, “Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks,”
2018
Cited alongside, same era.
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida, “Spectral normalization for generative adversarial networks,”
2018
Cited alongside, same era.
2019
Later among the works it cites.
Z. Huang, X. Wang, L. Huang, C. Huang, Y. Wei, and W. Liu, “Ccnet: Criss-cross attention for semantic segmentation,”
2019
Later among the works it cites.
F. Wu, A. Fan, A. Baevski, Y. Dauphin, and M. Auli, “Pay less attention with lightweight and dynamic convolutions,”
2019
Later among the works it cites.
2019
Later among the works it cites.
K. M. J. Yamagishi, C. Veaux, “Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),”
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Tanaka, H. Kameoka, T. Kaneko, and N. Hojo, “Atts2s-vc: Sequence-to-sequence voice conversion with attention and context preservation mechanisms,”
2019
Cited alongside, same era.
2019
Cited alongside, same era.
P. Tobing, Y. Wu, T. Hayashi, K. Kobayashi, and T. Toda, “Non-parallel voice conversion with cyclic variational autoencoder,”
2019
Cited alongside, same era.
T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo, “Stargan-vc2: Rethinking conditional methods for stargan-based voice conversion,”
2019
Cited alongside, same era.
T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo, “Cycleganvc2: Improved cyclegan-based non-parallel voice conversion,”
2019
Cited alongside, same era.
K. Kumar, R. Kumar, T. Boissiere, L. Gestin, W. Teoh, J. Sotelo, A. Brébisson, Y. Bengio, and A. Courville, “Melgan: Generative adversarial networks for conditional waveform synthesis,”
2019
Cited alongside, same era.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,”
2019
Cited alongside, same era.
Later among the works it cites.
——, “Cyclegan-vc3: Examining and improving cyclegan-vcs for mel-spectrogram conversion,”
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,”
2020
Later among the works it cites.