Fetching the paper…
Reading the bibliography…
We present a fully convolutional wav-to-wav network for converting between speakers' voices, without relying on text.
J. S. Garofolo
1993
Earlier work this paper cites.
T. J. Hazen, W. Shen, and C. White, “Query-by-example spoken term detection using phonetic posteriorgram templates,” in
2009
Earlier work this paper cites.
D. Erro, A. Moreno, and A. Bonafonte, “INCA algorithm for training voice conversion systems from nonparallel corpora,”
2010
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2010
Earlier work this paper cites.
F. P. Ribeiro, D. Florencio, C. Zhang, and M. Seltzer, “CROWDMOS: an approach for crowdsourcing mean opinion score studies,” in
2011
Earlier work this paper cites.
P. Song, Y. Jin, W. Zheng, and L. Zhao, “Text-independent voice conversion using speaker model alignment method from non-parallel speech,”
2014
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-Encoding Variational Bayes,” in
2014
Earlier work this paper cites.
X. Zhang, J. Trmal, D. Povey, and S. Khudanpur, “Improving deep neural network acoustic models using generalized maxout networks,” in
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: An ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
F.-L. Xie, F. K. Soong, and H. Li, “A KL divergence and DNN-based approach to voice conversion without parallel training sentences,” in
2016
Earlier work this paper cites.
2016
Cited alongside, same era.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: A vocoder-based high-quality speech synthesis system for real-time applications,”
2016
Cited alongside, same era.
2016
Cited alongside, same era.
T. Kinnunen, L. Juvela, P. Alku, and J. Yamagishi, “Non-parallel voice conversion using i-vector PLDA: towards unifying speaker verification and transformation,” in
2017
Cited alongside, same era.
M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein Generative Adversarial Networks,” in
2017
Cited alongside, same era.
Y.-C. Wu
2018
Later among the works it cites.
Google Cloud TTS robot, “US-WaveNet-E,”
2018
Later among the works it cites.
2018
Later among the works it cites.
Y. Taigman, L. Wolf, A. Polyak, and E. Nachmani, “VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop,” in
2018
Later among the works it cites.
E. Nachmani, A. Polyak, Y. Taigman, and L. Wolf, “Fitting new speakers based on a short untranscribed sample,”
2018
Later among the works it cites.
X. Tian, E. S. Chng, and H. Li, “A Vocoder-free WaveNet Voice Conversion with Non-Parallel Data,”
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural Discrete Representation Learning,” in
2017
Cited alongside, same era.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, “Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders,” in
2017
Cited alongside, same era.
C. Veaux,
2017
Cited alongside, same era.
L.-J. Liu
2018
Cited alongside, same era.
Y. Saito
2018
Cited alongside, same era.
Closest in time.
N. Mor, L. Wolf, A. Polyak, and Y. Taigman, “A Universal Music Translation Network,” in
2019
Closest in time.
Y. Adi, N. Zeghidour, R. Collobert, N. Usunier, V. Liptchinsky, and G. Synnaeve, “To reverse the gradient or not: An empirical comparison of adversarial and multi-task learning in speech recognition,” in
2019
Closest in time.
M. Ravanelli, T. Parcollet, and Y. Bengio, “The pytorch-kaldi speech recognition toolkit,” in
2019
Closest in time.