Fetching the paper…
Reading the bibliography…
End-to-end models for raw audio generation are a challenge, specially if they have to work with non-parallel data, which is a desirable setup in many situations.
Unsupervised learning
H. B. Barlow · 1989
Earlier work this paper cites.
Supervised factorial learning
A. N. Redlich · 1993
Earlier work this paper cites.
CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit, 2012
C. Veaux, J. Yamagishi, and K. MacDonald · 1994
Earlier work this paper cites.
Higher order statistical decorrelation without information loss
G. Deco and W. Brauer · 1995
Earlier work this paper cites.
Non-parallel training for voice conversion based on a parameter adaptation approach
A. Mouchtaris, J. Van der Spiegel, and P. Mueller · 2006
Earlier work this paper cites.
INCA algorithm for training voice conversion systems from nonparallel corpora
D. Erro, A. Moreno, and A. Bonafonte · 2010
Earlier work this paper cites.
Time durations of phonemes in Polish language for speech and speaker recognition
B. Ziolko and M. Ziolko · 2011
Earlier work this paper cites.
Mixture of factor analyzers using priors from non-parallel speech for voice conversion
Z. Wu, T. Kinnunen, E. S. Chang, and H. Li · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Earlier work this paper cites.
A family of non-parametric density estimation algorithms
E. G. Tabak and C. V. Turner · 2013
Earlier work this paper cites.
NICE: non-linear independent components estimation
L. Dinh, D. Krueger, and Y. Bengio · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
D. J. Rezende and S. Mohamed · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
WaveNet: a generative model for raw audio
A. Van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
A KL divergence and DNN-based approach to voice conversion without parallel training sentences
F.-L. Xie, F. K. Soong, and H. Li · 2016
Earlier work this paper cites.
Analysis of the voice conversion challenge 2016 evaluation results
M. Wester, Z. Wu, and J. Yamagishi · 2016
Earlier work this paper cites.
WORLD: a vocoder-based high-quality speech synthesis system for real-time applications
M. Morise, F. Yokomori, and K. Ozawa · 2016
Earlier work this paper cites.
SampleRNN: an unconditional end-to-end neural audio generation model
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio · 2017
Earlier work this paper cites.
SEGAN: speech enhancement generative adversarial network
S. Pascual, A. Bonafonte, and J. Serrà · 2017
Cited alongside, same era.
An overview of voice conversion systems
S. H. Mohammadi and A. Kain · 2017
Cited alongside, same era.
Voice conversion from unaligned corpora using variational autoencoding wasserstein generative adversarial networks
C. C. Hsu, H. T. Hwang, Y. C. Wu, Y. Tsao, and H. M. Wang · 2017
Cited alongside, same era.
HyperNetworks
D. Ha, A. Dai, and Q. V. Le · 2017
Cited alongside, same era.
Density estimation using Real NVP
L. Dinh, J. Sohl-Dickstein, and S. Bengio · 2017
Cited alongside, same era.
Non-parallel voice conversion using i-vector PLDA: towards unifying speaker verification and transformation
T. Kinnunen, L. Juvela, P. Alku, and J. Yamagishi · 2017
Cited alongside, same era.
Adaptive batch normalization for practical domain adaptation
Y. Li, N. Wang, J. Shi, H. Hou, and J. Liu · 2018
Later among the works it cites.
Neural voice cloning with a few samples
S. O. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou · 2018
Later among the works it cites.
Conditional end-to-end audio transforms
A. Haque, M. Guo, and P. Verma · 2018
Later among the works it cites.
TzK Flow - Conditional Generative Model
M. Livne and D. J. Fleet · 2018
Later among the works it cites.
Conditional recurrent flow: conditional generation of longitudinal samples with applications to neuroimaging
S. J. Hwang and W. H. Kim · 2018
Later among the works it cites.
Depthwise separable convolutions for neural machine translation
L. Kaiser, A. N. Gomez, and F. Chollet · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural discrete representation learning
A. Van den Oord, O. Vinyals, and K. Kavukcuoglu · 2017
Cited alongside, same era.
Neural audio synthesis of musical notes with WaveNet autoencoders
J. Engel, C. Resnick, A. Roberts, S. Dieleman, D. Eck, K. Simonyan, and M. Norouzi · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Cited alongside, same era.
Deep voice: real-time neural text-to-speech
S. O. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman, S. Sengupta, and M. Shoeybi · 2017
Cited alongside, same era.
Efficient neural audio synthesis
N. Kalchbrenner, E. Elsen, K. Simonyan, N. Casagrande, E. Lockhart, F. Stimberg, A. Van den Oord, S. Dieleman, and K. Kavukcuoglu · 2018
Cited alongside, same era.
WaveGlow: a flow-based generative network for speech synthesis
R. Prenger, R. Valle, and B. Catanzaro · 2018
Cited alongside, same era.
StarGAN: unified generative adversarial networks for multi-domain image-to-image translation
Y. Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo · 2018
Later among the works it cites.
Adversarial audio synthesis
C. Donahue, J. McAuley, and M. Puckette · 2019
Closest in time.
AdaFlow: domain-adaptive density estimator with application to anomaly detection and unpaired cross-domain translation
M. Yamaguchi, Y. Koizumi, and N. Harada · 2019
Closest in time.
A universal music translation network
N. Mor, L. Wolf, A. Polyak, and Y. Taigman · 2019
Closest in time.
Unsupervised singing voice conversion
E. Nachmani and L. Wolf · 2019
Closest in time.
FFJORD: free-form continuous dynamics for scalable reversible generative models
W. Grathwohl, R. T. Q. Chen, J. Bettencourt, I. Sutskever, and D. Duvenaud · 2019
Closest in time.
Flow++: improving flow-based generative models with variational dequantization and architecture design
J. Ho, X. Chen, A. Srinivas, R. Duan, and P. Abbeel · 2019
Closest in time.
A RAD approach to deep mixture models
L. Dinh, J. Sohl-Dickstein, R. Pascanu, and H. Larochelle · 2019
Closest in time.
Emerging convolutions for generative normalizing flows
E. Hoogeboom, R. Van den Berg, and M. Welling · 2019
Closest in time.
A style-based generator architecture for generative adversarial networks
T. Karras, S. Laine, and T. Aila · 2019
Closest in time.
Praat: doing phonetics by computer, 2019
P. Boersma and D. Weenink · 2019
Closest in time.