Fetching the paper…
Reading the bibliography…
An ability to model a generative process and learn a latent representation for speech in an unsupervised fashion will be crucial to process vast quantities of unlabelled speech data.
D. Wong, B. Juang, and D. Cheng, “Very low data rate speech compression with LPC vector and matrix quantization,” in
1983
Earlier work this paper cites.
V. Zue, S. Seneff, and J. Glass, “Speech database development at MIT: TIMIT and beyond,”
1990
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1,”
1993
Earlier work this paper cites.
A. Kain and M. W. Macon, “Spectral voice conversion for text-to-speech synthesis,” in
1998
Earlier work this paper cites.
T. Toda, Y. Ohtani, and K. Shikan, “Eigenvoice conversion based on gaussian mixture model,” in
2006
Earlier work this paper cites.
Y. Stylianou, “Voice transformation: A survey.” in
2009
Earlier work this paper cites.
N. Jaitly and G. E. Hinton, “Vocal tract length perturbation (VTLP) improves speech recognition,” in
2013
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”
2013
Earlier work this paper cites.
Z. Wu, E. S. Chng, and H. Li, “Conditional restricted boltzmann machine for voice conversion,” in
2013
Cited alongside, same era.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in
2014
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Cited alongside, same era.
X. Cui, V. Goel, and B. Kingsbury, “Data augmentation for deep neural network acoustic modeling,”
2015
Cited alongside, same era.
T. Nakashika, T. Takiguchi, and Y. Ariki, “Voice conversion using speaker-dependent conditional restricted boltzmann machine,”
2015
Cited alongside, same era.
T. Nakashika, T. Takiguchi, Y. Minami, T. Nakashika, T. Takiguchi, and Y. Minami, “Non-parallel training in voice conversion using an adaptive restricted boltzmann machine,”
2016
Later among the works it cites.
M. Blaauw and J. Bonada, “Modeling and transforming speech using variational autoencoders,”
2016
Later among the works it cites.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from non-parallel corpora using variational auto-encoder,” in
2016
Later among the works it cites.
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Later among the works it cites.
S. Tan and K. C. Sim, “Learning utterance-level normalisation using variational autoencoders for robust automatic speech recognition,” in
2016
Later among the works it cites.