Fetching the paper…
Reading the bibliography…
Vector Quantized Variational AutoEncoders (VQ-VAE) are a powerful representation learning framework that can discover discrete groups of features from a speech signal without supervision.
I. Rec, “P. 563: Single-ended method for objective speech quality assessment in narrow-band telephony applications,”
2004
Earlier work this paper cites.
International Telecommunication Union, Recommendation G.191: Software Tools and Audio Coding Standardization, Nov 11 2005
2005
Earlier work this paper cites.
M. Ekpenyong, E.-A. Urua, O. Watts, S. King, and J. Yamagishi, “Statistical parametric speech synthesis for ibibio,”
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in
2015
Earlier work this paper cites.
B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and music signal analysis in python,” in
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. van den Oord, O. Vinyals
2017
Earlier work this paper cites.
Y.-A. Chung and J. Glass, “Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,” in
2018
Earlier work this paper cites.
A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
W. Cai, J. Chen, and M. Li, “Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,”
2018
Cited alongside, same era.
T. Kosaka, Y. Aizawa, M. Kato, and T. Nose, “Acoustic model adaptation for emotional speech recognition using twitter-based emotional speech corpus,” in
X. Wang, S. Takaki, J. Yamagishi, S. King, and K. Tokuda, “A vector quantized variational autoencoder (vq-vae) autoregressive neural
2019
Later among the works it cites.
2019
Later among the works it cites.
C. Gârbacea, A. van den Oord, Y. Li, F. S. Lim, A. Luebs, O. Vinyals, and T. C. Walters, “Low bit-rate speech coding with vq-vae and a wavenet decoder,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
Y. Zhao, A. Ando, S. Takaki, J. Yamagishi, and S. Kobashikawa, “Does the Lombard Effect Improve Emotional Communication in Noise? — Analysis of Emotional Speech Acted in Noise,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
J. W. Kim, J. Salamon, P. Li, and J. P. Bello, “Crepe: A convolutional representation for pitch estimation,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord, “Unsupervised speech representation learning using wavenet autoencoders,”
2019
Cited alongside, same era.
S. Ding and R. Gutierrez-Osuna, “Group Latent Embedding for Vector Quantized Variational Autoencoder in Non-Parallel Voice Conversion,” in
2019
Cited alongside, same era.
2019
Later among the works it cites.
P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever, “Jukebox: A generative model for music,”
2020
Closest in time.
2020
Closest in time.
E. Cooper, C.-I. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and J. Yamagishi, “Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,”
2020
Closest in time.