Fetching the paper…
Reading the bibliography…
In this paper, we explore vector quantization for acoustic unit discovery.
A. Kain and M. W. Macon, “Spectral voice conversion for text-to-speech synthesis,” in
1998
Earlier work this paper cites.
B. Varadarajan, S. Khudanpur, and E. Dupoux, “Unsupervised learning of acoustic sub-word units,” in
2008
Earlier work this paper cites.
S. Sakti, R. Maia, S. Sakai, T. Shimizu, and S. Nakamura, “Development of HMM-based Indonesian speech synthesis,” in
2008
Earlier work this paper cites.
S. Sakti, E. Kelana, H. Riza, S. Sakai, K. Markov, and S. Nakamura, “Development of Indonesian large vocabulary continuous speech recognition system within A-STAR project,” in
2008
Earlier work this paper cites.
O. J. Räsänen, “Computational modeling of phonetic and lexical learning in early language acquisition: Existing models and future directions,”
2012
Earlier work this paper cites.
C.-y. Lee and J. R. Glass, “A nonparametric Bayesian approach to acoustic model discovery,” in
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
T. Schatz, V. Peddinti, F. Bach, A. Jansen, H. Hermansky, and E. Dupoux, “Evaluating speech features with the minimal-pair ABX task: Analysis of the classical MFC/PLP pipeline,” in
2013
Earlier work this paper cites.
M.-H. Siu, H. Gish, A. Chan, W. Belfield, and S. Lowe, “Unsupervised training of an HMM-based self-organizing unit recognizer with applications to topic classification and keyword discovery,”
2014
Earlier work this paper cites.
C.-y. Lee, T. O’Donnell, and J. R. Glass, “Unsupervised lexicon discovery from acoustic input,”
2015
Earlier work this paper cites.
L. Badino, A. Mereta, and L. Rosasco, “Discovering discrete subword units with binarized autoencoders and hidden-Markov-model encoders,” in
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Earlier work this paper cites.
M. Versteegh, X. Anguera, A. Jansen, and E. Dupoux, “The Zero Resource Speech Challenge 2015: Proposed approaches and results,” in
2016
Earlier work this paper cites.
N. Zeghidour, G. Synnaeve, N. Usunier, and E. Dupoux, “Joint learning of speaker and phonetic similarities with Siamese networks,” in
2016
Earlier work this paper cites.
L. Ondel, L. Burget, and J. Černockỳ, “Variational inference for acoustic unit discovery,”
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Z. Wu, O. Watts, and S. King, “Merlin: An open source neural network speech synthesis system.” in
2016
Cited alongside, same era.
E. Dunbar, X. N. Cao, J. Benjumea, J. Karadayi, M. Bernard, L. Besacier, X. Anguera, and E. Dupoux, “The Zero Resource Speech Challenge 2017,” in
2017
Cited alongside, same era.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in
2017
Cited alongside, same era.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An unsupervised autoregressive model for speech representation learning,” in
2019
Later among the works it cites.
C. Coupé, Y. M. Oh, D. Dediu, and F. Pellegrino, “Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche,”
2019
Later among the works it cites.
R. Eloff, A. Nortje, B. L. Van Niekerk, A. Govender, L. Nortje, A. Pretorius, E. Van Biljon, E. Van der Westhuizen, L. Van Staden, and H. Kamper, “Unsupervised acoustic unit discovery for speech synthesis using discrete latent-variable neural networks,” in
2019
Later among the works it cites.
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord, “Unsupervised speech representation learning using WaveNet autoencoders,”
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Schatz and N. H. Feldman, “Neural network vs. HMM speech recognition systems as models of human cross-linguistic phonetic perception,” in
2018
Cited alongside, same era.
J.-c. Chou, C.-c. Yeh, H.-y. Lee, and L.-s. Lee, “Multi-target voice conversion without parallel data by adversarially learning disentangled audio representations,” in
2018
Cited alongside, same era.
M. Heck, S. Sakti, and S. Nakamura, “Learning supervised feature transformations on zero resources for improved acoustic unit discovery,”
2018
Cited alongside, same era.
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh
2018
Cited alongside, same era.
L. Kaiser, S. Bengio, A. Roy, A. Vaswani, N. Parmar, J. Uszkoreit, and N. Shazeer, “Fast decoding in sequence models using discrete latent variables,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, “FFTNet: A real-time speaker-dependent neural vocoder,” in
2018
Cited alongside, same era.
2019
Later among the works it cites.
J. Lorenzo-Trueba, T. Drugman, J. Latorre, T. Merritt, B. Putrycz, R. Barra-Chicote, A. Moinet, and V. Aggarwal, “Towards achieving robust universal neural vocoding,” in
2019
Later among the works it cites.
W. Wang, Q. Tang, and K. Livescu, “Unsupervised pre-training of bidirectional speech encoders via masked reconstruction,” in
2020
Closest in time.
P.-J. Last, H. A. Engelbrecht, and H. Kamper, “Unsupervised feature learning for speech using correspondence and siamese networks,”
2020
Closest in time.
A. Baevski, S. Schneider, and M. Auli, “vq-wav2vec: Self-supervised learning of discrete speech representations,” in
2020
Closest in time.
M. Rivière, A. Joulin, P.-E. Mazaré, and E. Dupoux, “Unsupervised pretraining transfers well across languages,” in
2020
Closest in time.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen
2020
Closest in time.
2020
Closest in time.
M. Chen and T. Hain, “Unsupervised acoustic unit representation learning for voice conversion using WaveNet auto-encoders,” in
2020
Closest in time.