Fetching the paper…
Reading the bibliography…
We investigate segmenting and clustering speech into low-bitrate phone-like sequences without supervision.
M. A. Pitt, K. Johnson, E. Hume, S. Kiesling, and W. Raymond, “The Buckeye corpus of conversational speech: Labeling conventions and a test of transcriber reliability,” Speech Commun. , vol. 45, no. 1, pp. 89–95, 2005
2005
Earlier work this paper cites.
M. Johnson, T. L. Griffiths, and S. J. Goldwater, “Adaptor grammars: A framework for specifying compositional nonparametric Bayesian models,” in Proc. NIPS , 2006
2006
Earlier work this paper cites.
O. J. Räsänen, U. K. Laine, and T. Altosaar, “An improved speech segmentation quality measure: the r-value,” in Proc. Interspeech , 2009
2009
Earlier work this paper cites.
S. J. Goldwater, T. L. Griffiths, and M. Johnson, “A Bayesian framework for word segmentation: Exploring the effects of context,” Cognition , vol. 112, no. 1, pp. 21–54, 2009
2009
Earlier work this paper cites.
M. A. Carlin, S. Thomas, A. Jansen, and H. Hermansky, “Rapid evaluation of speech representations for spoken term discovery,” in Proc. Interspeech , 2011
2011
Earlier work this paper cites.
O. J. Räsänen, “Computational modeling of phonetic and lexical learning in early language acquisition: Existing models and future directions,” Speech Commun. , vol. 54, pp. 975–997, 2012
2012
Earlier work this paper cites.
A. Jansen et al. , “A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisition,” in Proc. ICASSP , 2013
2013
Earlier work this paper cites.
T. Schatz et al. , “Evaluating speech features with the minimal-pair ABX task: Analysis of the classical MFC/PLP pipeline,” in Proc. Interspeech , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
O. Räsänen, “Basic cuts revisited: Temporal segmentation of speech into phone-like units with statistical learning at a pre-linguistic level,” in Proc. CogSci , 2014
2014
Earlier work this paper cites.
G. Synnaeve, T. Schatz, and E. Dupoux, “Phonetics embedding learning with side information,” in Proc. SLT , 2014
2014
Earlier work this paper cites.
L. Badino, A. Mereta, and L. Rosasco, “Discovering discrete subword units with binarized autoencoders and hidden-Markov-model encoders,” in Proc. Interspeech , 2015
2015
Earlier work this paper cites.
D. Renshaw, H. Kamper, A. Jansen, and S. J. Goldwater, “A comparison of neural network methods for unsupervised representation learning on the Zero Resource Speech Challenge,” in Proc. Interspeech , 2015
2015
Earlier work this paper cites.
N. Zeghidour, G. Synnaeve, N. Usunier, and E. Dupoux, “Joint learning of speaker and phonetic similarities with Siamese networks,” in Proc. Interspeech , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Franke, M. Mueller, F. Hamlaoui, S. Stueker, and A. Waibel, “Phoneme boundary detection using deep bidirectional LSTMs,” in in Proc. Speech Commun. ITG Symposium . VDE, 2016, pp. 1–5
2016
Cited alongside, same era.
N. Zeghidour, G. Synnaeve, M. Versteegh, and E. Dupoux, “A deep scattering spectrum-deep Siamese network pipeline for unsupervised acoustic modeling,” in Proc. ICASSP , 2016
2016
Cited alongside, same era.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in Proc. NeurIPS , 2017
2017
Cited alongside, same era.
Y.-H. Wang, C.-T. Chung, and H.-y. Lee, “Gate activation signal analysis for gated recurrent neural networks and its correlation with phoneme boundaries,” in Proc. Interspeech , 2017
2017
Cited alongside, same era.
A. Saksida, A. Langus, and M. Nespor, “Co-occurrence statistics as a language-dependent cue for speech segmentation,” Devel. Sci. , vol. 20, no. 3, p. e12390, 2017
C. Coupé, Y. M. Oh, D. Dediu, and F. Pellegrino, “Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche,” Sci. Adv. , vol. 5, no. 9, 2019
2019
Later among the works it cites.
J. Chorowski et al. , “Unsupervised neural segmentation and clustering for unit discovery in sequential data,” in NeurIPS PGR Workshop , 2019
2019
Later among the works it cites.
C. Shain and M. Elsner, “Acquiring language from speech by learning to remember and predict,” in Proc. CoNLL , 2020
2020
Closest in time.
P.-J. Last, H. A. Engelbrecht, and H. Kamper, “Unsupervised feature learning for speech using correspondence and siamese networks,” IEEE Signal Proc. Let. , vol. 27, pp. 421–425, 2020
2020
Closest in time.
R. Algayres, M. S. Zaiem, B. Sagot, and E. Dupoux, “Evaluating the reliability of acoustic speech embeddings,” in Proc. Interspeech , 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
H. Kamper, K. Livescu, and S. Goldwater, “An embedded segmental K-means model for unsupervised segmentation and clustering of speech,” in Proc. ASRU , 2017
2017
Cited alongside, same era.
H. Kamper, A. Jansen, and S. Goldwater, “A segmental framework for fully-unsupervised large-vocabulary speech recognition,” Comput. Speech Lang. , vol. 46, pp. 154–174, 2017
2017
Cited alongside, same era.
E. Dupoux, “Cognitive science in the era of artificial intelligence: A roadmap for reverse-engineering the infant language-learner,” Cognition , vol. 173, pp. 43–59, 2018
2018
Cited alongside, same era.
M. Heck, S. Sakti, and S. Nakamura, “Learning supervised feature transformations on zero resources for improved acoustic unit discovery,” IEICE T. Inf. Syst. , vol. 101, no. 1, pp. 205–214, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
R. Menon, H. Kamper, E. Van Der Westhuizen, J. Quinn, and T. R. Niesler, “Feature exploration for almost zero-resource ASR-free keyword spotting using a multilingual bottleneck extractor and correspondence autoencoders,” in Proc. Interspeech , 2019
2019
Cited alongside, same era.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An unsupervised autoregressive model for speech representation learning,” in Proc. Interspeech , 2019
2019
Cited alongside, same era.
2020
Closest in time.
A. Baevski, S. Schneider, and M. Auli, “vq-wav2vec: Self-supervised learning of discrete speech representations,” in Proc. ICLR , 2020
2020
Closest in time.
M. Chen and T. Hain, “Unsupervised acoustic unit representation learning for voice conversion using WaveNet auto-encoders,” in Proc. Interspeech , 2020
2020
Closest in time.
Y.-A. Chung, H. Tang, and J. Glass, “Vector-quantized autoregressive predictive coding,” in Proc. Interspeech , 2020
2020
Closest in time.
E. Dunbar et al. , “The Zero Resource Speech Challenge 2020: Discovering discrete subword and word units,” in Proc. Interspeech , 2020
2020
Closest in time.
B. van Niekerk, L. Nortje, and H. Kamper, “Vector-quantized neural networks for acoustic unit discovery in the zerospeech 2020 challenge,” in Proc. Interspeech , 2020
2020
Closest in time.
O. Räsänen and M. A. C. Blandón, “Unsupervised discovery of recurring speech patterns using probabilistic adaptive metrics,” in Proc. Interspeech , 2020
2020
Closest in time.
2020
Closest in time.
F. Kreuk, J. Keshet, and Y. Adi, “Self-supervised contrastive learning for unsupervised phoneme segmentation,” in Proc. Interspeech , 2020
2020
Closest in time.
F. Kreuk, Y. Sheena, J. Keshet, and Y. Adi, “Phoneme boundary detection using learnable segmental features,” in Proc. ICASSP , 2020
2020
Closest in time.