Fetching the paper…
Reading the bibliography…
Recent work on unsupervised speech segmentation has used self-supervised models with phone and word segmentation modules that are trained jointly.
S. Roucos, R. Schwartz, and J. Makhoul, “Segment quantization for very-low-rate speech coding,” in Proc. ICASSP , 1982
1982
Earlier work this paper cites.
S. Roucos and M. O. Dunham, “A comparison of two methods for very-low-rate speech coding,” in Proc. MILCOM , 1985
1985
Earlier work this paper cites.
M. R. Brent, “An efficient, probabilistically sound algorithm for segmentation and word discovery,” Mach. Learn. , vol. 34, no. 1-3, pp. 71–105, 1999
1999
Earlier work this paper cites.
K. P. Murphy, “Hidden semi-Markov models (HSMMs),” 2002. [Online]. Available: http://www.cs.ubc.ca/~murphyk/mypapers.html
2002
Earlier work this paper cites.
M. A. Pitt, K. Johnson, E. Hume, S. Kiesling, and W. Raymond, “The Buckeye corpus of conversational speech: Labeling conventions and a test of transcriber reliability,” Speech Commun. , vol. 45, no. 1, pp. 89--95, 2005
2005
Earlier work this paper cites.
B. Varadarajan, S. Khudanpur, and E. Dupoux, “Unsupervised learning of acoustic sub-word units,” in Proc. ACL , 2008
2008
Earlier work this paper cites.
A. S. Park and J. R. Glass, “Unsupervised pattern discovery in speech,” IEEE Trans. Audio, Speech, Language Process. , vol. 16, no. 1, pp. 186–197, 2008
2008
Earlier work this paper cites.
S. J. Goldwater, T. L. Griffiths, and M. Johnson, “A Bayesian framework for word segmentation: Exploring the effects of context,” Cognition , vol. 112, no. 1, pp. 21–54, 2009
2009
Earlier work this paper cites.
M. Johnson and S. J. Goldwater, “Improving nonparameteric Bayesian inference: Experiments on unsupervised word segmentation with adaptor grammars,” in Proc. NAACL , 2009
2009
Earlier work this paper cites.
D. Mochihashi, T. Yamada, and N. Ueda, “Bayesian unsupervised word segmentation with nested Pitman-Yor language modeling,” in Proc. ACL , 2009
2009
Earlier work this paper cites.
H. Gish, M.-H. Siu, A. Chan, and B. Belfield, “Unsupervised training of an HMM-based speech recognizer for topic classification,” in Proc. Interspeech , 2009
2009
Earlier work this paper cites.
O. J. Räsänen, U. K. Laine, and T. Altosaar, “An improved speech segmentation quality measure: The R-value,” in Proc. Interspeech , 2009
2009
Earlier work this paper cites.
A. Jansen and B. Van Durme, “Efficient spoken term discovery using randomized algorithms,” in Proc. ASRU , 2011
2011
Earlier work this paper cites.
M. Elsner, S. J. Goldwater, and J. Eisenstein, “Bootstrapping a unified model of lexical and phonetic acquisition,” in Proc. ACL , 2012
2012
Earlier work this paper cites.
C.-y. Lee and J. R. Glass, “A nonparametric Bayesian approach to acoustic model discovery,” in Proc. ACL , 2012
2012
Earlier work this paper cites.
A. Jansen, E. Dupoux, S. J. Goldwater, M. Johnson, S. Khudanpur, K. Church, N. Feldman, H. Hermansky, F. Metze, R. Rose et al. , “A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisition,” in Proc. ICASSP , 2013
2013
Earlier work this paper cites.
M. Elsner, S. J. Goldwater, N. Feldman, and F. Wood, “A joint learning model of word segmentation, lexical acquisition and phonetic variability,” in Proc. EMNLP , 2013
2013
Earlier work this paper cites.
M. J. Johnson and A. S. Willsky, “Bayesian nonparametric hidden semi-markov models,” J. Mach. Learn. Res. , vol. 14, pp. 673–701, 2013
2013
Earlier work this paper cites.
C.-y. Lee, T. O’Donnell, and J. R. Glass, “Unsupervised lexicon discovery from acoustic input,” Trans. ACL , vol. 3, pp. 389–403, 2015
2015
Earlier work this paper cites.
O. J. Räsänen, G. Doyle, and M. C. Frank, “Unsupervised word discovery from speech using automatic segmentation into syllable-like units,” in Proc. Interspeech , 2015
2015
Earlier work this paper cites.
K. Uchiumi, H. Tsukahara, and D. Mochihashi, “Inducing word and part-of-speech with Pitman-Yor hidden semi-Markov models,” in Proc. ACL , 2015
2015
Earlier work this paper cites.
M. Versteegh, R. Thiollière, T. Schatz, X. N. Cao, X. Anguera, A. Jansen, and E. Dupoux, “The Zero Resource Speech Challenge 2015,” in Proc. Interspeech , 2015
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. ICLR , 2015
2015
Earlier work this paper cites.
O. Räsänen and H. Rasilo, “A joint model of word segmentation and meaning acquisition through cross-situational learning,” Psychol. Rev. , vol. 122, no. 4, pp. 792–829, 2015
2015
Cited alongside, same era.
L.-s. Lee, J. R. Glass, H.-y. Lee, and C.-a. Chan, “Spoken content retrieval—beyond cascading speech recognition with text retrieval,” IEEE Trans. Audio, Speech, Language Process. , vol. 23, no. 9, pp. 1389–1420, 2015
2015
Cited alongside, same era.
T. Taniguchi, S. Nagasaka, and R. Nakashima, “Nonparametric Bayesian double articulation analyzer for direct language acquisition from continuous speech signals,” IEEE Trans. Cogn. Developmental Syst. , vol. 8, no. 3, pp. 171–185, 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
J. Chorowski, N. Chen, R. Marxer, H. Dolfing, A. Łańcucki, G. Sanchez, T. Alumäe, and A. Laurent, “Unsupervised neural segmentation and clustering for unit discovery in sequential data,” in NeurIPS PGR Workshop , 2019
2019
Later among the works it cites.
K. Kawakami, C. Dyer, and P. Blunsom, “Learning to discover, ground and use words with segmental neural language models,” in Proc. ACL , 2019
2019
Later among the works it cites.
E. Dunbar, J. Karadayi, M. Bernard, X.-N. Cao, R. Algayres, L. Ondel, L. Besacier, S. Sakti, and E. Dupoux, “The Zero Resource Speech Challenge 2020: Discovering discrete subword and word units,” in Proc. Interspeech , 2020
2020
Later among the works it cites.
B. van Niekerk, L. Nortje, and H. Kamper, “Vector-quantized neural networks for acoustic unit discovery in the ZeroSpeech 2020 challenge,” in Proc. Interspeech , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Franke, M. Mueller, F. Hamlaoui, S. Stueker, and A. Waibel, “Phoneme boundary detection using deep bidirectional LSTMs,” in in Proc. Speech Commun. ITG Symposium , 2016
2016
Cited alongside, same era.
H. Kamper, A. Jansen, and S. J. Goldwater, “Unsupervised word segmentation and lexicon discovery using acoustic word embeddings,” IEEE Trans. Audio, Speech, Language Process. , vol. 24, no. 4, pp. 669–679, 2016
2016
Cited alongside, same era.
D. Harwath, A. Torralba, and J. R. Glass, “Unsupervised learning of spoken language with visual context,” in Proc. NIPS , 2016
2016
Cited alongside, same era.
L. Gelderloos and G. Chrupała, “From phonemes to images: Levels of representation in a recurrent neural model of visually-grounded language learning,” in Proc. COLING , 2016
2016
Cited alongside, same era.
H. Kamper, A. Jansen, and S. J. Goldwater, “A segmental framework for fully-unsupervised large-vocabulary speech recognition,” Comput. Speech Lang. , vol. 46, pp. 154–174, 2017
2017
Cited alongside, same era.
H. Kamper, K. Livescu, and S. J. Goldwater, “An embedded segmental K-means model for unsupervised segmentation and clustering of speech,” in Proc. ASRU , 2017
2017
Cited alongside, same era.
M. Elsner and C. Shain, “Speech segmentation with a neural encoder model of working memory,” in Proc. EMNLP , 2017
2017
Cited alongside, same era.
E. Dunbar, X. N. Cao, J. Benjumea, J. Karadayi, M. Bernard, L. Besacier, X. Anguera, and E. Dupoux, “The Zero Resource Speech Challenge 2017,” in Proc. ASRU , 2017
2017
Cited alongside, same era.
A. Baevski, S. Schneider, and M. Auli, “vq-wav2vec: Self-supervised learning of discrete speech representations,” in Proc. ICLR , 2020
2020
Later among the works it cites.
T. A. Nguyen, M. de Seyssel, P. Rozé, M. Rivière, E. Kharitonov, A. Baevski, E. Dunbar, and E. Dupoux, “The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,” in NeurIPS SAS Workshop , 2020
2020
Later among the works it cites.
F. Kreuk, J. Keshet, and Y. Adi, “Self-supervised contrastive learning for unsupervised phoneme segmentation,” in Proc. Interspeech , 2020
2020
Later among the works it cites.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen et al. , “Libri-light: A benchmark for ASR with limited or no supervision,” in Proc. ICASSP , 2020
2020
Later among the works it cites.
F. Kreuk, Y. Sheena, J. Keshet, and Y. Adi, “Phoneme boundary detection using learnable segmental features,” in Proc. ICASSP , 2020
2020
Later among the works it cites.
O. Räsänen and M. A. C. Blandón, “Unsupervised discovery of recurring speech patterns using probabilistic adaptive metrics,” in Proc. Interspeech , 2020
2020
Later among the works it cites.
S. Bhati, J. Villalba, P. Zelasko, and N. Dehak, “Self-expressing autoencoders for unsupervised spoken term discovery,” in Proc. Interspeech , 2020
2020
Later among the works it cites.
2021
Later among the works it cites.
S. Bhati, J. Villalba, P. Żelasko, L. Moro-Velazquez, and N. Dehak, “Segmental contrastive predictive coding for unsupervised word segmentation,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Kamper and B. van Niekerk, “Towards unsupervised phone and word segmentation using self-supervised vector-quantized neural networks,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
J. Chorowski, G. Ciesielski, J. Dzikowski, A. Łancucki, R. Marxer, M. Opala, P. Pusz, P. Rychlikowski, and M. Stypułkowski, “Aligned contrastive predictive coding,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
B. van Niekerk, L. Nortje, M. Baas, and H. Kamper, “Analyzing speaker information in self-supervised models to improve zero-resource speech processing,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE Trans. Audio, Speech, Language Process. , 2021
2021
Later among the works it cites.
R. Sanabria, H. Tang, and S. J. Goldwater, “On the difficulty of segmenting words with attention,” in EMNLP Insights Workshop , 2021
2021
Later among the works it cites.
R. Algayres, T. Ricoul, J. Karadayi, H. Laurençon, S. Zaiem, A. Mohame, B. Sagot, and E. Dupoux, “DP-Parse: Finding word boundaries from raw speech with an instance lexicon,” Trans. ACL , vol. 10, pp. 1051–1065, 2022
2022
Closest in time.
K. Olaleye, D. Onea t
2022
Closest in time.
P. Peng and D. Harwath, “Word discovery in visually grounded, self-supervised speech models,” in Proc. Interspeech , 2022
2022
Closest in time.