Fetching the paper…
Reading the bibliography…
We propose a self-supervised representation learning model for the task of unsupervised phoneme boundary detection.
J. S. Garofolo, “Timit acoustic phonetic continuous speech corpus,”
1993
Earlier work this paper cites.
H. G. Tillmann and B. Pompino-Marschall, “Theoretical principles concerning segmentation, labelling strategies and levels of categorical annotation for spoken language database systems,” in
1993
Earlier work this paper cites.
F. Kubala, T. Anastasakos, H. Jin, L. Nguyen, and R. Schwartz, “Transcribing radio news,” in
1996
Earlier work this paper cites.
J. Keshet, S. Shalev-Shwartz, Y. Singer, and D. Chazan, “Phoneme alignment based on discriminative learning,” in
2005
Earlier work this paper cites.
S. Dusan and L. Rabiner, “On the relation between maximum spectral transition positions and phone boundaries,” in
2006
Earlier work this paper cites.
M. A. Pitt, L. Dilley, K. Johnson, S. Kiesling, W. Raymond, E. Hume, and E. Fosler-Lussier, “Buckeye corpus of conversational speech (2nd release),”
2007
Earlier work this paper cites.
Y. P. Estevan, V. Wan, and O. Scharenborg, “Finding maximum margin segments in speech,” in
2007
Earlier work this paper cites.
G. Almpanidis and C. Kotropoulos, “Phonemic segmentation using the generalised gamma distribution and small sample bayesian information criterion,”
2008
Earlier work this paper cites.
D. Rybach, C. Gollan, R. Schluter, and H. Ney, “Audio segmentation for speech recognition using segment features,” in
2009
Earlier work this paper cites.
J. Keshet, D. Grangier, and S. Bengio, “Discriminative keyword spotting,”
2009
Earlier work this paper cites.
O. J. Räsänen, U. K. Laine, and T. Altosaar, “An improved speech segmentation quality measure: the r-value,” in
2009
Earlier work this paper cites.
M. Gutmann and A. Hyvärinen, “Noise-contrastive estimation: A new estimation principle for unnormalized statistical models,” in
2010
Cited alongside, same era.
O. Räsänen, U. K. Laine, and T. Altosaar, “Blind segmentation of speech using non-linear filtering methods,”
2011
Cited alongside, same era.
M. H. Moattar and M. M. Homayounpour, “A review on speaker diarization systems and approaches,”
2012
Cited alongside, same era.
S. King and M. Hasegawa-Johnson, “Accurate speech segmentation by mimicking human auditory processing,” in
2013
Cited alongside, same era.
A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier nonlinearities improve neural network acoustic models,” in
2013
Cited alongside, same era.
2016
Later among the works it cites.
M. Goldrick, J. Keshet, E. Gustafson, J. Heller, and J. Needle, “Automatic analysis of slips of the tongue: Insights into the cognitive architecture of speech production,”
2016
Later among the works it cites.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using kaldi.” in
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
O. Rasanen, “Basic cuts revisited: Temporal segmentation of speech into phone-like units with statistical learning at a pre-linguistic level,” in
2014
Cited alongside, same era.
Y. Adi, J. Keshet, and M. Goldrick, “Vowel duration measurement using deep neural networks,” in
2015
Cited alongside, same era.
D.-T. Hoang and H.-C. Wang, “Blind phone segmentation based on spectral change detection using legendre polynomial approximation,”
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Cited alongside, same era.
Y. Adi, J. Keshet, E. Cibelli, E. Gustafson, C. Clopper, and M. Goldrick, “Automatic measurement of vowel duration via structured prediction,”
2016
Cited alongside, same era.
J. Franke, M. Mueller, F. Hamlaoui, S. Stueker, and A. Waibel, “Phoneme boundary detection using deep bidirectional lstms,” in
2016
Cited alongside, same era.
A. Ben-Shalom, J. Keshet, D. Modan, and A. Laufer, “Automatic tools for analyzing spoken hebrew.”
Cited in the paper.
2018
Later among the works it cites.
2018
Later among the works it cites.
A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”
2018
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2020
Closest in time.