Fetching the paper…
Reading the bibliography…
This work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition.
M. Gales, “Maximum likelihood linear transformations for HMM-based speech recognition,”
1997
Earlier work this paper cites.
J. J. Godfrey and E. Holliman, “Switchboard-1 release 2,”
1997
Earlier work this paper cites.
L. D. Consortium
1997
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “Fisher English training speech parts 1 and 2,”
2004
Earlier work this paper cites.
D. Liu, D. Kiecza, A. Srivastava, and F. Kubala, “Online speaker adaptation and tracking for real-time speech recognition,” in
2005
Earlier work this paper cites.
J. G. Fiscus, J. Ajot, M. Michel, and J. S. Garofolo, “The Rich Transcription 2006 spring meeting recognition evaluation,” in
2006
Earlier work this paper cites.
U. Von Luxburg, “A tutorial on spectral clustering,”
2007
Earlier work this paper cites.
P. Georgiou, M. Black, and S. Narayanan, “Behavioral signal processing for understanding (distressed) dyadic interactions: Some recent developments,” in
2011
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Cited alongside, same era.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz
2011
Cited alongside, same era.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,”
2011
Cited alongside, same era.
T. Hain, L. Burget, J. Dines, P. N. Garner, F. Grézl, A. El Hannani, M. Huijbregts, M. Karafiat, M. Lincoln, and V. Wan, “Transcribing meetings with the AMIDA systems,”
2012
Cited alongside, same era.
S. Narayanan and P. Georgiou, “Behavioral signal processing: Deriving human behavioral informatics from speech and language,”
D. Dimitriadis and P. Fousek, “Developing on-line speaker diarization system,” in
2017
Later among the works it cites.
M. À. India Massana, J. A. Rodríguez Fonollosa, and F. J. Hernando Pericás, “LSTM neural network-based speaker segmentation using acoustic and language modelling,” in
2017
Later among the works it cites.
H. Bredin, “Tristounet: Triplet loss for speaker turn embedding,” in
2017
Later among the works it cites.
Z. Zajíc, M. Hrúz, and L. Müller, “Speaker diarization using convolutional neural network for statistics accumulation refinement,” in
2017
Later among the works it cites.
T. J. Park and P. Georgiou, “Multimodal speaker segmentation and diarization using lexical and acoustic cues via sequence to sequence neural networks,” in
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
P. Cerva, J. Silovsky, J. Zdansky, J. Nouza, and L. Seps, “Speaker-adaptive speech recognition using speaker diarization for improved transcription of large spoken archives,”
2013
Cited alongside, same era.
2014
Cited alongside, same era.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in
2018
Later among the works it cites.
Q. Wang, C. Downey, L. Wan, P. A. Mansfield, and I. L. Moreno, “Speaker diarization with LSTM,” in
2018
Later among the works it cites.