Fetching the paper…
Reading the bibliography…
Speaker diarization for real-life scenarios is an extremely challenging problem.
H. Ning, M. Liu, H. Tang, and T. Huang, “A spectral clustering approach to speaker diarization,” in
2006
Earlier work this paper cites.
S. J. D. Prince and J. H. Elder, “Probabilistic linear discriminant analysis for inferences about identity,” in
2007
Earlier work this paper cites.
C. Fredouille, S. Bozonnet, and N. Evans, “The LIA-EURECOM RT‘09 speaker diarization system,” in
2009
Earlier work this paper cites.
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlíček, Y. Qian, P. Schwarz, J. Silovský, G. Stemmer, and K. Vesel, “The kaldi speech recognition toolkit,” in
2011
Earlier work this paper cites.
T. Yoshioka and T. Nakatani, “Generalization of multi-channel linear prediction methods for blind MIMO impulse response shortening,”
2012
Earlier work this paper cites.
E. Variani, X. Lei, E. McDermott, I. Moreno, and J. Gonzalez-Dominguez, “Deep neural networks for small footprint text-dependent speaker verification,” in
2014
Earlier work this paper cites.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory recurrent neural network architectures for large scale acoustic modeling,” in
2014
Earlier work this paper cites.
G. Sell and D. Garcia-Romero, “Diarization resegmentation in the factor analysis subspace,” in
2015
Earlier work this paper cites.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,”
2017
Earlier work this paper cites.
D. Dimitriadis and P. Fousek, “Developing on-line speaker diarization system,” in
2017
Earlier work this paper cites.
D. Garcia-Romero, D. Snyder, G. Sell, D. Povey, and A. McCree, “Speaker diarization using deep neural network embeddings,” in
2017
Earlier work this paper cites.
K. Žmolíková, M. Delcroix, K. Kinoshita, T. Higuchi, A. Ogawa, and T. Nakatani, “Speaker-aware neural network based beamformer for speaker extraction in speech mixtures,” in
2017
Cited alongside, same era.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “Mixup: Beyond empirical risk minimization,”
2017
Cited alongside, same era.
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Willaba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, and S. Khudanpur, “Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,” in
2018
Cited alongside, same era.
M. Diez, F. Landini, L. Burget, J. Rohdin, A. Silnova, K. Žmolíková, O. Novotny, K. Veselý, O. Glembek, O. Plchot, L. Mošner, and P. Matejka, “BUT system for DIHARD speech diarization challenge 2018,” in
2018
Cited alongside, same era.
2019
Later among the works it cites.
——, “The second DIHARD diarization challenge: Dataset, task, and baselines,”
2019
Later among the works it cites.
Y. Fujita, N. Kanda, S. Horiguchi, Y. Xue, K. Nagamatsu, and S. Watanabe, “End-to-end neural speaker diarization with self-attention,” in
2019
Later among the works it cites.
N. Kanda, S. Horiguchi, R. Takashima, Y. Fujita, K. Nagamatsu, and S. Watanabe, “Auxiliary interference speaker loss for target-speaker speech recognition,” in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” in
2018
Cited alongside, same era.
N. Ryant, K. Church, C. Cieri, A. Cristia, J. Du, S. Ganapathy, and M. Liberman, “First DIHARD challenge evaluation plan,” Tech. Rep., 2018
2018
Cited alongside, same era.
J. Barker, S. Watanabe, E. Vincent, and J. Trmal, “The fifth ’CHiME’ speech separation and recognition challenge: Dataset, task and baselines,” in
2018
Cited alongside, same era.
M. Delcroix, K. Zmolikova, K. Kinoshita, A. Ogawa, and T. Nakatani, “Single channel target speaker extraction and recognition with speaker beam,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
I. Medennikov, Y. Khokhlov, A. Romanenko, D. Popov, N. Tomashenko, I. Sorokin, and A. Zatvornitskiy, “An investigation of mixup training strategies for acoustic models in ASR,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
T. Park, K. Han, M. Kumar, and S. Narayanan, “Auto-tuning spectral clustering for speaker diarization using normalized maximum eigengap,”
2019
Later among the works it cites.
2020
Closest in time.
2020
Closest in time.
I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Korenevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, I. Podluzhny, A. Laptev, and A. Romanenko, “The STC system for the CHiME-6 challenge,” in
2020
Closest in time.