Fetching the paper…
Reading the bibliography…
This paper describes the system developed by the BUT team for the fourth track of the VoxCeleb Speaker Recognition Challenge, focusing on diarization on the VoxConverse dataset.
“The ISL meeting corpus: The impact of meeting type on speech style,”
S. Burger, V. MacLaren, and H. Yu, · 2002
Earlier work this paper cites.
“The ICSI meeting corpus,”
A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stolcke, et al., · 2003
Earlier work this paper cites.
“The ami meeting corpus: A pre-announcement,”
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, et al., · 2005
Earlier work this paper cites.
“Efficient use of overlap information in speaker diarization,”
S. Otterson and M. Ostendorf, · 2007
Earlier work this paper cites.
“The NIST Rich Transcription 2009 (RT’09) evaluation,” http://www.itl.nist.gov/iad/mig/tests/rt/2009/docs/rt09-meeting-eval-plan-v2.pdf , 2009
2009
Earlier work this paper cites.
“Speech dereverberation based on variance-normalized delayed linear prediction,”
T. Nakatani, T. Yoshioka, K. Kinoshita, M. Miyoshi, and B.-H. Juang, · 2010
Earlier work this paper cites.
“Bayesian Speaker Verification with Heavy-Tailed Priors,”
P. Kenny, · 2010
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, et al., · 2011
Earlier work this paper cites.
“Analysis of i-vector length normalization in speaker recognition systems,”
D. Garcia-Romero and C. Y. Espy-Wilson, · 2011
Earlier work this paper cites.
“The etape corpus for the evaluation of speech-based tv content processing in the french language,”
G. Gravier, G. Adda, N. Paulson, M. Carré, A. Giraudel, and O. Galibert, · 2012
Cited alongside, same era.
“The mgb challenge: Evaluating multi-genre broadcast media recognition,”
P. Bell, M. J. Gales, T. Hain, J. Kilgour, P. Lanchantin, X. Liu, A. McParland, S. Renals, O. Saz, M. Wester, et al., · 2015
Cited alongside, same era.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Cited alongside, same era.
“Voxceleb2: Deep speaker recognition,”
J. S. Chung, A. Nagrani, and A. Zisserman, · 2018
Cited alongside, same era.
“Speaker diarization with enhancing speech for the first dihard challenge.,”
L. Sun, J. Du, C. Jiang, X. Zhang, S. He, B. Yin, and C.-H. Lee, · 2018
Cited alongside, same era.
“Cosface: Large margin cosine loss for deep face recognition,”
H. Wang, Y. Wang, Z. Zhou, X. Ji, D. Gong, J. Zhou, Z. Li, and W. Liu, · 2018
Later among the works it cites.
“Albayzin 2018 evaluation: the iberspeech-rtve challenge on speech technologies for spanish broadcast media,”
E. Lleida, A. Ortega, A. Miguel, V. Bazán-Gil, C. Pérez, M. Gómez, and A. de Prada, · 2019
Later among the works it cites.
“The second dihard diarization challenge: Dataset, task, and baselines,”
N. Ryant, K. Church, C. Cieri, A. Cristia, J. Du, S. Ganapathy, and M. Liberman, · 2019
Later among the works it cites.
“Spot the conversation: speaker diarisation in the wild,”
J. S. Chung, J. Huh, A. Nagrani, T. Afouras, and A. Zisserman, · 2020
Closest in time.
“But system for the second dihard speech diarization challenge,”
F. Landini, S. Wang, M. Diez, L. Burget, P. Matějka, K. Žmolíková, L. Mošner, A. Silnova, O. Plchot, O. Novotnỳ, et al., · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Nara-wpe: A python package for weighted prediction error dereverberation in numpy and tensorflow for online and offline processing,”
L. Drude, J. Heymann, C. Boeddeker, and R. Haeb-Umbach, · 2018
Cited alongside, same era.
“TED-LIUM 3: twice as much data and corpus repartition for experiments on speaker adaptation,”
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Estève, · 2018
Cited alongside, same era.
“X-vectors: Robust dnn embeddings for speaker recognition,”
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, · 2018
Cited alongside, same era.
“VBHMM x-vectors Diarization (aka VBx),” https://github.com/BUTSpeechFIT/VBx/tree/v1.1_VoxConverse2020
Cited in the paper.
Closest in time.
“Optimizing bayesian hmm based x-vector clustering for the second dihard speech diarization challenge,”
M. Diez, L. Burget, F. Landini, S. Wang, and H. Černockỳ, · 2020
Closest in time.
“pyannote.audio: neural building blocks for speaker diarization,”
H. Bredin, R. Yin, J. M. Coria, G. Gelly, P. Korshunov, M. Lavechin, D. Fustes, H. Titeux, W. Bouaziz, and M.-P. Gill, · 2020
Closest in time.
“Overlap-aware diarization: resegmentation using neural end-to-end overlapped speech detection,”
L. Bullock, H. Bredin, and L. P. Garcia-Perera, · 2020
Closest in time.