Fetching the paper…
Reading the bibliography…
The recently proposed VBx diarization method uses a Bayesian hidden Markov model to find speaker clusters in a sequence of x-vectors.
G. Sun, C. Zhang, P. C. Woodland, Speaker diarisation using 2D self-attentive combination of embeddings , CoRR abs/1902.03190 · 1902
Earlier work this paper cites.
1907
Earlier work this paper cites.
doi:10.21437/Interspeech.2020-1908
Q. Lin, Y. Hou, M. Li, Self-Attentive Similarity Measurement Strategies in Speaker Diarization , in: Proc. Interspeech 2020, 2020, pp. 284–288 · 1908
Earlier work this paper cites.
1910
Earlier work this paper cites.
L. Bullock, H. Bredin, L. P. Garcia-Perera, Overlap-aware diarization: resegmentation using neural end-to-end overlapped speech detection (2019) · 1910
Earlier work this paper cites.
NIST SRE 2000 Evaluation Plan, https://www.nist.gov/sites/default/files/documents/2017/09/26/spk-2000-plan-v1.0.htm_.pdf
2000
Earlier work this paper cites.
A. F. Martin, M. A. Przybocki, Stream-based speaker segmentation using speaker factors and eigenvoices, in: 7th European Conference on Speech Communication and Technology, Eurospeech, Vol. 7, num. 2, 2001, pp. 787–790
2001
Earlier work this paper cites.
Y. Fujita, S. Watanabe, S. Horiguchi, Y. Xue, K. Nagamatsu, End-to-End Neural Diarization: Reformulating Speaker Diarization as Simple Multi-label Classification (2020) · 2003
Earlier work this paper cites.
M. Beal, Variational Algorithms for Approximate Bayesian Inference, Ph.D. thesis, Gatsby Unit, University College London (2003)
2003
Earlier work this paper cites.
S. Watanabe, M. Mandel, J. Barker, E. Vincent, A. Arora, X. Chang, S. Khudanpur, V. Manohar, D. Povey, D. Raj, D. Snyder, A. S. Subramanian, J. Trmal, B. B. Yair, C. Boeddeker, Z. Ni, Y. Fujita, S. Horiguchi, N. Kanda, T. Yoshioka, N. Ryant, CHiME-6 Challenge: Tackling Multispeaker Speech Recognition for Unsegmented Recordings (2020) · 2004
Earlier work this paper cites.
S. Horiguchi, Y. Fujita, S. Watanabe, Y. Xue, K. Nagamatsu, End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors (2020) · 2005
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, et al., The AMI meeting corpus: A pre-announcement, in: International workshop on machine learning for multimodal interaction, Springer, 2006, pp. 28–39
2006
Earlier work this paper cites.
C. M. Bishop, Pattern Recognition and Machine Learning, Springer-Verlag New York, Inc., Secaucus, NJ, USA, 2006
2006
Earlier work this paper cites.
2007
Earlier work this paper cites.
E. B Fox, E. B Sudderth, M. Jordan, A. S Willsky, The sticky HDP-HMM: Bayesian nonparametric hidden Markov models with persistent states, Technical Report P-2777, MIT LIDS (01 2007)
2007
Earlier work this paper cites.
M. Pal, M. Kumar, R. Peri, T. J. Park, S. H. Kim, C. Lord, S. Bishop, S. Narayanan, Meta-learning with Latent Space Clustering in Generative Adversarial Network for Speaker Diarization (2020) · 2007
Earlier work this paper cites.
doi:10.1109/TASL.2007.902460
X. Anguera, C. Wooters, J. Hernando, Acoustic Beamforming for Speaker Diarization of Meetings, Audio, Speech, and Language Processing, IEEE Transactions on 15 (2007) 2011 – 2022 · 2007
Earlier work this paper cites.
P. Kenny, Bayesian Analysis of Speaker Diarization with Eigenvoice Priors, Tech. rep., Montreal: CRIM (2008)
2008
Earlier work this paper cites.
K. Kinoshita, M. Delcroix, N. Tawara, Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds (2020) · 2010
Earlier work this paper cites.
X. Xiao, N. Kanda, Z. Chen, T. Zhou, T. Yoshioka, S. Chen, Y. Zhao, G. Liu, Y. Wu, J. Wu, S. Liu, J. Li, Y. Gong, Microsoft Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2020 (2020) · 2010
Cited alongside, same era.
doi:10.1109/TASL.2010.2064307
N. D. et al., Front-End Factor Analysis for Speaker Verification, IEEE Transactions on Audio, Speech, and Language Processing 19 (4) (2011) 788–798 · 2010
Cited alongside, same era.
P. Kenny, Bayesian Speaker Verification with Heavy-Tailed Priors, in: Proc. Odyssey-10, Brno, Czech Republic, 2010
2010
Cited alongside, same era.
N. Brummer, E. Villiers, The speaker partitioning problem, Proc. of Odyssey 2010
2010
Cited alongside, same era.
G. Sun, C. Zhang, P. Woodland, Combination of Deep Speaker Embeddings for Diarisation (2020) · 2010
Cited alongside, same era.
H. Zeinali, H. Sameti, T. Stafylakis, DeepMine Speech Processing Database: Text-Dependent and Independent Speaker Verification and Speech Recognition in Persian and English., in: Odyssey, 2018, pp. 386–392
2018
Later among the works it cites.
doi:10.1109/ICASSP.2018.8461546
M. Maciejewski, D. Snyder, V. Manohar, N. Dehak, S. Khudanpur, Characterizing Performance of Speaker Diarization Systems on Far-Field Speech Using Standard Methods, in: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 5244–5248 · 2018
Later among the works it cites.
N. R. et. al., The Second DIHARD Diarization Challenge: Dataset, task, and baselines., in: Proceedings of Interspeech, 2019
2019
Later among the works it cites.
doi:10.21437/Interspeech.2019-2813
M. Diez, L. Burget, S. Wang, J. Rohdin, H. Černocký, Bayesian HMM based x-vector clustering for Speaker Diarization, in: Proc. Interspeech 2019, 2019, pp. 346–350 · 2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al., Pytorch: An imperative style, high-performance deep learning library, in: Advances in neural information processing systems, 2019, pp. 8026–8037
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, et al., The Kaldi speech recognition toolkit, in: IEEE 2011 workshop on automatic speech recognition and understanding, IEEE Signal Processing Society, 2011
2011
Cited alongside, same era.
D. Garcia-Romero, C. Espy-Wilson, Analysis of i-vector Length Normalization in Speaker Recognition Systems., 2011, pp. 249–252
2011
Cited alongside, same era.
D. Raj, Z. Huang, S. Khudanpur, Multi-class Spectral Clustering with Overlaps for Speaker Diarization (2020) · 2011
Cited alongside, same era.
D. Raj, L. P. Garcia-Perera, Z. Huang, S. Watanabe, D. Povey, A. Stolcke, S. Khudanpur, DOVER-Lap: A Method for Combining Overlap-aware Diarization Outputs (2020) · 2011
Cited alongside, same era.
doi:10.1109/TASL.2013.2264673
S. H. Shum, N. Dehak, R. Dehak, J. R. Glass, Unsupervised Methods for Speaker Diarization: An Integrated and Iterative Approach, IEEE Transactions on Audio, Speech, and Language Processing 21 (10) (2013) 2015–2028 · 2013
Cited alongside, same era.
doi:10.1109/TASLP.2013.2285474
M. Senoussaoui, P. Kenny, T. Stafylakis, P. Dumouchel, A Study of the Cosine Distance-Based Mean Shift for Telephone Speech Diarization , IEEE/ACM Trans. Audio, Speech and Lang. Proc. 22 (1) (2014) 217–227 · 2013
Cited alongside, same era.
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778
2016
Cited alongside, same era.
2019
Later among the works it cites.
J. Deng, J. Guo, N. Xue, S. Zafeiriou, Arcface: Additive angular margin loss for deep face recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 4690–4699
2019
Later among the works it cites.
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, C. Wang, Fully supervised speaker diarization, in: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2019, pp. 6301–6305
2019
Later among the works it cites.
Z. Huang, S. Watanabe, Y. Fujita, P. García, Y. Shao, D. Povey, S. Khudanpur, Speaker diarization with region proposal network, in: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2020, pp. 6514–6518
2020
Closest in time.
I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Korenevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, I. Podluzhny, et al., Target-Speaker Voice Activity Detection: A Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario , Interspeech 2020 doi:10.21437/interspeech.2020-1602
2020
Closest in time.
F. Landini, S. Wang, M. Diez, L. Burget, P. Matějka, K. Žmolíková, L. Mošner, A. Silnova, O. Plchot, O. Novotný, et al., BUT System for the Second DIHARD Speech Diarization Challenge, in: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2020, pp. 6529–6533
2020
Closest in time.
M. Diez, L. Burget, F. Landini, S. Wang, H. Černocký, Optimizing Bayesian HMM based x-vector clustering for the second DIHARD speech diarization challenge, in: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2020, pp. 6519–6523
2020
Closest in time.
M. Diez, L. Burget, F. Landini, J. Černocký, Analysis of Speaker Diarization Based on Bayesian HMM With Eigenvoice Priors, IEEE/ACM Transactions on Audio, Speech, and Language Processing 28 (2020) 355–368
2020
Closest in time.
Y. Fan, J. Kang, L. Li, K. Li, H. Chen, S. Cheng, P. Zhang, Z. Zhou, Y. Cai, D. Wang, CN-CELEB: a challenging Chinese speaker recognition dataset, in: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2020, pp. 7604–7608
2020
Closest in time.
H. Bredin, R. Yin, J. M. Coria, G. Gelly, P. Korshunov, M. Lavechin, D. Fustes, H. Titeux, W. Bouaziz, M.-P. Gill, pyannote.audio: neural building blocks for speaker diarization, in: ICASSP 2020, IEEE International Conference on Acoustics, Speech, and Signal Processing, Barcelona, Spain, 2020
2020
Closest in time.
doi:10.21437/Interspeech.2020-1879
H. Aronowitz, W. Zhu, M. Suzuki, G. Kurata, R. Hoory, New Advances in Speaker Diarization , in: Interspeech 2020, 21st Annual Conference of the International Speech Communication Association, Virtual Event, Shanghai, China, 25-29 October 2020, ISCA, 2020, pp. 279–283 · 2020
Closest in time.
doi:10.21437/Odyssey.2020-15
Q. Lin, W. Cai, L. Yang, J. Wang, J. Zhang, M. Li, DIHARD II is Still Hard: Experimental Results and Discussions from the DKU-LENOVO Team , in: Proc. Odyssey 2020 The Speaker and Language Recognition Workshop, 2020, pp. 102–109 · 2020
Closest in time.
doi:10.1109/ICASSP40776.2020.9053373
Y. Fathullah, C. Zhang, P. Woodland, Improved Large-Margin Softmax Loss for Speaker Diarisation, 2020, pp. 7104–7108 · 2020
Closest in time.
M. Diez, L. Burget, CSL VBx derivations, http://www.fit.vutbr.cz/~mireia/CSL_VBHMM_tech_report.pdf , Technical Report, Brno University of Technology (2021)
2021
Closest in time.