Fetching the paper…
Reading the bibliography…
Speaker segmentation consists in partitioning a conversation between one or more speakers into speaker turns.
J. Carletta, “Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus,” Language Resources and Evaluation , vol. 41, no. 2, 2007
2007
Earlier work this paper cites.
S. Otterson and M. Ostendorf, “Efficient use of overlap information in speaker diarization,” in 2007 IEEE Workshop on Automatic Speech Recognition & Understanding (ASRU) . IEEE, 2007, pp. 683–686
2007
Earlier work this paper cites.
D. Charlet, C. Barras, and J. Liénard, “Impact of overlapping speech detection on speaker diarization for broadcast news and debates,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing , May 2013, pp. 7707–7711
2013
Earlier work this paper cites.
D. Snyder, G. Chen, and D. Povey, “MUSAN: A Music, Speech, and Noise Corpus,” 2015
2015
Earlier work this paper cites.
R. Yin, H. Bredin, and C. Barras, “Speaker Change Detection in Broadcast TV Using Bidirectional Long Short-Term Memory Networks,” in Proc. Interspeech 2017 , 2017
2017
Earlier work this paper cites.
H. Bredin, “pyannote.metrics: a toolkit for reproducible evaluation, diagnostic, and error analysis of speaker diarization systems,” in Proc. Interspeech 2017 , Stockholm, Sweden, August 2017. [Online]. Available: http://pyannote.github.io/pyannote-metrics
2017
Earlier work this paper cites.
G. Gelly and J.-L. Gauvain, “Optimization of RNN-Based Speech Activity Detection,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 3, pp. 646–656, March 2018
2018
Earlier work this paper cites.
M. Ravanelli and Y. Bengio, “Speaker recognition from raw waveform with sincnet,” in Proc. SLT 2018 , 2018
2018
Cited alongside, same era.
M. Kunešová, M. Hrúz, Z. Zajíc, and V. Radová, “Detection of overlapping speech for the purposes of speaker diarization,” in Speech and Computer , A. A. Salah, A. Karpov, and R. Potapova, Eds. Cham: Springer International Publishing, 2019, pp. 247–257
2019
Cited alongside, same era.
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, and S. Watanabe, “End-to-End Neural Speaker Diarization with Permutation-free Objectives,” in Interspeech , 2019, pp. 4300–4304
2019
Cited alongside, same era.
Y. Fujita, N. Kanda, S. Horiguchi, Y. Xue, K. Nagamatsu, and S. Watanabe, “End-to-end neural speaker diarization with self-attention,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2019, pp. 296–303
2019
Cited alongside, same era.
J. S. Chung, J. Huh, A. Nagrani, T. Afouras, and A. Zisserman, “Spot the Conversation: Speaker Diarisation in the Wild,” in Proc. Interspeech 2020 , 2020, pp. 299–303. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-2337
2020
Later among the works it cites.
F. Landini, J. Profant, M. Diez, and L. Burget, “Bayesian hmm clustering of x-vector sequences (vbx) in speaker diarization: theory, implementation and analysis on standard tasks,” 2020
2020
Later among the works it cites.
H. Bredin, R. Yin, J. M. Coria, G. Gelly, P. Korshunov, M. Lavechin, D. Fustes, H. Titeux, W. Bouaziz, and M.-P. Gill, “pyannote.audio: neural building blocks for speaker diarization,” in Proc. ICASSP 2020 , 2020
2020
Later among the works it cites.
S. Horiguchi, P. Garcia, Y. Fujita, S. Watanabe, and K. Nagamatsu, “End-to-end speaker diarization as post-processing,” 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Bullock, H. Bredin, and L. P. Garcia-Perera, “Overlap-aware diarization: Resegmentation using neural end-to-end overlapped speech detection,” in Proc. ICASSP 2020 , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Y. Takashima, Y. Fujita, S. Watanabe, S. Horiguchi, P. García, and K. Nagamatsu, “End-to-end speaker diarization conditioned on speech activity and overlap detection,” in 2021 IEEE Spoken Language Technology Workshop (SLT) , 2021, pp. 849–856
2021
Closest in time.
F. Landini, O. Glembek, P. Matějka, J. Rohdin, L. Burget, M. Diez, and A. Silnova, “Analysis of the BUT Diarization System for VoxConverse Challenge,” in 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021
2021
Closest in time.
Silero Team, “Silero VAD: pre-trained enterprise-grade Voice Activity Detector (VAD), Number Detector and Language Classifier,” https://github.com/snakers4/silero-vad , 2021
2021
Closest in time.