Fetching the paper…
Reading the bibliography…
This paper investigates the utilization of an end-to-end diarization model as post-processing of conventional clustering-based diarization.
“Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus,”
J. Carletta, · 2007
Earlier work this paper cites.
“Speaker diarization: A review of recent research,”
X. Anguera, S. Bozonnet, N. Evans, C. Fredouille, G. Friedland, and O. Vinyals, · 2012
Earlier work this paper cites.
“Speaker diarization with PLDA i-vector scoring and unsupervised calibration,”
G. Sell and D. Garcia-Romero, · 2014
Earlier work this paper cites.
“MUSAN: A music, speech, and noise corpus,” arXiv:1510.08484, 2015
D. Snyder, G. Chen, and D. Povey, · 2015
Earlier work this paper cites.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, · 2017
Earlier work this paper cites.
“A study on data augmentation of reverberant speech for robust speech recognition,”
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Front-end processing for the CHiME-5 dinner party scenario,”
C. Boeddeker, J. Heitkaemper, J. Schmalenstoeer, L. Drude, J. Heymann, and R. Haeb-Umbach, · 2018
Earlier work this paper cites.
“Speaker diarization based on Bayesian HMM with eigenvoice priors,”
M. Diez, L. Burget, and P. Matejka, · 2018
Earlier work this paper cites.
“Listening to each speaker one by one with recurrent selective hearing networks,”
K. Kinoshita, L. Drude, M. Delcroix, and T. Nakatani, · 2018
Earlier work this paper cites.
“Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,”
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, and S. Khudanpur, · 2018
Cited alongside, same era.
“The Second DIHARD Diarization Challenge: Dataset, task, and baselines,”
N. Ryant, K. Church, C. Cieri, A. Cristia, J. Du, S. Ganapathy, and M. Liberman, · 2019
Cited alongside, same era.
“End-to-end neural speaker diarization with permutation-free objectives,”
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, and S. Watanabe, · 2019
Cited alongside, same era.
“End-to-end neural speaker diarization with self-attention,”
Y. Fujita, N. Kanda, S. Horiguchi, Y. Xue, K. Nagamatsu, and S. Watanabe, · 2019
Cited alongside, same era.
“Fully supervised speaker diarization,”
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, and C. Wang, · 2019
Cited alongside, same era.
“Neural speaker diarization with speaker-wise chain rule,” arXiv:2006.01796, 2020
Y. Fujita, S. Watanabe, S. Horiguchi, Y. Xue, J. Shi, and K. Nagamatsu, · 2020
Closest in time.
“BUT system for the Second DIHARD Speech Diarization Challenge,”
F. Landini, S. Wang, M. Diez, L. Burget, P. Matějka, K. Žmolíková, L. Mošner, A. Silnova, O. Plchot, O. Novotnỳ, H. Zeinali, and J. Rohdin, · 2020
Closest in time.
“DIHARD II is still hard: Experimental results and discussions from the DKU-LENOVO team,”
Q. Lin, W. Cai, L. Yang, J. Wang, J. Zhang, and M. Li, · 2020
Closest in time.
“Speaker diarization with region proposal network,”
Z. Huang, S. Watanabe, Y. Fujita, P. García, Y. Shao, D. Povey, and S. Khudanpur, · 2020
Closest in time.
“Tackling real noisy reverberant meetings with all-neural source separation, counting, and diarization system,”
K. Kinoshita, M. Delcroix, S. Araki, and T. Nakatani, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Landini, S. Wang, M. Diez, L. Burget, P. Matějka, K. Žmolíková, L. Mošner, O. Plchot, O. Novotnỳ, H. Zeinali, and J. Rohdin, · 2019
Cited alongside, same era.
“Bayesian HMM based x-vector clustering for speaker diarization,”
M. Diez, L. Burget, S. Wang, J. Rohdin, and J. Černockỳ, · 2019
Cited alongside, same era.
“The STC system for the CHiME-6 Challenge,”
I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Korenevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, I. Podluzhny, A. Laptev, and A. Romanenko, · 2020
Cited alongside, same era.
Y. Fujita, S. Watanabe, S. Horiguchi, Y. Xue, and K. Nagamatsu, · 2020
Cited alongside, same era.
“End-to-end speaker diarization for an unknown number of speakers with encoder-decoder based attractors,”
S. Horiguchi, Y. Fujita, S. Wananabe, Y. Xue, and K. Nagamatsu, · 2020
Cited alongside, same era.
“Target-speaker voice activity detection: a novel approach for multi-speaker diarization in a dinner party scenario,”
I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Korenevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, Podluzhny, et al., · 2020
Closest in time.
“Personal VAD: Speaker-conditioned voice activity detection,”
S. Ding, Q. Wang, S.-y. Chang, L. Wan, and I. L. Moreno, · 2020
Closest in time.
“VoiceFilter-Lite: Streaming targeted voice separation for on-device speech recognition,”
Q. Wang, I. L. Moreno, M. Saglam, K. Wilson, A. Chiao, R. Liu, Y. He, W. Li, J. Pelecanos, M. Nika, and A. Gruenstein, · 2020
Closest in time.
“Speaker detection in the wild: Lessons learned from JSALT 2019,”
P. Garcia, J. Villalba, H. Bredin, J. Du, D. Castan, A. Cristia, L. Bullock, L. Guo, K. Okabe, P. S. Nidadavolu, et al., · 2020
Closest in time.
“Overlap-aware diarization: Resegmentation using neural end-to-end overlapped speech detection,”
L. Bullock, H. Bredin, and L. P. Garcia-Perera, · 2020
Closest in time.