Fetching the paper…
Reading the bibliography…
Recent diarization technologies can be categorized into two approaches, i.e., clustering and end-to-end neural approaches, which have different pros and cons.
“Constrained k-means clustering with background knowledge,”
K. Wagstaff, C. Cardie, S. Rogers, and S S. Schroedl, · 2001
Earlier work this paper cites.
“Wavesplit: End-to-end speech separation by speaker clustering,” 2020,
N. Zeghidour and D. Grangier, · 2002
Earlier work this paper cites.
S. Horiguchi, Y. Fujita, S. Watanabe, Y. Xue, and K. Nagamatsu, · 2005
Earlier work this paper cites.
“The AMI meeting corpus: A pre-announcement,”
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, G. Lathoud, M. Lincoln, A. Lisowska, I. McCowan, W. Post, D. Reidsma, , and P. Wellner, · 2006
Earlier work this paper cites.
“Online end-to-end neural diarization with speaker-tracing buffer,” 2020,
Y. Xue, S. Horiguchi, Y. Fujita, S. Watanabe, and K. Nagamatsu, · 2006
Earlier work this paper cites.
“Online speaker diarization with relation network,” 2020,
X. Li, Y. Zhao, C. Luo, and W. Zeng, · 2009
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, , and P. Ouellet, · 2011
Earlier work this paper cites.
“Speaker diarization: A review of recent research,”
X. Anguera, S. Bozonnet, N. Evans, C. Fredouille, G. Friedland, and O. Vinyals, · 2012
Earlier work this paper cites.
“Librispeech: An asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Cited alongside, same era.
“MUSAN: A music, speech, and noise corpus,,” 2015,
D. Snyder, G. Chen, and D. Povey, · 2015
Cited alongside, same era.
“Deep neural network-based speaker embeddings for end-to-end speaker verification,”
D. Snyder, P. Ghahremani, D. Povey, D. Garcia-Romero, Y. Carmiel, , and S. Khudanpur, · 2016
Cited alongside, same era.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
M. Kolbæk, D. Yu, Z. Tan, and J. Jensen, · 2017
Cited alongside, same era.
“A study on data augmentation of reverberant speech for robust speech recognition,”
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, · 2017
Cited alongside, same era.
“BUT system for DIHARD speech diarization challenge 2018,”
M. Diez, F. Landini, L. Burget, J. Rohdin, A. Silnova, K. Zmolikova, O. Novotný, K. Veselý, O. Glembek, O. Plchot, L. Mošner, and P. Matějka, · 2018
Later among the works it cites.
“Listening to each speaker one by one with recurrent selective hearing networks,”
K. Kinoshita, L. Drude, M. Delcroix, and T. Nakatani, · 2018
Later among the works it cites.
“Fully supervised speaker diarization,”
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, and C. Wang, · 2019
Later among the works it cites.
“All-neural online source separation, counting, and diarization for meeting analysis,”
T. von Neumann, K. Kinoshita, M. Delcroix, S. Araki, T. Nakatani, and R. Haeb-Umbach, · 2019
Later among the works it cites.
“End-to-end neural speaker diarization with permutation-free objectives,”
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, and S. Watanabe, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Cited alongside, same era.
First DIHARD Challenge Evaluation Plan
N. Ryant, K. Church, C. Cieri, A. Cristia, J. Du, S. Ganapathy, and M. Liberman, · 2018
Cited alongside, same era.
“Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,”
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, and S. Khudanpur, · 2018
Cited alongside, same era.
Y. Fujita, N. Kanda, S. Horiguchi, Y. Xue, K. Nagamatsu, and S. Watanabe, · 2019
Later among the works it cites.
“Low-latency speaker-independent continuous speech separation,”
T. Yoshioka, Z. Chen, C. Liu, X. Xiao, H. Erdogan, and D. Dimitriadis, · 2019
Later among the works it cites.