Fetching the paper…
Reading the bibliography…
This paper investigates an end-to-end neural diarization (EEND) method for an unknown number of speakers.
“2000 NIST Speaker Recognition Evaluation,” https://catalog.ldc.upenn.edu/LDC2001S97
2000
Earlier work this paper cites.
K. Maekawa, “Corpus of spontaneous Japanese: Its design and evaluation,” in Proc. ISCA & IEEE Workshop on Spontaneous Speech Process. Recognit. , 2003, pp. 7–12
2003
Earlier work this paper cites.
J. Carletta, “Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus,” Lang. Resour. and Eval. , vol. 41, no. 2, pp. 181–190, 2007
2007
Earlier work this paper cites.
S. H. Shum, N. Dehak, R. Dehak, and J. R. Glass, “Unsupervised methods for speaker diarization: An integrated and iterative approach,” IEEE Trans. Audio, Speech, Lang. Process. , vol. 21, no. 10, pp. 2015–2028, 2013
2013
Earlier work this paper cites.
G. Sell and D. Garcia-Romero, “Speaker diarization with PLDA i-vector scoring and unsupervised calibration,” in Proc. IEEE Spoken Lang. Technol. Workshop , 2014, pp. 413–417
2014
Earlier work this paper cites.
M. Senoussaoui, P. Kenny, T. Stafylakis, and P. Dumouchel, “A study of the cosine distance-based mean shift for telephone speech diarization,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 22, no. 1, pp. 217–227, 2014
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. Int. Conf. Learn. Representations , 2015
2015
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2016, pp. 31–35
2016
Earlier work this paper cites.
M. À. India Massana, J. A. Rodríguez Fonollosa, and F. J. Hernando Pericás, “Lstm neural network-based speaker segmentation using acoustic and language modelling,” in Proc. Interspeech , 2017, pp. 2834–2838
2017
Earlier work this paper cites.
D. Dimitriadis and P. Fousek, “Developing on-line speaker diarization system,” in Proc. Interspeech , 2017, pp. 2739–2743
2017
Earlier work this paper cites.
D. Garcia-Romero, D. Snyder, G. Sell, D. Povey, and A. McCree, “Speaker diarization using deep neural network embeddings,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2017, pp. 4930–4934
2017
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2017, pp. 241–245
2017
Earlier work this paper cites.
Y. Luo, Z. Chen, J. R. Hershey, J. Le Roux, and N. Mesgarani, “Deep clustering and conventional networks for music separation: Stronger together,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2017, pp. 61–65
2017
Earlier work this paper cites.
Z. Chen, Y. Luo, and N. Mesgarani, “Deep attractor network for single-microphone speaker separation,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2017, pp. 246–250
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. Advances Neural Inf. Process. Syst. , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
C. Boeddeker, J. Heitkaemper, J. Schmalenstoeer, L. Drude, J. Heymann, and R. Haeb-Umbach, “Front-end processing for the CHiME-5 dinner party scenario,” in Proc. 5th Int. Workshop Speech Process. Everyday Environ. (CHiME-5) , 2018
2018
Earlier work this paper cites.
Q. Wang, C. Downey, L. Wan, P. Andrew Mansfield, and I. Lopez Moreno, “Speaker diarization with LSTM,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2018, pp. 5239–5243
2018
Earlier work this paper cites.
K. Kinoshita, L. Drude, M. Delcroix, and T. Nakatani, “Listening to each speaker one by one with recurrent selective hearing networks,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2018, pp. 5064–5068
2018
Earlier work this paper cites.
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, and S. Khudanpur, “Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,” in Proc. Interspeech , 2018, pp. 2808–2812
2018
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2018, pp. 5329–5333
2018
Earlier work this paper cites.
M. Maciejewski, D. Snyder, V. Manohar, N. Dehak, and S. Khudanpur, “Characterizing performance of speaker diarization systems on far-field speech using standard methods,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2018, pp. 5244–5248
2018
Earlier work this paper cites.
Y. Luo and N. Mesgarani, “TasNet: Time-domain audio separation network for real-time, single-channel speech separation,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2018, pp. 696–700
2018
Earlier work this paper cites.
Y. Luo, Z. Chen, and N. Mesgarani, “Speaker-independent speech separation with deep attractor network,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 26, no. 4, pp. 787–796, 2018
2018
Cited alongside, same era.
B. B. Meier, I. Elezi, M. Amirian, O. Dürr, and T. Stadelmann, “Learning neural models for end-to-end clustering,” in Proc. IAPR Workshop Artif. Neural Netw. Pattern Recognit. , 2018, pp. 126–138
2018
Cited alongside, same era.
L. Drude, T. von Neumann, and R. Haeb-Umbach, “Deep attractor networks for speaker re-identification and blind source separation,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2018, pp. 11–15
2018
Cited alongside, same era.
T. von Neumann, K. Kinoshita, M. Delcroix, S. Araki, T. Nakatani, and R. Haeb-Umback, “All-neural online source separation, counting, and diarization for meeting analysis,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2019, pp. 91–95
2019
Cited alongside, same era.
M. Diez, L. Burget, F. Landini, and J. Černocký, “Analysis of speaker diarization based on bayesian HMM with eigenvoice priors,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 28, pp. 355–368, 2020
2020
Later among the works it cites.
Z. Huang, S. Watanabe, Y. Fujita, P. García, Y. Shao, D. Povey, and S. Khudanpur, “Speaker diarization with region proposal network,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2020, pp. 6514–6518
2020
Later among the works it cites.
I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Korenevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, I. Podluzhny, A. Laptev, and A. Romanenko, “Target-speaker voice activity detection: a novel approach for multi-speaker diarization in a dinner party scenario,” in Proc. Interspeech , 2020, pp. 274–278
2020
Later among the works it cites.
F. Landini, S. Wang, M. Diez, L. Burget, P. Matějka, K. Žmolíková, L. Mošner, A. Silnova, O. Plchot, O. Novotnỳ, H. Zeinali, and J. Rohdin, “BUT system for the Second DIHARD Speech Diarization Challenge,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2020, pp. 6529–6533
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, and S. Watanabe, “End-to-end neural speaker diarization with permutation-free objectives,” in Proc. Interspeech , 2019, pp. 4300–4304
2019
Cited alongside, same era.
Y. Fujita, N. Kanda, S. Horiguchi, Y. Xue, K. Nagamatsu, and S. Watanabe, “End-to-end neural speaker diarization with self-attention,” in Proc. IEEE Autom. Speech Recognit. Understanding Workshop , 2019, pp. 296–303
2019
Cited alongside, same era.
T. J. Park, K. J. Han, J. Huang, X. He, B. Zhou, P. Georgiou, and S. Narayanan, “Speaker diarization with lexical information,” in Proc. Interspeech , 2019, pp. 391–395
2019
Cited alongside, same era.
M. Diez, L. Burget, S. Wang, J. Rohdin, and J. Černockỳ, “Bayesian HMM based x-vector clustering for speaker diarization,” in Proc. Interspeech , 2019, pp. 346–350
2019
Cited alongside, same era.
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, and C. Wang, “Fully supervised speaker diarization,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2019, pp. 6301–6305
2019
Cited alongside, same era.
N. Ryant, K. Church, C. Cieri, A. Cristia, J. Du, S. Ganapathy, and M. Liberman, “The Second DIHARD Diarization Challenge: Dataset, task, and baselines,” in Proc. Interspeech , 2019, pp. 978–982
2019
Cited alongside, same era.
——, “Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Cited alongside, same era.
N. Takahashi, S. Parthasaarathy, N. Goswami, and Y. Mitsufuji, “Recursive speech separation for unknown number of speakers,” in Proc. Interspeech , 2019, pp. 1348–1352
2019
Cited alongside, same era.
2020
Later among the works it cites.
M. Delcroix, K. Zmolikova, T. Ochiai, K. Kinoshita, and T. Nakatani, “Speaker activity driven neural speech extraction,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2021, pp. 6099–6103
2021
Closest in time.
Y. Takashima, Y. Fujita, S. Watanabe, S. Horiguchi, P. Garcia, and K. Nagamatsu, “End-to-end speaker diarization conditioned on speech activity and overlap detection,” in Proc. IEEE Spoken Lang. Technol. Workshop , 2021, pp. 849–856
2021
Closest in time.
S. Horiguchi, N. Yalta, P. Garcia, Y. Takashima, Y. Xue, D. Raj, Z. Huang, Y. Fujita, S. Watanabe, and S. Khudanpur, “The Hitachi-JHU DIHARD III system: Competitive end-to-end neural diarization and x-vector clustering systems combined by DOVER-Lap,” in Proc. 3rd DIHARD Speech Diarization Challenge Workshop , 2021
2021
Closest in time.
E. Han, C. Lee, and A. Stolcke, “BW-EDA-EEND: Streaming end-to-end neural speaker diarization for a variable number of speakers,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2021, pp. 7193–7197
2021
Closest in time.
Y. Xue, S. Horiguchi, Y. Fujita, Y. Takashima, S. Watanabe, P. Garcia, and K. Nagamatsu, “Online streaming end-to-end neural diarization handling overlapping speech and flexible numbers of speakers,” in Proc. Interspeech , 2021, pp. 3116–3120
2021
Closest in time.
X. Xiao, N. Kanda, Z. Chen, T. Zhou, T. Yoshioka, S. Chen, Y. Zhao, G. Liu, Y. Wu, J. Wu, S. Liu, J. Li, and Y. Gong, “Microsoft speaker diarization system for the VoxCeleb speaker recognition challenge 2020,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2021, pp. 5824–5828
2021
Closest in time.
D. Raj, Z. Huang, and S. Khudanpur, “Multi-class spectral clustering with overlaps for speaker diarization,” in Proc. IEEE Spoken Lang. Technol. Workshop , 2021, pp. 582–589
2021
Closest in time.
Q. Li, F. L. Kreyssig, C. Zhang, and P. C. Woodland, “Discriminative neural clustering for speaker diarisation,” in Proc. IEEE Spoken Lang. Technol. Workshop , 2021, pp. 574–581
2021
Closest in time.
N. Ryant, P. Singh, V. Krishnamohan, R. Varma, K. Church, C. Cieri, J. Du, S. Ganapathy, and M. Liberman, “The Third DIHARD Diarization Challenge,” in Proc. Interspeech , 2021, pp. 3570–3574
2021
Closest in time.
N. Zeghidour and D. Grangier, “Wavesplit: End-to-end speech separation by speaker clustering,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 29, pp. 2840–2849, 2021
2021
Closest in time.
S. Horiguchi, P. Garcia, Y. Fujita, S. Watanabe, and K. Nagamatsu, “End-to-end speaker diarization as post-processing,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2021, pp. 7188–7192
2021
Closest in time.
D. Raj, L. P. Garcia-Perera, Z. Huang, S. Watanabe, D. Povey, A. Stolcke, and S. Khudanpur, “DOVER-Lap: A method for combining overlap-aware diarization outputs,” in Proc. IEEE Spoken Lang. Technol. Workshop , 2021, pp. 881–888
2021
Closest in time.
F. Landini, A. Lozano-Diez, L. Burget, M. Diez, A. Silnova, K. Žmolíková, O. Glembek, P. Matějka, T. Stafylakis, and N. Brümmer, “BUT system description for the Third DIHARD Speech Diarization Challenge,” in Proc. 3rd DIHARD Speech Diarization Challenge Workshop , 2021
2021
Closest in time.
S. Maiti, H. Erdogan, K. Wilson, S. Wisdom, S. Watanabe, and J. R. Hershey, “End-to-end diarization for variable number of speakers with local-global networks and discriminative speaker embeddings,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2021, pp. 7183–7187
2021
Closest in time.
S. Horiguchi, P. García, S. Watanabe, Y. Xue, Y. Takashima, and Y. Kawaguchi, “Towards neural diarization for unlimited numbers of speakers using global and local attractors,” in Proc. IEEE Autom. Speech Recognit. Understanding Workshop , 2021, pp. 98–105
2021
Closest in time.
Y. C. Liu, E. Han, C. Lee, and A. Stolcke, “End-to-end neural diarization: From transformer to conformer,” in Proc. Interspeech , 2021, pp. 3081–3085
2021
Closest in time.
T. J. Park, N. Kanda, D. Dimitriadis, K. J. Han, S. Watanabe, and S. Narayanan, “A review of speaker diarization: Recent advances with deep learning,” Comput. Speech Lang. , vol. 72, p. 101317, 2022
2022
Closest in time.
F. Landini, J. Profant, M. Diez, and L. Burget, “Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks,” Comput. Speech Lang. , vol. 71, p. 101254, 2022
2022
Closest in time.