Fetching the paper…
Reading the bibliography…
The most common approach to speaker diarization is clustering of speaker embeddings.
NIST, “2000 speaker recognition evaluation plan,” https://www.nist.gov/sites/default/files/documents/2017/09/26/spk-2000-plan-v1.0.htm_.pdf, 2000
2000
Earlier work this paper cites.
A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stolcke, and C. Wooters, “The ICSI meeting corpus,” in Proc. ICASSP , vol. I, 2003, pp. 364–367
2003
Earlier work this paper cites.
K. Maekawa, “Corpus of spontaneous japanese: Its design and evaluation,” in ISCA & IEEE Workshop on Spontaneous Speech Processing and Recognition , 2003
2003
Earlier work this paper cites.
A. Graves and J. Schmidhuber, “Framewise phoneme classification with bidirectional lstm and other neural network architectures,” Neural Networks , vol. 18, no. 5, pp. 602 – 610, 2005, iJCNN 2005
2005
Earlier work this paper cites.
NIST, “The 2009 (RT-09) rich transcription meeting recognition evaluation plan,” http://www.itl.nist.gov/iad/mig/tests/rt/2009/docs/rt09-meeting-eval-plan-v2.pdf, 2009
2005
Earlier work this paper cites.
S. E. Tranter and D. A. Reynolds, “An overview of automatic speaker diarization systems,” IEEE Trans. on ASLP , vol. 14, no. 5, pp. 1557–1565, 2006
2006
Earlier work this paper cites.
Ö. Çetin and E. Shriberg, “Overlap in meetings: ASR effects and analysis by dialog factors, speakers, and collection site,” in Proc. MLMI , 2006, pp. 212–224
2006
Earlier work this paper cites.
S. Renals, T. Hain, and H. Bourlard, “Interpretation of multiparty meetings the AMI and Amida projects,” in 2008 Hands-Free Speech Communication and Microphone Arrays , 2008, pp. 115–118
2008
Earlier work this paper cites.
S. Meignier, “LIUM_SPKDIARIZATION: An open source toolkit for diarization,” in CMU SPUD Workshop , 2010
2010
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Trans. on ASLP , vol. 19, no. 4, pp. 788–798, 2011
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The Kaldi speech recognition toolkit,” in Proc. ASRU , 2011
2011
Earlier work this paper cites.
X. Anguera, S. Bozonnet, N. Evans, C. Fredouille, G. Friedland, and O. Vinyals, “Speaker diarization: A review of recent research,” IEEE Trans. on ASLP , vol. 20, no. 2, pp. 356–370, 2012
2012
Earlier work this paper cites.
S. H. Shum, N. Dehak, R. Dehak, and J. R. Glass, “Unsupervised methods for speaker diarization: An integrated and iterative approach,” IEEE Trans. on ASLP , vol. 21, no. 10, pp. 2015–2028, 2013
2013
Earlier work this paper cites.
G. Sell and D. Garcia-Romero, “Speaker diarization with PLDA i-vector scoring and unsupervised calibration,” in Proc. SLT , 2014, pp. 413–417
2014
Earlier work this paper cites.
M. Senoussaoui, P. Kenny, T. Stafylakis, and P. Dumouchel, “A study of the cosine distance-based mean shift for telephone speech diarization,” IEEE/ACM Trans. on ASLP , vol. 22, no. 1, pp. 217–227, 2014
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Proc. NIPS , 2014, pp. 3104–3112
2014
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in Proc. ICLR , 2015
2015
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Proc. NIPS , 2015, pp. 577–585
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” in Proc. ICLR , 2015
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in Proc. ICASSP , 2016, pp. 4960–4964
2016
Cited alongside, same era.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in Proc. ICASSP , 2016, pp. 31–35
2016
Cited alongside, same era.
2016
Cited alongside, same era.
D. Dimitriadis and P. Fousek, “Developing on-line speaker diarization system,” in Proc. Interspeech , 2017, pp. 2739–2743
2017
Cited alongside, same era.
D. Garcia-Romero, D. Snyder, G. Sell, D. Povey, and A. McCree, “Speaker diarization using deep neural network embeddings,” in Proc. ICASSP , 2017, pp. 4930–4934
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in Proc. ICASSP , 2018, pp. 4879–4883
2018
Later among the works it cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” in Proc. ICASSP , 2018, pp. 5329–5333
2018
Later among the works it cites.
2018
Later among the works it cites.
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, and S. Khudanpur, “Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,” in Proc. Interspeech , 2018, pp. 2808–2812
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Z. Lin, M. Feng, C. Nogueira dos Santos, M. Yu, B. Xiang, B. Zhou, and Y. Bengio, “A structured self-attentive sentence embedding,” in Proc. ICLR , 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NIPS , 2017, pp. 5998–6008
2017
Cited alongside, same era.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid CTC/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing , vol. 11, no. 8, pp. 1240–1253, 2017
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in Proc. Interspeech , 2017, pp. 4006–4010
2017
Cited alongside, same era.
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, “Char2wav: End-to-end speech synthesis,” in ICLR Workshop , 2017
2017
Cited alongside, same era.
D. Yu, M. Kolbæk, Z. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in Proc. ICASSP , 2017, pp. 241–245
2017
Cited alongside, same era.
M. Kolbæk, D. Yu, Z. Tan, and J. Jensen, “Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,” IEEE/ACM Trans. on ASLP , vol. 25, no. 10, pp. 1901–1913, 2017
2017
Cited alongside, same era.
M. Diez, F. Landini, L. Burget, J. Rohdin, A. Silnova, K. Z̆molíková, O. Novotný, K. Veselý, O. Glembek, O. Plchot, L. Mos̆ner, and P. Matĕjka, “BUT system for DIHARD speech diarization challenge 2018,” in Proc. Interspeech , 2018, pp. 2798–2802
2018
Later among the works it cites.
L. Sun, J. Du, C. Jiang, X. Zhang, S. He, B. Yin, and C.-H. Lee, “Speaker diarization with enhancing speech for the first DIHARD challenge,” in Proc. Interspeech , 2018, pp. 2793–2797
2018
Later among the works it cites.
V. A. Miasato Filho, D. A. Silva, and L. G. Depra Cuozzo, “Joint discriminative embedding learning, speech activity and overlap detection for the dihard speaker diarization challenge,” in Proc. Interspeech , 2018, pp. 2818–2822
2018
Later among the works it cites.
X. Wang, R. B. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proc. CVPR , 2018, pp. 7794–7803
2018
Later among the works it cites.
M. Sperber, J. Niehues, G. Neubig, S. Stüker, and A. Waibel, “Self-attentional acoustic models,” in Proc. Interspeech , 2018, pp. 3723–3727
2018
Later among the works it cites.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,” Proc. ICASSP , pp. 5884–5888, 2018
2018
Later among the works it cites.
W. Jun and L. Shengchen, “Self-attention mechanism based system for DCASE2018 challenge task1 and task4,” in DCASE2018 Challenge , 2018
2018
Later among the works it cites.
Y. Zhu, T. Ko, D. Snyder, B. Mak, and D. Povey, “Self-attentive speaker embeddings for text-independent speaker verification,” in Proc. Interspeech , 2018, pp. 3573–3577
2018
Later among the works it cites.
N. Kanda, Y. Fujita, S. Horiguchi, R. Ikeshita, K. Nagamatsu, and S. Watanabe, “Acoustic modeling for distant multi-talker speech recognition with single- and multi-channel branches,” in Proc. ICASSP , 2019, pp. 6630–6634
2019
Later among the works it cites.
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, and S. Watanabe, “End-to-end neural speaker diarization with permutation-free objectives,” in Proc. Interspeech , 2019 (to appear)
2019
Later among the works it cites.
Y. Fujita, N. Kanda, S. Horiguchi, Y. Xue, K. Nagamatsu, and S. Watanabe, “End-to-end neural speaker diarization with self-attention,” in Proc. ASRU , 2019 (submitted)
2019
Later among the works it cites.
D. Snyder, D. Garcia-Romero, G. Sell, A. McCree, D. Povey, and S. Khudanpur, “Speaker recognition for multi-speaker conversations using x-vectors,” in Proc. ICASSP , 2019, pp. 5796–5800
2019
Later among the works it cites.
V. S. Narayanaswamy, J. J. Thiagarajan, H. Song, and A. Spanias, “Designing an effective metric learning pipeline for speaker diarization,” in Proc. ICASSP , 2019, pp. 5806–5810
2019
Later among the works it cites.
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, and C. Wang, “Fully supervised speaker diarization,” in Proc. ICASSP , 2019, pp. 6301–6305
2019
Later among the works it cites.
L. Ye, M. Rochan, Z. Liu, and Y. Wang, “Cross-modal self-attention network for referring image segmentation,” in Proc. CVPR , 2019, pp. 10 502–10 511
2019
Later among the works it cites.