Fetching the paper…
Reading the bibliography…
Speaker diarization has been mainly developed based on the clustering of speaker embeddings.
“2000 speaker recognition evaluation plan,” https://www.nist.gov/sites/default/files/documents/2017/09/26/spk-2000-plan-v1.0.htm_.pdf, 2000
NIST, · 2000
Earlier work this paper cites.
“The ICSI meeting corpus,”
A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stolcke, and C. Wooters, · 2003
Earlier work this paper cites.
“Corpus of spontaneous japanese: Its design and evaluation,”
K. Maekawa, · 2003
Earlier work this paper cites.
“Framewise phoneme classification with bidirectional lstm and other neural network architectures,”
A. Graves and J. Schmidhuber, · 2005
Earlier work this paper cites.
“An overview of automatic speaker diarization systems,”
S. E. Tranter and D. A. Reynolds, · 2006
Earlier work this paper cites.
“Overlap in meetings: ASR effects and analysis by dialog factors, speakers, and collection site,”
Ö. Çetin and E. Shriberg, · 2006
Earlier work this paper cites.
“Interpretation of multiparty meetings the AMI and Amida projects,”
S. Renals, T. Hain, and H. Bourlard, · 2008
Earlier work this paper cites.
“The 2009 (RT-09) rich transcription meeting recognition evaluation plan,” http://www.itl.nist.gov/iad/mig/tests/rt/2009/docs/rt09-meeting-eval-plan-v2.pdf, 2009
NIST, · 2009
Earlier work this paper cites.
“LIUM_SPKDIARIZATION: An open source toolkit for diarization,”
S. Meignier, · 2010
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, · 2011
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, · 2011
Earlier work this paper cites.
“Speaker diarization: A review of recent research,”
X. Anguera, S. Bozonnet, N. Evans, C. Fredouille, G. Friedland, and O. Vinyals, · 2012
Earlier work this paper cites.
“Unsupervised methods for speaker diarization: An integrated and iterative approach,”
S. H. Shum, N. Dehak, R. Dehak, and J. R. Glass, · 2013
Earlier work this paper cites.
“Speaker diarization with PLDA i-vector scoring and unsupervised calibration,”
G. Sell and D. Garcia-Romero, · 2014
Earlier work this paper cites.
“A study of the cosine distance-based mean shift for telephone speech diarization,”
M. Senoussaoui, P. Kenny, T. Stafylakis, and P. Dumouchel, · 2014
Earlier work this paper cites.
“Sequence to sequence learning with neural networks,”
I. Sutskever, O. Vinyals, and Q. V. Le, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“Attention-based models for speech recognition,”
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“MUSAN: A music, speech, and noise corpus,”
D. Snyder, G. Chen, and D. Povey, · 2015
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization,”
D. P. Kingma and J. Ba, · 2015
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Cited alongside, same era.
J. Lei Ba, J. R. Kiros, and G. E. Hinton, · 2016
Cited alongside, same era.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, · 2016
Cited alongside, same era.
“Developing on-line speaker diarization system,”
D. Dimitriadis and P. Fousek, · 2017
Cited alongside, same era.
“Speaker diarization using deep neural network embeddings,”
D. Garcia-Romero, D. Snyder, G. Sell, D. Povey, and A. McCree, · 2017
“Speaker diarization with LSTM,”
Q. Wang, C. Downey, L. Wan, P. A. Mansfield, and I. L. Moreno, · 2018
Later among the works it cites.
“Generalized end-to-end loss for speaker verification,”
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, · 2018
Later among the works it cites.
“X-vectors: Robust DNN embeddings for speaker recognition,”
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, · 2018
Later among the works it cites.
“Links: A high-dimensional online clustering method,”
P. A. Mansfield, Q. Wang, C. Downey, L. Wan, and I. L. Moreno, · 2018
Later among the works it cites.
“Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,”
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, and S. Khudanpur, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“A structured self-attentive sentence embedding,”
Z. Lin, M. Feng, C. Nogueira dos Santos, M. Yu, B. Xiang, B. Zhou, and Y. Bengio, · 2017
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Cited alongside, same era.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, · 2017
Cited alongside, same era.
“Tacotron: Towards end-to-end speech synthesis,”
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, · 2017
Cited alongside, same era.
“Char2wav: End-to-end speech synthesis,”
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, · 2017
Cited alongside, same era.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
D. Yu, M. Kolbæk, Z. Tan, and J. Jensen, · 2017
Cited alongside, same era.
“BUT system for DIHARD speech diarization challenge 2018,”
M. Diez, F. Landini, L. Burget, J. Rohdin, A. Silnova, K. Z̆molíková, O. Novotný, K. Veselý, O. Glembek, O. Plchot, L. Mos̆ner, and P. Matĕjka, · 2018
Later among the works it cites.
“Speaker diarization with enhancing speech for the first DIHARD challenge,”
L. Sun, J. Du, C. Jiang, X. Zhang, S. He, B. Yin, and C.-H. Lee, · 2018
Later among the works it cites.
“Joint discriminative embedding learning, speech activity and overlap detection for the dihard speaker diarization challenge,”
V. A. Miasato Filho, D. A. Silva, and L. G. Depra Cuozzo, · 2018
Later among the works it cites.
“Non-local neural networks,”
X. Wang, R. B. Girshick, A. Gupta, and K. He, · 2018
Later among the works it cites.
“Self-attentional acoustic models,”
M. Sperber, J. Niehues, G. Neubig, S. Stüker, and A. Waibel, · 2018
Later among the works it cites.
“Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,”
L. Dong, S. Xu, and B. Xu, · 2018
Later among the works it cites.
“Self-attention mechanism based system for DCASE2018 challenge task1 and task4,”
W. Jun and L. Shengchen, · 2018
Later among the works it cites.
“Self-attentive speaker embeddings for text-independent speaker verification,”
Y. Zhu, T. Ko, D. Snyder, B. Mak, and D. Povey, · 2018
Later among the works it cites.
“Acoustic modeling for distant multi-talker speech recognition with single- and multi-channel branches,”
N. Kanda, Y. Fujita, S. Horiguchi, R. Ikeshita, K. Nagamatsu, and S. Watanabe, · 2019
Closest in time.
“End-to-end neural speaker diarization with permutation-free objectives,”
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, and S. Watanabe, · 2019
Closest in time.
“Speaker recognition for multi-speaker conversations using x-vectors,”
D. Snyder, D. Garcia-Romero, G. Sell, A. McCree, D. Povey, and S. Khudanpur, · 2019
Closest in time.
“Designing an effective metric learning pipeline for speaker diarization,”
V. S. Narayanaswamy, J. J. Thiagarajan, H. Song, and A. Spanias, · 2019
Closest in time.
“Fully supervised speaker diarization,”
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, and C. Wang, · 2019
Closest in time.
“Cross-modal self-attention network for referring image segmentation,”
L. Ye, M. Rochan, Z. Liu, and Y. Wang, · 2019
Closest in time.