Fetching the paper…
Reading the bibliography…
Significant progress has recently been made in speaker diarisation after the introduction of d-vectors as speaker embeddings extracted from neural network (NN) speaker classifiers for clustering speech segments.
LSTM based similarity measurement with spectral clustering for speaker diarization
Lin, Q., Yin, R., Li, M., Bredin, H., Barras, C., 2019 · 1907
Earlier work this paper cites.
End-to-end neural speaker diarization with permutation-free objectives
Fujita, Y., Kanda, N., Horiguchi, S., Nagamatsu, K., Watanabe, S., 2019 · 1909
Earlier work this paper cites.
Arobust algorithm for accurate endpointing of speech signals
Savoji, M., 1989 · 1989
Earlier work this paper cites.
Phoneme recognition using time-delay neural networks
Waibel, A., Hanazawa, T., Hinton, G., Shikano, K., Lang, K., 1989 · 1989
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., Schmidhuber, J., 1997 · 1997
Earlier work this paper cites.
Automatic segmentation, classification and clustering of broadcast news audio, in: Proceedings of the DARPA speech recognition workshop
Siegler, M., Jain, U., Raj, B., Stern, R., 1997 · 1997
Earlier work this paper cites.
Speaker, environment and channel change detection and clustering via the Bayesian information criterion, in: Proceedings of the DARPA Broadcast News Transcription and Understanding Workshop, pp. 1–6
Chen, S., Gopalakrishnan, P., 1998 · 1998
Earlier work this paper cites.
Cluster adaptive training of hidden Markov models
Gales, M., 2000 · 2000
Earlier work this paper cites.
Separating style and content with bilinear models
Tenenbaum, J., Freeman, W., 2000 · 2000
Earlier work this paper cites.
Generating and evaluating segmentations for automatic speech recognition of conversational telephone speech, in: Proceedings of the 29th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 753–756
Tranter, S., Yu, K., Evermann, G., Woodland, P., 2004 · 2004
Earlier work this paper cites.
Speaker diarization for multi-party meetings using acoustic fusion, in: Proceedings of the 5th IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), pp. 426–431
Anguera, X., Woofers, C., Hernando, J., 2005 · 2005
Earlier work this paper cites.
Evaluation of BIC-based algorithms for audio segmentation
Cettolo, M., Vescovi, M., Rizzi, R., 2005 · 2005
Earlier work this paper cites.
The Cambridge University March 2005 speaker diarisation system, in: Proceedings of the 6th Conference of the International Speech Communication Association (Interspeech), pp. 2437–2440
Sinha, R., Tranter, S., Gales, M., Woodland, P., 2005 · 2005
Earlier work this paper cites.
Multistage speaker diarization of broadcast news
Barras, C., Zhu, X., Meignier, S., Gauvain, J.L., 2006 · 2006
Earlier work this paper cites.
Step-by-step and integrated approaches in broadcast news speaker diarization
Meignier, S., Moraru, D., Fredouille, C., Bonastre, J.F., Besacier, L., 2006 · 2006
Earlier work this paper cites.
A spectral clustering approach to speaker diarization, in: Proceedings of the 7th Conference of the International Speech Communication Association (Interspeech), pp. 921–924
Ning, H., Liu, M., Tang, H., Huang, T., 2006 · 2006
Earlier work this paper cites.
A tutorial on spectral clustering
Luxburg, U.v., 2007 · 2007
Earlier work this paper cites.
Efficient speaker change detection using adapted Gaussian mixture models
Malegaonkar, A., Ariyaeeinia, A., Sivakumaran, P., 2007 · 2007
Earlier work this paper cites.
Meta-learning with latent space clustering in generative adversarial network for speaker diarization
Pal, M., Kumar, M., Peri, R., Park, T., Kim, S., Lord, C., Bishop, S., Narayanan, S., 2020a · 2007
Earlier work this paper cites.
The ICSI RT07s speaker diarization system, in: Proceedings of the International Evaluation Workshops CLEAR 2007 and RT 2007, pp. 509–519
Wooters, C., Huijbregts, M., 2007 · 2007
Earlier work this paper cites.
Stream-based speaker segmentation using speaker factors and eigenvoices, in: Proceedings of the 33rd IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4133–4136
Castaldo, F., Colibro, D., Dalmasso, E., Laface, P., Vair, C., 2008 · 2008
Earlier work this paper cites.
The IIR-NTU speaker diarization systems, in: Proceedings of the RT’09, NIST Rich Transcription Workshop
Nguyen, T., Sun, H., Zhao, S., Khine, S., Tran, H., Ma, T., Ma, B., Chng, E., Li, H., 2009 · 2009
Earlier work this paper cites.
Bilinear classifiers for visual recognition, in: Advances in Neural Information Processing Systems 22 (NIPS), pp. 1–9
Pirsiavash, H., Ramanan, D., Fowlkes, C., 2009 · 2009
Earlier work this paper cites.
Diarization of telephone conversations using factor analysis
Kenny, P., Reynolds, D., Castaldo, F., 2010 · 2010
Earlier work this paper cites.
Front-end factor analysis for speaker verification
Dehak, N., Kenny, P., Dehak, R., Dumouchel, P., Ouellet, P., 2011 · 2011
Earlier work this paper cites.
A sticky HDP-HMM with application to speaker diarization
Fox, E., Sudderth, E., Jordan, M., Willsky, A., 2011 · 2011
Earlier work this paper cites.
Analysis of i-vector length normalization in speaker recognition systems, in: Proceedings of the 12th Conference of the International Speech Communication Association (Interspeech), pp. 249–252
Garcia-Romero, D., Espy-Wilson, C.Y., 2011 · 2011
Earlier work this paper cites.
Speaker diarization using PLDA-based speaker clustering, in: Proceedings of the 6th IEEE International Conference on Intelligent Data Acquisition and Advanced Computing Systems (IDAACS), pp. 347–350
Prazak, J., Silovsky, J., 2011 · 2011
Earlier work this paper cites.
Exploiting intra-conversation variability for speaker diarization, in: Proceedings of the 12th Conference of the International Speech Communication Association (Interspeech), pp. 945–948
Shum, S., Dehak, N., Chuangsuwanich, E., Reynolds, D., Glass, K., 2011 · 2011
Earlier work this paper cites.
Speaker diarization: A review of recent research
Anguera, X., Bozonnet, S., Evans, N., Fredouille, C., Friedland, G., Vinyals, O., 2012 · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G., Deng, L., Yu, D., Dahl, G., Mohamed, A.r., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T., Kingsbury, B., 2012 · 2012
Earlier work this paper cites.
End-to-end speaker diarization as post-processing
Horiguchi, S., Garcia, P., Fujita, Y., Watanabe, S., Nagamatsu, K., 2020 · 2012
Cited alongside, same era.
Recurrent neural networks for voice activity detection, in: Proceedings of the 38th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7378–7382
Hughes, T., Mierle, K., 2013 · 2013
Cited alongside, same era.
Deep belief networks based voice activity detection
Zhang, X.L., Wu, J., 2013 · 2013
Cited alongside, same era.
Speaker diarization with PLDA i-vector scoring and unsupervised calibration, in: Proceedings of the 5th IEEE Spoken Language Technology Workshop (SLT), pp. 413–417
Sell, G., Garcia-Romero, D., 2014 · 2014
Cited alongside, same era.
Improving DNN speaker independence with i-vector inputs, in: Proceedings of the 39th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 225–229
Angular softmax for short-duration text-independent speaker verification, in: Proceedings of the 19th Conference of the International Speech Communication Association (Interspeech), pp. 3623–3627
Huang, Z., Wang, S., Yu, K., 2018 · 2018
Later among the works it cites.
Characterizing performance of speaker diarization systems on far-field speech using standard methods, in: Proceedings of the 43rd IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5244–5248
Maciejewski, M., Snyder, D., Manohar, V., Dehak, N., Khudanpur, S., 2018 · 2018
Later among the works it cites.
Attentive statistics pooling for deep speaker embedding
Okabe, K., Koshinaka, T., Shinoda, K., 2018 · 2018
Later among the works it cites.
Multistream diarization fusion using the minimum variance Bayesian information criterion, in: Proceedings of the 43rd IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5224–5228
Park, T., Georgy, P., 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Senior, A., Lopez-Moreno, I., 2014 · 2014
Cited alongside, same era.
A study of the cosine distance-based mean shift for telephone speech diarization
Senoussaoui, M., Kenny, P., Stafylakis, T., Dumouchel, P., 2014 · 2014
Cited alongside, same era.
Deep neural networks for small footprint text-dependent speaker verification, in: Proceedings of the 39th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4052–4056
Variani, E., Lei, X., McDermott, E., Moreno, I., Gonzalez-Dominguez, J., 2014 · 2014
Cited alongside, same era.
Speaker change point detection using deep neural nets, in: Proceedings of the 40th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4420–4424
Gupta, V., 2015 · 2015
Cited alongside, same era.
A time delay neural network architecture for efficient modeling of long temporal contexts, in: Proceedings of the 16th Conference of the International Speech Communication Association (Interspeech), pp. 3214–3218
Peddinti, V., Povey, D., Khudanpur, S., 2015 · 2015
Cited alongside, same era.
Diarization resegmentation in the factor analysis subspace, in: Proceedings of the 40th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4794–4798
Sell, G., Garcia-Romero, D., 2015 · 2015
Cited alongside, same era.
A comparison of neural network feature transforms for speaker diarization, in: Proceedings of the 16th Conference of the International Speech Communication Association (Interspeech), pp. 3026–3030
Yella, S., Stolcke, A., 2015 · 2015
Cited alongside, same era.
The HTK Book (for HTK version 3.5)
Young, S., Evermann, G., Gales, M., Hain, T., Kershaw, D., Liu, X., Moore, G., Odell, J., Ollason, D., Povey, D., Ragni, A., Valtchev, V., Woodland, P., Zhang, C., 2015 · 2015
Cited alongside, same era.
Later among the works it cites.
Keyword-based speaker localization: Localizing a target speaker in a multi-speaker environment, in: Proceedings of the 19th Conference of the International Speech Communication Association (Interspeech), pp. 2703–2707
Sivasankaran, S., Vincent, E., Fohr, D., 2018 · 2018
Later among the works it cites.
X-vectors: Robust DNN embeddings for speaker recognition, in: Proceedings of the 43rd IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5329–5333
Snyder, D., Garcia-Romero, D., Sell, G., Povey, D., Khudanpur, S., 2018 · 2018
Later among the works it cites.
Generalized end-to-end loss for speaker verification, in: Proceedings of the 43rd IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4879–4883
Wan, L., Wang, Q., Papir, A., Moreno, I., 2018 · 2018
Later among the works it cites.
High order recurrent neural networks for acoustic modelling, in: Proceedings of the 43rd IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5849–5853
Zhang, C., Woodland, P., 2018 · 2018
Later among the works it cites.
Self-attentive speaker embeddings for text-independent speaker verification, in: Proceedings of the 19th Conference of the International Speech Communication Association (Interspeech), pp. 3573–3577
Zhu, Y., Ko, T., Snyder, D., Mak, B., Povey, D., 2018 · 2018
Later among the works it cites.
BLOCK: Bilinear superdiagonal fusion for visual question answering and visual relationship detection, in: Proceedings of the 33rd AAAI Conference on Artificial Intelligence (AAAI), pp. 8102–8109
Ben-younes, H., Cadene, R., Thome, N., Cord, M., 2019 · 2019
Later among the works it cites.
Speaker characterization using TDNN-LSTM based speaker embedding, in: Proceedings of the 44th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6211–6215
Chen, C., Zhang, S., Yeh, C., Wang, J., Wang, T., Huang, C., 2019 · 2019
Later among the works it cites.
Incremental transfer learning in two-pass information bottleneck based speaker diarization system for meetings, in: Proceedings of the 44th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6291–6295
Dawalatabad, N., Madikeri, S., Sekhar, C., Murthy, H., 2019 · 2019
Later among the works it cites.
ArcFace: Additive angular margin loss for deep face recognition, in: Proceedings of the 32nd IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4685–4694
Deng, J., Guo, J., Xue, N., Zafeiriou, S., 2019 · 2019
Later among the works it cites.
Bayesian HMM based x-vector clustering for speaker diarization
Diez, M., Burget, L., Wang, S., Rohdin, J., Černockỳ, J., 2019 · 2019
Later among the works it cites.
Phoneme dependent speaker embedding and model factorization for multi-speaker speech synthesis and adaptation, in: Proc. ICASSP
Fu, R., Tao, J., Wen, Z., Zheng, Y., 2019 · 2019
Later among the works it cites.
Large margin softmax loss for speaker verification, in: Proceedings of the 20th Conference of the International Speech Communication Association (Interspeech), pp. 2873–2877
Liu, Y., He, L., Liu, J., 2019 · 2019
Later among the works it cites.
Designing an effective metric learning pipeline for speaker diarization, in: Proceedings of the 44th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5806–5810
Narayanaswamy, V., Thiagarajan, J., Song, H., Spanias, A., 2019 · 2019
Later among the works it cites.
Speaker diarisation using 2D self-attentive combination of embeddings, in: Proceedings of the 44th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5801–5805
Sun, G., Zhang, C., Woodland, P., 2019 · 2019
Later among the works it cites.
Multi-PLDA diarization on children’s speech, in: Proceedings of the 20th Conference of the International Speech Communication Association (Interspeech), pp. 376–380
Xie, J., Garcia-Perera, L., Povey, D., Khudanpur, S., 2019 · 2019
Later among the works it cites.
Ensemble additive margin softmax for speaker verification, in: Proceedings of the 44th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6046–6050
Yu, Y., Fan, L., Li, W., 2019 · 2019
Later among the works it cites.
Phase-aware speech enhancement based on deep neural networks
Zheng, N., Zhang, X.L., 2019 · 2019
Later among the works it cites.
Overlap-aware diarization: Resegmentation using neural end-to-end overlapped speech detection, in: Proceedings of the 45th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7114–7118
Bullock, L., Bredin, H., Garcia-Perera, L., 2020 · 2020
Closest in time.
Improved large-margin softmax loss for speaker diarisation, in: Proceedings of the 45th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7104–7108
Fathullah, Y., Zhang, C., Woodland, P., 2020 · 2020
Closest in time.
Cosine-distance virtual adversarial training for semi-supervised speaker-discriminative acoustic embeddings, in: Proceedings of the 21st Conference of the International Speech Communication Association (Interspeech)
Kreyssig, F.L., Woodland, P.C., 2020 · 2020
Closest in time.
H-vectors: Utterance-level speaker embedding using a hierarchical attention model, in: Proceedings of the 45th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7579–7583
Shi, Y., Huang, Q., Hain, T., 2020 · 2020
Closest in time.
Speaker diarization with session-level speaker embedding refinement using graph neural networks, in: Proceedings of the 45th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7109–7113
Wang, J., Xiao, X., Wu, J., Ramamurthy, R., Rudzicz, F., Brudno, M., 2020 · 2020
Closest in time.
Multimodal intelligence: Representation learning, information fusion, and applications
Zhang, C., Yang, Z., He, X., Deng, L., 2020 · 2020
Closest in time.
Discriminative neural clustering for speaker diarisation, in: Proceedings of the 8th IEEE Spoken Language Technology Workshop (SLT)
Li, Q., Kreyssig, F., Zhang, C., Woodland, P., 2021 · 2021
Closest in time.
Acoustic beamforming for speaker diarization of meetings
Anguera, X., Wooters, C., Hernando, J., 2007 · 2022
Closest in time.
Unsupervised methods for speaker diarization: An integrated and iterative approach
Shum, S., Dehak, G., Dehak, R., Glass, J., 2013 · 2028
Closest in time.