Fetching the paper…
Reading the bibliography…
Learning robust speaker embeddings is a crucial step in speaker diarization.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, G. Lathoud, M. Lincoln, A. Lisowska, I. McCowan, W. Post, D. Reidsma, and P. Wellner, “The AMI meeting corpus: A pre-announcement,” in Proc. the Second International Conference on Machine Learning for Multimodal Interaction , ser. MLMI’05, 28–39, 2006
2006
Earlier work this paper cites.
S. Ioffe, “Probabilistic linear discriminant analysis,” in ECCV , 2006, pp. 531–542
2006
Earlier work this paper cites.
2007
Earlier work this paper cites.
U. von Luxburg, “A tutorial on spectral clustering,” Statistics and Computing , vol. 17, no. 4, 2007
2007
Earlier work this paper cites.
X. Anguera, C. Wooters, and J. Hernando, “Acoustic beamforming for speaker diarization of meetings,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 7, pp. 2011–2022, 2007
2007
Earlier work this paper cites.
X. Anguera, S. Bozonnet, N. Evans, C. Fredouille, G. Friedland, and O. Vinyals, “Speaker diarization: A review of recent research,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 2, pp. 356–370, 2012
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” Proc. ICLR , 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 770–778
2016
Cited alongside, same era.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: A large-scale speaker identification dataset,” in Proc. Interspeech , 2017, pp. 2616–2620
2017
Cited alongside, same era.
L. N. Smith, “Cyclical learning rates for training neural networks,” in IEEE WACV , 2017, pp. 464–472
2017
Cited alongside, same era.
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, and S. Khudanpur, “Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,” in Proc. Interspeech , 2808–2812, 2018
2018
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” in Proc. Interspeech , 2019, pp. 2613–2617
2019
Later among the works it cites.
2019
Later among the works it cites.
M. Diez, L. Burget, S. Wang, J. Rohdin, and J. Černocký, “Bayesian HMM based x-vector clustering for speaker diarization,” in Proc. Interspeech , 2019, pp. 346–350
2019
Later among the works it cites.
T. J. Park, K. J. Han, M. Kumar, and S. Narayanan, “Auto-tuning spectral clustering for speaker diarization using normalized maximum eigengap,” IEEE Signal Processing Letters , vol. 27, pp. 381–385, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Wan, Q. Wang, A. Papir, and I. Lopez-Moreno, “Generalized end-to-end loss for speaker verification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 4879–4883, 2018
2018
Cited alongside, same era.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 5329–5333, 2018
2018
Cited alongside, same era.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-Excitation networks,” in IEEE/CVF CVPR , 2018, pp. 7132–7141
2018
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in Proc. Interspeech , 2018, pp. 1086–1090
2018
Cited alongside, same era.
G. Sun, C. Zhang, and P. C. Woodland, “Speaker diarisation using 2d self-attentive combination of embeddings,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 5801–5805, 2019
2019
Cited alongside, same era.
S. Gao, M.-M. Cheng, K. Zhao, X. Zhang, M.-H. Yang, and P. H. S. Torr, “Res2Net: A new multi-scale backbone architecture,” IEEE TPAMI , pp. 652–662, 2019
2019
Cited alongside, same era.
Z. Gao, Y. Song, I. McLoughlin, P. Li, Y. Jiang, and L.-R. Dai, “Improving Aggregation and Loss Function for Better Embedding Learning in End-to-End Speaker Verification System,” in Proc. Interspeech , 2019, pp. 361–365
2019
Cited alongside, same era.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “ArcFace: Additive angular margin loss for deep face recognition,” in IEEE/CVF CVPR , 2019, pp. 4685–4694
2019
Cited alongside, same era.
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,” in Proc. Interspeech , 2020, pp. 3830–3834
2020
Later among the works it cites.
H. Zeinali, K. A. Lee, J. Alam, and L. Burget, “SdSV challenge 2020: Large-scale evaluation of short-duration speaker verification,” in Proc. Interspeech , 2020, pp. 731–735
2020
Later among the works it cites.
J. Thienpondt, B. Desplanques, and K. Demuynck, “Cross-lingual speaker verification with domain-balanced hard prototype mining and language-dependent score normalization,” in Proc. Interspeech , 2020, pp. 756–760
2020
Later among the works it cites.
D. Garcia-Romero, G. Sell, and A. McCree, “Magneto: X-vector magnitude estimation network plus offset for improved speaker recognition,” in Proc. Odyssey , 2020, pp. 1–8
2020
Later among the works it cites.
2021
Closest in time.
N. Dawalatabad, S. Madikeri, C. C. Sekhar, and H. A. Murthy, “Novel architectures for unsupervised information bottleneck based speaker diarization of meetings,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 14–27, 2021
2021
Closest in time.
M. Ravanelli, T. Parcollet, A. Rouhe, P. Plantinga, E. Rastorgueva, L. Lugosch, N. Dawalatabad, C. Ju-Chieh, A. Heba, F. Grondin, W. Aris, C.-F. Liao, S. Cornell, S.-L. Yeh, H. Na, Y. Gao, S.-W. Fu, C. Subakan, R. De Mori, and Y. Bengio, “Speechbrain,” https://github.com/speechbrain/speechbrain
2021
Closest in time.
Xavier Anguera, “Diarization Error Rate,” http://www.xavieranguera.com/phdthesis/node108.html
2021
Closest in time.