Fetching the paper…
Reading the bibliography…
The performance of most speaker diarization systems with x-vector embeddings is both vulnerable to noisy environments and lacks domain robustness.
A. F. Martin and M. A. Przybocki, “Speaker recognition in a multi-speaker environment,” in Proc. Speech Commun. and Tech. , 2001
2001
Earlier work this paper cites.
A. Janin et al. , “The ICSI meeting corpus,” in Proc. ICASSP , vol. 1, 2003, pp. I–I
2003
Earlier work this paper cites.
S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” in Proc. CVPR , vol. 1. IEEE, 2005, pp. 539–546
2005
Earlier work this paper cites.
H. Ning et al. , “A spectral clustering approach to speaker diarization,” in Proc. ICSLP , 2006
2006
Earlier work this paper cites.
J. G. Fiscus et al. , “The Rich Transcription 2006 spring meeting recognition evaluation,” in International Workshop on Machine Learning for Multimodal Interaction . Springer, 2006, pp. 309–322
2006
Earlier work this paper cites.
K. J. Han, S. Kim, and S. S. Narayanan, “Strategies to improve the robustness of agglomerative hierarchical clustering under data source variation for speaker diarization,” IEEE Trans. Audio, Speech, Lang. Process. , vol. 16, no. 8, pp. 1590–1601, 2008
2008
Earlier work this paper cites.
D. Vijayasenan, F. Valente, and H. Bourlard, “An information theoretic approach to speaker diarization of meeting data,” IEEE Trans. Audio, Speech, and Lang. Process. , vol. 17, no. 7, pp. 1382–1393, 2009
2009
Earlier work this paper cites.
N. Dehak et al. , “Front-end factor analysis for speaker verification,” IEEE Trans. Audio, Speech, Lang. Process. , vol. 19, no. 4, pp. 788–798, 2010
2010
Earlier work this paper cites.
K. Chen and A. Salman, “Learning speaker-specific characteristics with a deep neural architecture,” IEEE Trans. on Neural Networks , vol. 22, no. 11, pp. 1744–1756, 2011
2011
Earlier work this paper cites.
X. Anguera et al. , “Speaker diarization: A review of recent research,” IEEE Trans. Audio, Speech, Lang. Process. , vol. 20, no. 2, pp. 356–370, 2012
2012
Earlier work this paper cites.
S. H. Shum, N. Dehak, R. Dehak, and J. R. Glass, “Unsupervised methods for speaker diarization: An integrated and iterative approach,” IEEE Trans. Audio, Speech, Lang. Process. , vol. 21, no. 10, pp. 2015–2028, 2013
2013
Earlier work this paper cites.
M. Senoussaoui, P. Kenny, T. Stafylakis, and P. Dumouchel, “A study of the cosine distance-based mean shift for telephone speech diarization,” IEEE/ACM Trans. Audio, Speech, and Lang. Process. , vol. 22, no. 1, pp. 217–227, 2014
2014
Earlier work this paper cites.
I. Goodfellow et al. , “Generative adversarial nets,” in Proc. NIPS , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
G. Koch, R. Zemel, and R. Salakhutdinov, “Siamese neural networks for one-shot image recognition,” in ICML deep learning workshop , vol. 2. Lille, 2015
2015
Earlier work this paper cites.
F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in Proc. CVPR , 2015, pp. 815–823
2015
Earlier work this paper cites.
J. Xie, R. Girshick, and A. Farhadi, “Unsupervised deep embedding for clustering analysis,” in Proc. ICML , 2016, pp. 478–487
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
X. Chen et al. , “InfoGAN: Interpretable representation learning by information maximizing generative adversarial nets,” in Proc. NIPS , 2016, pp. 2172–2180
2016
Earlier work this paper cites.
J. Yang, D. Parikh, and D. Batra, “Joint unsupervised learning of deep representations and image clusters,” in Proc. CVPR , 2016, pp. 5147–5156
2016
Earlier work this paper cites.
S. Ravi and H. Larochelle, “Optimization as a model for few-shot learning,” 2016
2016
Earlier work this paper cites.
A. Santoro et al. , “Meta-learning with memory-augmented neural networks,” in Proc. ICML , 2016, pp. 1842–1850
2016
Earlier work this paper cites.
O. Vinyals et al. , “Matching networks for one shot learning,” in Proc. NIPS , 2016, pp. 3630–3638
2016
Cited alongside, same era.
R. Grzadzinski et al. , “Measuring changes in social communication behaviors: preliminary development of the brief observation of social communication change (BOSCC),” Journal of autism and developmental disorders , vol. 46, no. 7, pp. 2464–2479, 2016
2016
Cited alongside, same era.
D. Garcia-Romero et al. , “Speaker diarization using deep neural network embeddings,” in Proc. ICASSP , 2017, pp. 4930–4934
2017
Cited alongside, same era.
Z. Zajíc, M. Hrúz, and L. Müller, “Speaker diarization using convolutional neural network for statistics accumulation refinement,” in Proc. Interspeech , 2017, pp. 3562–3566
2017
Cited alongside, same era.
J. Snell, K. Swersky, and R. Zemel, “Prototypical networks for few-shot learning,” in Proc. NIPS , 2017, pp. 4077–4087
A. Zhang et al. , “Fully supervised speaker diarization,” in Proc. ICASSP , 2019, pp. 6301–6305
2019
Later among the works it cites.
T. J. Park, K. J. Han, M. Kumar, and S. Narayanan, “Auto-tuning spectral clustering for speaker diarization using normalized maximum eigengap,” IEEE Sign. Process. Lett. , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
G. Sun, C. Zhang, and P. C. Woodland, “Speaker diarisation using 2D self-attentive combination of embeddings,” in Proc. ICASSP , 2019, pp. 5801–5805
2019
Later among the works it cites.
T. J. Park et al. , “The Second DIHARD challenge: System Description for USC-SAIL Team,” in Proc. Interspeech , 2019, pp. 998–1002
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
B. Yang et al. , “Towards k-means-friendly spaces: Simultaneous deep learning and clustering,” in Proc. ICML . JMLR. org, 2017, pp. 3861–3870
2017
Cited alongside, same era.
H. Bredin, “Tristounet: triplet loss for speaker turn embedding,” in Proc. ICASSP , 2017, pp. 5430–5434
2017
Cited alongside, same era.
M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein GAN,” arXiv preprint arXiv:1701.07875 , 2017
2017
Cited alongside, same era.
I. Gulrajani et al. , “Improved training of Wasserstein GANs,” in Proc. NIPS , 2017, pp. 5767–5777
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Q. Wang et al. , “Speaker diarization with LSTM,” in Proc. ICASSP , 2018, pp. 5239–5243
2018
Cited alongside, same era.
G. Sell et al. , “Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge,” in Proc. Interspeech , 2018, pp. 2808–2812
2018
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
Y. Fujita et al. , “End-to-End Neural Speaker Diarization with Permutation-Free Objectives,” in Proc. Interspeech , 2019, pp. 4300–4304
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Mukherjee, H. Asnani, E. Lin, and S. Kannan, “ClusterGAN: Latent space clustering in generative adversarial networks,” in Proc. AAAI , vol. 33, 2019, pp. 4610–4617
2019
Later among the works it cites.
J. Wang et al. , “Centroid-based deep metric learning for speaker recognition,” in Proc. ICASSP . IEEE, 2019, pp. 3652–3656
2019
Later among the works it cites.
K. Ghasedi et al. , “Balanced self-paced learning for generative adversarial clustering network,” in Proc. CVPR , 2019, pp. 4391–4400
2019
Later among the works it cites.
2019
Later among the works it cites.
M. Diez et al. , “Analysis of speaker diarization based on Bayesian HMM with eigenvoice priors,” IEEE/ACM Trans. Audio, Speech, and Lang. Process. , vol. 28, pp. 355–368, 2019
2019
Later among the works it cites.
N. Ryant et al. , “Second DIHARD challenge evaluation plan,” Linguistic Data Consortium, Tech. Rep , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
E. Fini and A. Brutti, “Supervised online diarization with sample mean loss for multi-domain data,” in Proc. ICASSP , 2020, pp. 7134–7138
2020
Closest in time.
M. Pal et al. , “Speaker diarization using latent space clustering in generative adversarial network,” in Proc. ICASSP , 2020, pp. 6504–6508
2020
Closest in time.
2020
Closest in time.
N. R. Koluguri et al. , “Meta-learning for robust child-adult classification from speech,” in Proc. ICASSP , 2020, pp. 8094–8098
2020
Closest in time.