Fetching the paper…
Reading the bibliography…
The classical i-vectors and the latest end-to-end deep speaker embeddings are the two representative categories of utterance-level representations in automatic speaker verification systems.
D. Reynolds and R. Rose, “Robust text-independent speaker identification using gaussian mixture speaker models,”
1995
Earlier work this paper cites.
D. Reynolds, T. Quatieri, and R. Dunn, “Speaker verification using adapted gaussian mixture models,” in
2000
Earlier work this paper cites.
W. Campbell, D. Sturim, and D. Reynolds, “Support vector machines using gmm supervectors for speaker verification,”
2006
Earlier work this paper cites.
S. Prince and J. Elder, “Probabilistic linear discriminant analysis for inferences about identity,” in
2007
Earlier work this paper cites.
T. Kinnunen and H. Li, “An overview of text-independent speaker recognition: From features to supervectors,”
2010
Earlier work this paper cites.
P. Kenny, “Bayesian speaker verification with heavy tailed priors,” in
2010
Earlier work this paper cites.
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
D. Garcia-Romero and C. Y. Espy-Wilson, “Analysis of i-vector length normalization in speaker recognition systems.” in
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The kaldi speech recognition toolkit,” in
2011
Earlier work this paper cites.
P. Bousquet, A. Larcher, D. Matrouf, J. Bonastre, and O. Plchot, “Variance-spectra based normalization for i-vector standard and probabilistic linear discriminant analysis,” in
2012
Cited alongside, same era.
J. Gonzalez-Dominguez, I. Lopez-Moreno, H. Sak, J. Gonzalez-Rodriguez, and P. J. Moreno, “Automatic language identification using long short-term memory recurrent neural networks,” in
2014
Cited alongside, same era.
J. H. Hansen and T. Hasan, “Speaker recognition by machines and humans: A tutorial review,”
2015
Cited alongside, same era.
D. Garcia-Romero and A. McCree, “Insights into deep neural networks for speaker recognition,” in
2015
Cited alongside, same era.
F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in
2015
Cited alongside, same era.
C. Zhang and K. Koishida, “End-to-end text-independent speaker verification with triplet loss on short utterances,” in
2017
Later among the works it cites.
W. Liu, Y. Wen, Z. Yu, M. Li, B. Raj, and L. Song, “Sphereface: Deep hypersphere embedding for face recognition,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
F. Wang, X. Xiang, J. Cheng, and A. L. Yuille, “Normface: L2 hypersphere embedding for face verification,” in
2017
Later among the works it cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: a large-scale speaker identification dataset,” in
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Cited alongside, same era.
L. Chao, M. Xiaokong, J. Bing, L. Xiangang, Z. Xuewei, L. Xiao, C. Ying, K. Ajay, and Z. Zhenyao, “Deep speaker: an end-to-end neural speaker embedding system,” 2017
2017
Cited alongside, same era.
D. Snyder, P. Ghahremani, D. Povey, D. Garcia-Romero, Y. Carmiel, and S. Khudanpur, “Deep neural network-based speaker embeddings for end-to-end speaker verification,” in
2017
Cited alongside, same era.
H. Bredin, “Tristounet: triplet loss for speaker turn embedding,” in
2017
Cited alongside, same era.
W. Cai, J. Chen, and M. Li, “Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,” in
2018
Closest in time.
W. Cai, Z. Cai, W. Liu, X. Wang, and M. Li, “Insights into end-to-end learning scheme for language identification,” in
2018
Closest in time.
W. Cai, Z. Cai, X. Zhang, and M. Li, “A novel learnable dictionary encoding layer for end-to-end language identification,” in
2018
Closest in time.
D. Snyder, G. Garcia-Romero, D. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in
2018
Closest in time.