Fetching the paper…
Reading the bibliography…
Recently, speaker embeddings extracted from a speaker discriminative deep neural network (DNN) yield better performance than the conventional methods such as i-vector.
S. Ioffe, “Probabilistic linear discriminant analysis,” in
2006
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz
2011
Earlier work this paper cites.
D. Snyder, P. Ghahremani, D. Povey, D. Garcia-Romero, Y. Carmiel, and S. Khudanpur, “Deep neural network-based speaker embeddings for end-to-end speaker verification,” in
2016
Earlier work this paper cites.
G. Heigold, I. Moreno, S. Bengio, and N. Shazeer, “End-to-end text-dependent speaker verification,” in
2016
Earlier work this paper cites.
S.-X. Zhang, Z. Chen, Y. Zhao, J. Li, and Y. Gong, “End-to-end attention based text-dependent speaker verification,” in
2016
Earlier work this paper cites.
Y. Wen, K. Zhang, Z. Li, and Y. Qiao, “A discriminative feature learning approach for deep face recognition,” in
2016
Earlier work this paper cites.
M. McLaren, L. Ferrer, D. Castan, and A. Lawson, “The speakers in the wild (SITW) speaker recognition database,” in
2016
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, D. Povey, and S. Khudanpur, “Deep neural network embeddings for text-independent speaker verification,” in
2017
Earlier work this paper cites.
C. Zhang and K. Koishida, “End-to-end text-independent speaker verification with triplet loss on short utterances.” in
2017
Cited alongside, same era.
W. Liu, Y. Wen, Z. Yu, M. Li, B. Raj, and L. Song, “Sphereface: Deep hypersphere embedding for face recognition,” in
2017
Cited alongside, same era.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in PyTorch,” in
2017
Cited alongside, same era.
F. Wang, X. Xiang, J. Cheng, and A. L. Yuille, “Normface: L2 hypersphere embedding for face verification,” in
2017
Cited alongside, same era.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: A large-scale speaker identification dataset,” in
2017
Cited alongside, same era.
F. Wang, J. Cheng, W. Liu, and H. Liu, “Additive margin softmax for face verification,”
2018
Later among the works it cites.
W. Cai, J. Chen, and M. Li, “Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,” in
2018
Later among the works it cites.
Z. Huang, S. Wang, and K. Yu, “Angular softmax for short-duration text-independent speaker verification,”
2018
Later among the works it cites.
K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statistics pooling for deep speaker embedding,” in
2018
Later among the works it cites.
M. Hajibabaei and D. Dai, “Unified hypersphere embedding for speaker recognition,”
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in
2018
Cited alongside, same era.
Z. Huang, S. Wang, and Y. Qian, “Joint i-vector with end-to-end system for short duration text-independent speaker verification,” in
2018
Cited alongside, same era.
H. Wang, Y. Wang, Z. Zhou, X. Ji, D. Gong, J. Zhou, Z. Li, and W. Liu, “Cosface: Large margin cosine loss for deep face recognition,” in
2018
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in
2018
Later among the works it cites.
A. Sergeev and M. Del Balso, “Horovod: fast and easy distributed deep learning in tensorflow,”
2018
Later among the works it cites.
J. Deng, J. Guo, X. Niannan, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in
2019
Closest in time.
W. Xie, A. Nagrani, J. S. Chung, and A. Zisserman, “Utterance-level aggregation for speaker recognition in the wild,” in
2019
Closest in time.