Fetching the paper…
Reading the bibliography…
In this paper, we propose a new pooling method called spatial pyramid encoding (SPE) to generate speaker embeddings for text-independent speaker verification.
J. Sivic and A. Zisserman, “Video google: A text retrieval approach to object matching in videos,” in
2003
Earlier work this paper cites.
S. Ioffe, “Probabilistic linear discriminant analysis,” in
2006
Earlier work this paper cites.
S. Lazebnik, C. Schmid, and J. Ponce, “Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories,” in
2006
Earlier work this paper cites.
P. Kenny, “Bayesian speaker verification with heavy tailed priors,” in
2010
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
D. Garcia-Romero and C. Espy-Wilson, “Analysis of ivector length normalization in speaker recognition systems,” in
2011
Earlier work this paper cites.
D. Garcia-Romero and C. Y. Espy-Wilson, “Analysis of i-vector length normalization in speaker recognition systems,” in
2011
Earlier work this paper cites.
Y. Kamishima, N. Inoue, and K. Shinoda, “Event detection in consumer videos using gmm supervectors and svms,”
2013
Earlier work this paper cites.
P. Kenny, V. Gupta, T. Stafylakis, P. Ouellet, and J. Alam, “Deep neural networks for extracting baum-welch statistics for speaker recognition,” in
2014
Earlier work this paper cites.
Y. Lei, N. Scheffer, L. Ferrer, and M. McLaren, “A novel scheme for speaker recognition using a phonetically-aware deep neural network,” in
2014
Earlier work this paper cites.
E. Variani, X. Lei, E. McDermott, I. Moreno, and J. GonzalezDominguez, “Deep neural networks for small footprint text-dependent speaker verification,” in
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” in
2014
Earlier work this paper cites.
Y. Chen, I. Lopez-Moreno, T. N. Sainath, M. Visontai, R. Alvarez, and C. Parada, “Locally-connected and convolutional neural networks for small footprint speaker recognition,” in
2015
Earlier work this paper cites.
J. Hansen and T. Hasan, “Speaker recognition by machines and humans: A tutorial review,”
2015
Cited alongside, same era.
D. Snyder, P. Ghahremani, D. Povey, D. Garcia-Romero, Y. Carmiel, and S. Khudanpur, “Deep neural network based speaker embeddings for end-to-end speaker verification,” in
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Cited alongside, same era.
W. Liu, Y. Wen, Z. Yu, and M. Yang, “Large-margin softmax loss for convolutional neural networks,” in
2016
Cited alongside, same era.
Y. Wen, K. Zhang, Z. Li, and Y. Qiao, “A discriminative feature learning approach for deep face recognition,” in
2016
Cited alongside, same era.
C. Zhang, K. Koishida, and J. Hansen, “Text-independent speaker verification based on triplet convolutional neural network embeddings,”
2018
Later among the works it cites.
N. Li, D. Tuo, D. Su, Z. Li, and D. Yu, “Deep discriminative embeddings for duration robust speaker verification,” in
2018
Later among the works it cites.
Z. Huang, S. Wang, and K. Yu, “Angular softmax for short duration text-independent speaker verification,” in
2018
Later among the works it cites.
K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statistics pooling for deep speaker embedding,” in
2018
Later among the works it cites.
W. Cai, J. Chen, and M. Li, “Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,” in
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: A large-scale speaker identification dataset,” in
2017
Cited alongside, same era.
D. Snyder, D. Garcia-Romero, D. Povey, and S. Khudanpur, “Deep neural network embeddings for text-independent speaker verification,” in
2017
Cited alongside, same era.
W. Liu, Y. Wen, Z. Yu, M. Li, B. Raj, and L. Song, “Sphereface: Deep hypersphere embedding for face recognition,” in
2017
Cited alongside, same era.
H. Zhang, J. Xue, and K. Dana, “Deep ten: Texture encoding network,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in pytorch,” in
2017
Cited alongside, same era.
W. Cai, J. Chen, and M. Li, “Analysis of length normalization in end-to-end speaker verification system,” in
2018
Later among the works it cites.
Y. Zheng, D. K. Pal, and M. Savvides, “Ring loss: Convex feature normalization for face recognition,” in
2018
Later among the works it cites.
M. Ravanelli and Y. Bengio, “Speaker recognition from raw waveform with sincnet,” in
2018
Later among the works it cites.
N. Le and J. Odobez, “Robust and discriminative speaker embedding via intra-class distance variance regularization,” in
2018
Later among the works it cites.
S. Shon, H. Tang, and J. Glass, “Frame-level speaker embeddings for text-independent speaker recognition and analysis of end-to-end model,” in
2018
Later among the works it cites.
S. Yadav and A. Rai, “Learning discriminative features for speaker identification and verification,” in
2018
Later among the works it cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in
2018
Later among the works it cites.