Fetching the paper…
Reading the bibliography…
In this paper, we explore the encoding/pooling layer and loss function in the end-to-end speaker and language recognition system.
“Report: A vector quantization approach to speaker recognition,”
F. Soong, A. E. Rosenberg, J. Bling‐Hwang, and L. R. Rabiner, · 1985
Earlier work this paper cites.
“Robust text-independent speaker identification using gaussian mixture speaker models,”
D.A. Reynolds and R.C. Rose, · 1995
Earlier work this paper cites.
“Speaker verification using adapted gaussian mixture models,”
D.A. Reynolds, T.F. Quatieri, and R.B. Dunn, · 2000
Earlier work this paper cites.
“Support vector machines using gmm supervectors for speaker verification,”
W.M. Campbell, D.E. Sturim, and DA Reynolds, · 2006
Earlier work this paper cites.
“Dimensionality reduction by learning an invariant mapping,”
R. Hadsell, S. Chopra, and Y. Lecun, · 2006
Earlier work this paper cites.
“Probabilistic linear discriminant analysis for inferences about identity,”
S.J.D. Prince and J.H. Elder, · 2007
Earlier work this paper cites.
“The 2007 NIST Language Recognition Evaluation Plan,”
NIST, · 2007
Earlier work this paper cites.
“An overview of text-independent speaker recognition: From features to supervectors,”
T. Kinnunen and H. Li, · 2010
Earlier work this paper cites.
“Bayesian speaker verification with heavy tailed priors,”
P. Kenny, · 2010
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, · 2011
Earlier work this paper cites.
“The NIST 2012 Speaker Recognition Evaluation Plan,”
C.S. Greenberg, · 2012
Earlier work this paper cites.
“Spoken language recognition: From fundamentals to practice,”
H. Li, B. Ma, and K. Lee, · 2013
Earlier work this paper cites.
“Neural network bottleneck features for language identification,”
P. Matejka, L. Zhang, T. Ng, H. Mallidi, O. Glembek, J. Ma, and B. Zhang, · 2014
Earlier work this paper cites.
“Speaker verification and spoken language identification using a generalized i-vector framework with phonetic tokenizations and tandem features,”
M. Li and W. Liu, · 2014
Cited alongside, same era.
“A novel scheme for speaker recognition using a phonetically-aware deep neural network,”
Y. Lei, N. Scheffer, L. Ferrer, and M. McLaren, · 2014
Cited alongside, same era.
“Automatic language identification using deep neural networks,”
I. Lopez-Moreno, J. Gonzalez-Dominguez, O. Plchot, D. Martinez, J. Gonzalez-Rodriguez, and P. Moreno, · 2014
Cited alongside, same era.
“Automatic language identification using long short-term memory recurrent neural networks,”
J. Gonzalez-Dominguez, I. Lopez-Moreno, H. Sak, J. Gonzalez-Rodriguez, and P. J Moreno, · 2014
Cited alongside, same era.
“Deep learning face representation by joint identification-verification,”
Y. Chen, Y. Chen, X. Wang, and X. Tang, · 2014
Cited alongside, same era.
“Hierarchical attention networks for document classification,”
Z. Yang, D. Yang, C. Dyer, X. He, A. Smola, and E. Hovy, · 2016
Later among the works it cites.
“A discriminative feature learning approach for deep face recognition,”
Y. Wen, K. Zhang, Z. Li, and Y. Qiao, · 2016
Later among the works it cites.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Later among the works it cites.
“Deep neural network-based speaker embeddings for end-to-end speaker verification,”
D. Snyder, P. Ghahremani, D. Povey, D. Garcia-Romero, Y. Carmiel, and S. Khudanpur, · 2017
Later among the works it cites.
“Deep speaker: an end-to-end neural speaker embedding system,” 2017
L. Chao, M. Xiaokong, J. Bing, L. Xiangang, Z. Xuewei, L. Xiao, C. Ying, K. Ajay, and Z. Zhenyao, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Deep bottleneck network based i-vector representation for language identification,”
Y. Song, X. Hong, B. Jiang, R. Cui, I. Mcloughlin, and L. Dai, · 2015
Cited alongside, same era.
“A unified deep neural network for speaker and language recognition,”
Fred Richardson, Douglas Reynolds, and Najim Dehak, · 2015
Cited alongside, same era.
“Deep neural network approaches to speaker and language recognition,”
F. Richardson, D. Reynolds, and N. Dehak, · 2015
Cited alongside, same era.
“Facenet: A unified embedding for face recognition and clustering,”
F. Schroff, D. Kalenichenko, and J. Philbin, · 2015
Cited alongside, same era.
“Language recognition via i-vectors and dimensionality reduction,”
N. Dehak, P.A. Torres-Carrasquillo, D. Reynolds, and R. Dehak, · 2016
Cited alongside, same era.
“Generalized i-vector representation with phonetic tokenizations and tandem features for both text independent and text dependent speaker verification,”
M. Li, L. Liu, W. Cai, and W. Liu, · 2016
Cited alongside, same era.
“Time delay deep neural network-based universal background models for speaker recognition,”
D. Snyder, D. Garcia-Romero, and D. Povey, · 2016
Cited alongside, same era.
“Deep speaker embeddings for short-duration speaker verification,”
G. Bhattacharya, J. Alam, and P. Kenny, · 2017
Later among the works it cites.
“Attention-based models for text-dependent speaker verification,”
FA Chowdhury, Q. Wang, I. L. Moreno, and L. Wan, · 2017
Later among the works it cites.
“Sphereface: Deep hypersphere embedding for face recognition,”
W. Liu, Y. Wen, Z. Yu, M. Li, B. Raj, and L. Song, · 2017
Later among the works it cites.
“Voxceleb: a large-scale speaker identification dataset,”
A. Nagrani, J. S. Chung, and A. Zisserman, · 2017
Later among the works it cites.
“Lid-senones and their statistics for language identification,”
M. Jin, Y. Song, I. McLoughlin, and L. Dai, · 2018
Closest in time.
“Insights into end-to-end learning scheme for language identification,”
W. Cai, Z. Cai, W. Liu, X. Wang, and M. Li, · 2018
Closest in time.
“A novel learnable dictionary encoding layer for end-to-end language identification,”
W. Cai, Z. Cai, X. Zhang, X. Wang, and M. Li, · 2018
Closest in time.