Fetching the paper…
Reading the bibliography…
The objective of this paper is speaker recognition "in the wild"-where utterances may be of variable length and also contain irrelevant signals.
“Residual enhanced visual vectors for on-device image matching,”
D. Chen, S. Tsai, V. Chandrasekhar, G. Takacs, H. Chen, R. Vedantham, R. Grzeszczuk, and B. Girod, · 2011
Earlier work this paper cites.
“Automatic language identification using deep neural networks,”
I. Lopez-Moreno, J. Gonzalez-Dominguez, O. Plchot, D. Martinez, J. Gonzalez-Rodriguez, and P. Moreno, · 2014
Earlier work this paper cites.
“Matconvnet: Convolutional neural networks for matlab,”
A. Vedaldi and K. Lenc, · 2015
Earlier work this paper cites.
“The speakers in the wild (SITW) speaker recognition database,”
M. McLaren, L. Ferrer, D. Castan, and A. Lawson, · 2016
Earlier work this paper cites.
“Tensorflow: Large-scale machine learning on heterogeneous distributed systems,”
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G.S. Corrado, A. Davis, J. Dean, M. Devin, et al., · 2016
Earlier work this paper cites.
“Speaker identification and clustering using convolutional neural networks,”
Y. Lukic, C. Vogt, O. Dürr, and T. Stadelmann, · 2016
Earlier work this paper cites.
“NetVLAD: CNN architecture for weakly supervised place recognition,”
R. Arandjelović, P. Gronat, A. Torii, T. Pajdla, and J. Sivic, · 2016
Earlier work this paper cites.
“VoxCeleb: a large-scale speaker identification dataset,”
A. Nagrani, J. S. Chung, and A. Zisserman, · 2017
Earlier work this paper cites.
“Automatic differentiation in pytorch,”
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, · 2017
Earlier work this paper cites.
“Deep speaker: an end-to-end neural speaker embedding system,”
C. Li, X. Ma, B. Jiang, X. Li, X. Zhang, X. Liu, Y. Cao, A. Kannan, and Z. Zhu, · 2017
Cited alongside, same era.
“Deep neural network embeddings for text-independent speaker verification,”
D. Snyder, D. Garcia-Romero, D. Povey, and S. Khudanpur, · 2017
Cited alongside, same era.
“Deep speaker embeddings for short-duration speaker verification,”
G. Bhattacharya, J. Alam, and P. Kenny, · 2017
Cited alongside, same era.
“Attention-based models for text-dependent speaker verification,”
FA Chowdhury, Quan Wang, Ignacio Lopez Moreno, and Li Wan, · 2017
Cited alongside, same era.
“VoxCeleb2: Deep speaker recognition,”
J. S. Chung, A. Nagrani, and A. Zisserman, · 2018
Cited alongside, same era.
“A novel learnable dictionary encoding layer for end-to-end language identification,”
W. Cai, Z. Cai, X. Zhang, X. Wang, and M. Li, · 2018
Later among the works it cites.
W. Cai, J. Chen, and M. Li, · 2018
Later among the works it cites.
“GhostVLAD for set-based face recognition,”
Y. Zhong, R. Arandjelović, and A. Zisserman, · 2018
Later among the works it cites.
“Analysis of length normalization in end-to-end speaker verification system,”
W. Cai, J. Chen, and M. Li, · 2018
Later among the works it cites.
“Unified hypersphere embedding for speaker recognition,”
M. Hajibabaei and D. Dai, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Generalized end-to-end loss for speaker verification,”
L. Wan, Q. Wang, A. Papir, and I.L. Moreno, · 2018
Cited alongside, same era.
S. Shon, H. Tang, and J. Glass, · 2018
Cited alongside, same era.
“Attentive statistics pooling for deep speaker embedding,”
K. Okabe, T. Koshinaka, and K. Shinoda, · 2018
Cited alongside, same era.
Later among the works it cites.
“X-vectors: Robust dnn embeddings for speaker recognition,”
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, · 2018
Later among the works it cites.
“End-to-end language identification using netfv and netvlad,”
Jinkun Chen, Weicheng Cai, Danwei Cai, Zexin Cai, Haibin Zhong, and Ming Li, · 2018
Later among the works it cites.
“Additive margin softmax for face verification,”
F. Wang, W. Liu, H. Liu, and J. Cheng, · 2018
Later among the works it cites.