Fetching the paper…
Reading the bibliography…
The VoxCeleb Speaker Recognition Challenge 2019 aimed to assess how well current speaker recognition technology is able to identify speakers in unconstrained or `in the wild' data.
“Phoneme recognition using time-delay neural networks,”
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, · 1989
Earlier work this paper cites.
“NIST speaker recognition evaluation chronicles,”
M. P. Alvin and A. Martin, · 2004
Earlier work this paper cites.
“Multicondition training of gaussian plda models in i-vector space for noise and reverberation robust speaker recognition,”
D. Garcia-Romero, X. Zhou, and C. Y. Espy-Wilson, · 2012
Earlier work this paper cites.
“The REPERE corpus: a multimodal corpus for person recognition.,”
A. Giraudel, M. Carré, V. Mapelli, J. Kahn, O. Galibert, and L. Quintard, · 2012
Earlier work this paper cites.
“The speakers in the wild speaker recognition challenge plan,”
M. McLaren, A. Lawson, L. Ferrer, D. Castan, and M. Graciarena, · 2015
Earlier work this paper cites.
“Musan: A music, speech, and noise corpus,”
D. Snyder, G. Chen, and D. Povey, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Earlier work this paper cites.
“Asvspoof: The automatic speaker verification spoofing and countermeasures challenge,”
Z. Wu, J. Yamagishi, T. Kinnunen, C. Hanilçi, M. Sahidullah, A. Sizov, N. Evans, M. Todisco, and H. Delgado, · 2017
Cited alongside, same era.
“VoxCeleb: a large-scale speaker identification dataset,”
A. Nagrani, J. S. Chung, and A. Zisserman, · 2017
Cited alongside, same era.
“Sphereface: Deep hypersphere embedding for face recognition,”
W. Liu, Y. Wen, Z. Yu, M. Li, B. Raj, and L. Song, · 2017
Cited alongside, same era.
“Voxceleb2: Deep speaker recognition,”
J. S. Chung, A. Nagrani, and A. Zisserman, · 2018
Cited alongside, same era.
“Additive margin softmax for face verification,”
F. Wang, J. Cheng, W. Liu, and H. Liu, · 2018
Cited alongside, same era.
“Deepmine speech processing database: Text-dependent and independent speaker verification and speech recognition in persian and english.,”
“Seeing voices and hearing faces: Cross-modal biometric matching,”
A. Nagrani, S. Albanie, and A. Zisserman, · 2018
Later among the works it cites.
“BUT system description to voxceleb speaker recognition challenge 2019,” http://www.robots.ox.ac.uk/~vgg/data/voxceleb/data_workshop/BUT_Zeinali_VoxSRC.pdf , 2019
H. Zeinali, S. Wang, A. Silnova, P. Matˇejka, and O. Plchot, · 2019
Closest in time.
“JHU-HLTCOE system description for voxsrc,” http://www.robots.ox.ac.uk/~vgg/data/voxceleb/data_workshop/JHU-HLTCOE_VoxSRC.pdf , 2019
D. Garcia-Romero, A. McCree, D. Snyder, and G. Sell, · 2019
Closest in time.
“CNN with phonetic attention for text-independent speaker verification,” http://www.robots.ox.ac.uk/~vgg/data/voxceleb/data_workshop/VoxSRC_TZ_microsoft.pdf , 2019
T. Zhou, Y. Zhao, J. Li, and J. Wu, · 2019
Closest in time.
“The DKU-TVM-SYSU system for the voxceleb speaker recognition challenge,” http://www.robots.ox.ac.uk/~vgg/data/voxceleb/data_workshop/DKU-TVM-SYSU.pdf , 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Zeinali, H. Sameti, and T. Stafylakis, · 2018
Cited alongside, same era.
“Audio-visual person recognition in multimedia data from the IARPA Janus program,”
G. Sell, K. Duh, D. Snyder, D. Etter, and D. Garcia-Romero, · 2018
Cited alongside, same era.
X. Qin, W. Cai, Y. Gong, and M. Li, · 2019
Closest in time.
“Arcface: Additive angular margin loss for deep face recognition,”
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, · 2019
Closest in time.
“Voxceleb: Large-scale speaker verification in the wild,”
A. Nagrani, J. S. Chung, W. Xie, and A. Zisserman, · 2020
Closest in time.