Fetching the paper…
Reading the bibliography…
In this paper we propose and analyse a large margin fine-tuning strategy and a quality-aware score calibration in text-independent speaker verification.
“Phoneme recognition using time-delay neural networks,”
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, · 1989
Earlier work this paper cites.
“Probabilistic linear discriminant analysis,”
S. Ioffe, · 2006
Earlier work this paper cites.
“SPRAAK: an open source ”SPeech recognition and automatic annotation kit”,”
K. Demuynck, J. Roelens, D. V. Compernolle, and P. Wambacq, · 2008
Earlier work this paper cites.
“Comparison of speaker recognition approaches for real applications.,”
S. Cumani, P. Batzu, D. Colibro, C. Vair, P. Laface, and V. Vasilakakis, · 2011
Earlier work this paper cites.
“The BOSARIS toolkit: Theory, algorithms and code for surviving the new DCF,” 2013
N. Brümmer and E. de Villiers, · 2013
Earlier work this paper cites.
“Quality measure functions for calibration of speaker recognition systems in various duration conditions,”
M. I. Mandasari, R. Saeidi, M. McLaren, and D. A. van Leeuwen, · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“MUSAN: A music, speech, and noise corpus,” 2015
D. Snyder, G. Chen, and D. Povey, · 2015
Earlier work this paper cites.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Earlier work this paper cites.
“Robustness of quality-based score calibration of speaker recognition systems with respect to low-snr and short-duration conditions,”
A. Nautsch, R. Saeidi, C. Rathgeb, and C. Busch, · 2016
Earlier work this paper cites.
“Analysis and description of ABC submission to NIST SRE 2016,”
O. Plchot, P. Matějka, A. Silnova, O. Novotný, M. D. Sánchez, J. Rohdin, O. Glembek, N. Brümmer, A. Swart, J. Jorrín-Prieto, P. García, L. Buera, P. Kenny, J. Alam, and G. Bhattacharya, · 2017
Cited alongside, same era.
“Cyclical learning rates for training neural networks,”
L. N. Smith, · 2017
Cited alongside, same era.
“A study on data augmentation of reverberant speech for robust speech recognition,”
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, · 2017
Cited alongside, same era.
“VoxCeleb: A large-scale speaker identification dataset,”
A. Nagrani, J. S. Chung, and A. Zisserman, · 2017
Cited alongside, same era.
“X-vectors: Robust DNN embeddings for speaker recognition,”
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, · 2018
Cited alongside, same era.
“Self-attentive speaker embeddings for text-independent speaker verification,”
“Speaker recognition for multi-speaker conversations using x-vectors,”
D. Snyder, D. Garcia-Romero, G. Sell, A. McCree, D. Povey, and S. Khudanpur, · 2019
Later among the works it cites.
“ArcFace: Additive angular margin loss for deep face recognition,”
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, · 2019
Later among the works it cites.
“Res2Net: A new multi-scale backbone architecture,”
S. Gao, M.-M. Cheng, K. Zhao, X. Zhang, M.-H. Yang, and P. H. S. Torr, · 2019
Later among the works it cites.
“x-Vector DNN Refinement with Full-Length Recordings for Speaker Recognition,”
D. Garcia-Romero, D. Snyder, G. Sell, A. McCree, D. Povey, and S. Khudanpur, · 2019
Later among the works it cites.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Zhu, T. Ko, D. Snyder, B. K.-W. Mak, and D. Povey, · 2018
Cited alongside, same era.
“Attentive statistics pooling for deep speaker embedding,”
K. Okabe, T. Koshinaka, and K. Shinoda, · 2018
Cited alongside, same era.
“Squeeze-and-Excitation networks,”
J. Hu, L. Shen, and G. Sun, · 2018
Cited alongside, same era.
“Ring loss: Convex feature normalization for face recognition,”
Y. Zheng, D. K. Pal, and M. Savvides, · 2018
Cited alongside, same era.
“VoxCeleb2: Deep speaker recognition,”
J. S. Chung, A. Nagrani, and A. Zisserman, · 2018
Cited alongside, same era.
B. Desplanques, J. Thienpondt, and K. Demuynck, · 2020
Closest in time.
“Magneto: X-vector magnitude estimation network plus offset for improved speaker recognition,”
D. Garcia-Romero, G. Sell, and A. McCree, · 2020
Closest in time.
“Cross-lingual speaker verification with domain-balanced hard prototype mining and language-dependent score normalization,”
J. Thienpondt, B. Desplanques, and K. Demuynck, · 2020
Closest in time.
“The IDLAB VoxCeleb speaker recognition challenge 2020 system description,”
J. Thienpondt, B. Desplanques, and K. Demuynck, · 2020
Closest in time.
“Voxsrc 2020: The second voxceleb speaker recognition challenge,” 2020
A. Nagrani, J. S. Chung, J. Huh, A. Brown, E. Coto, W. Xie, M. McLaren, D. A. Reynolds, and A. Zisserman, · 2020
Closest in time.