Fetching the paper…
Reading the bibliography…
Effective fusion of multi-scale features is crucial for improving speaker verification performance.
W. M. Campbell, J. P. Campbell, D. A. Reynolds, E. Singer, and P. A. Torres-Carrasquillo, “Support vector machines for speaker and language recognition,” Comput. Speech Lang. , vol. 20, no. 2-3, pp. 210–229, 2006
2006
Earlier work this paper cites.
P. Kenny, G. Boulianne, P. Ouellet, and P. Dumouchel, “Joint factor analysis versus eigenchannels in speaker recognition,” IEEE Trans. Speech Audio Process. , vol. 15, no. 4, pp. 1435–1447, 2007
2007
Earlier work this paper cites.
S. J. D. Prince and J. H. Elder, “Probabilistic linear discriminant analysis for inferences about identity,” in ICCV 2007 . IEEE Computer Society, 2007, pp. 1–8
2007
Earlier work this paper cites.
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Trans. Speech Audio Process. , vol. 19, no. 4, pp. 788–798, 2011
2011
Earlier work this paper cites.
J. Christie, J. P. Ginsberg, J. Steedman, J. Fridriksson, L. Bonilha, and C. Rorden, “Global versus local processing: seeing the left side of the forest and the right side of the trees,” Frontiers in human neuroscience , vol. 6, p. 28, 2012
2012
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML 2015 , ser. JMLR Workshop and Conference Proceedings, F. R. Bach and D. M. Blei, Eds., vol. 37. JMLR.org, 2015, pp. 448–456
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR 2016 . IEEE Computer Society, 2016, pp. 770–778
2016
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, D. Povey, and S. Khudanpur, “Deep neural network embeddings for text-independent speaker verification,” in Interspeech 2017 , F. Lacerda, Ed. ISCA, 2017, pp. 999–1003
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: A large-scale speaker identification dataset,” in Interspeech 2017 , F. Lacerda, Ed. ISCA, 2017, pp. 2616–2620
2017
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in ICASSP 2017 . IEEE, 2017, pp. 5220–5224
2017
Earlier work this paper cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in Interspeech 2018 , B. Yegnanarayana, Ed. ISCA, 2018, pp. 1086–1090
2018
Cited alongside, same era.
Y. Tang, G. Ding, J. Huang, X. He, and B. Zhou, “Deep speaker embedding learning with multi-level pooling for text-independent speaker verification,” in ICASSP 2019 . IEEE, 2019, pp. 6116–6120
2019
Cited alongside, same era.
S. Seo, D. J. Rim, M. Lim, D. Lee, H. Park, J. Oh, C. Kim, and J. Kim, “Shortcut connections based deep speaker embeddings for end-to-end speaker verification system,” in Interspeech 2019 , G. Kubin and Z. Kacic, Eds. ISCA, 2019, pp. 2928–2932
2019
Cited alongside, same era.
A. Hajavi and A. Etemad, “A deep neural network for short-segment speaker recognition,” in Interspeech 2019 , G. Kubin and Z. Kacic, Eds. ISCA, 2019, pp. 2878–2882
2019
Cited alongside, same era.
T. Zhou, Y. Zhao, and J. Wu, “Resnext and res2net structures for speaker verification,” in IEEE Spoken Language Technology Workshop, SLT 2021 . IEEE, 2021, pp. 301–307
2021
Later among the works it cites.
S. Gao, M. Cheng, K. Zhao, X. Zhang, M. Yang, and P. H. S. Torr, “Res2net: A new multi-scale backbone architecture,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 43, no. 2, pp. 652–662, 2021
2021
Later among the works it cites.
J. Thienpondt, B. Desplanques, and K. Demuynck, “The idlab voxsrc-20 submission: Large margin fine-tuning and quality-aware score calibration in DNN based speaker verification,” in ICASSP 2021 . IEEE, 2021, pp. 5814–5818
2021
Later among the works it cites.
Y. Yu, S. Zheng, H. Suo, Y. Lei, and W. Li, “Cam: Context-aware masking for robust speaker verification,” in ICASSP 2021 . IEEE, 2021, pp. 6703–6707
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification,” in Interspeech 2020 , H. Meng, B. Xu, and T. F. Zheng, Eds. ISCA, 2020, pp. 3830–3834
2020
Cited alongside, same era.
Y. Jung, S. M. Kye, Y. Choi, M. Jung, and H. Kim, “Improving multi-scale aggregation using feature pyramid module for robust speaker verification of variable-duration utterances,” in Interspeech 2020 , H. Meng, B. Xu, and T. F. Zheng, Eds. ISCA, 2020, pp. 1501–1505
2020
Cited alongside, same era.
W. Cai, J. Chen, J. Zhang, and M. Li, “On-the-fly data loader and utterance-level aggregation for speaker and language recognition,” IEEE ACM Trans. Audio Speech Lang. Process. , vol. 28, pp. 1038–1051, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Y. Chen, W. Guo, and B. Gu, “Improved meta-learning training for speaker verification,” in Interspeech 2021 , H. Hermansky, H. Cernocký, L. Burget, L. Lamel, O. Scharenborg, and P. Motlícek, Eds. ISCA, 2021, pp. 1049–1053
2021
Cited alongside, same era.
B. Liu, Z. Chen, S. Wang, H. Wang, B. Han, and Y. Qian, “Df-resnet: Boosting speaker verification performance with depth-first design,” in Interspeech 2022 , H. Ko and J. H. L. Hansen, Eds. ISCA, 2022, pp. 296–300
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Zhang, Z. Lv, H. Wu, S. Zhang, P. Hu, Z. Wu, H. Lee, and H. Meng, “Mfa-conformer: Multi-scale feature aggregation conformer for automatic speaker verification,” in Interspeech 2022 , H. Ko and J. H. L. Hansen, Eds. ISCA, 2022, pp. 306–310
2022
Later among the works it cites.
J. Deng, J. Guo, J. Yang, N. Xue, I. Kotsia, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 44, no. 10, pp. 5962–5979, 2022
2022
Later among the works it cites.
B. Liu, H. Wang, Z. Chen, S. Wang, and Y. Qian, “Self-knowledge distillation via feature enhancement for speaker verification,” in ICASSP 2022 . IEEE, 2022, pp. 7542–7546
2022
Later among the works it cites.
B. Gu, W. Guo, and J. Zhang, “Memory storable network based feature aggregation for speaker representation learning,” IEEE ACM Trans. Audio Speech Lang. Process. , vol. 31, pp. 643–655, 2023
2023
Closest in time.