Fetching the paper…
Reading the bibliography…
Time delay neural network (TDNN) has been proven to be efficient for speaker verification.
1910
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 770–778
2016
Earlier work this paper cites.
G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, p. 2261–2269
2017
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 5220–5224
2017
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5329–5333
2018
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 7132–7141
2018
Earlier work this paper cites.
K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statistics pooling for deep speaker embedding,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2018, pp. 2252–2256
2018
Earlier work this paper cites.
Y. Zhu, T. Ko, D. Snyder, B. Mak, and D. Povey, “Self-attentive speaker embeddings for text independent speaker verification,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2018, pp. 3573–3577
2018
Earlier work this paper cites.
S. Zheng, G. Liu, H. Suo, and Y. Lei, “Autoencoder-based semi-supervised curriculum learning for out-of-domain speaker verification,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2019, pp. 4360–4364
2019
Cited alongside, same era.
M. Indiay, P. Safariy, and J. Hernando, “Self multi-head attention for speaker recognition,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2019, pp. 4305–4309
2019
Cited alongside, same era.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arc-face: Additive angular margin loss for deep face recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 4690–4699
2019
Cited alongside, same era.
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2020, pp. 3830–3834
Z. Bai and X. Zhang, “Speaker recognition based on deep learning: An overview,” Neural Networks , vol. 140, pp. 65–99, 2021
2021
Later among the works it cites.
Y.-Q. Yu, S. Zheng, H. Suo, Y. Lei, and W.-J. Li, “Cam: Context-aware masking for robust speaker verification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 6703–6707
2021
Later among the works it cites.
J. Thienpondt, B. Desplanques, and K. Demuynck, “Integrating frequency translational invariance in tdnns and frequency positional information in 2d resnets to enhance speaker verification,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2021, pp. 2302–2306
2021
Later among the works it cites.
B. Liu, Z. Chen, S. Wang, H. Wang, B. Han, and Y. Qian, “Df-resnet: Boosting speaker verification performance withdepth-first design,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2022, pp. 296–300
2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Y.-Q. Yu and W.-J. Li, “Densely connected time delay neural network for speaker verification,” in Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2020, pp. 921–925
2020
Cited alongside, same era.
S. Zheng, Y. Lei, and H. Suo, “Phonetically-aware coupled network for short duration text-independent speaker verification,” in Interspeech 2020, 21st Annual Conference of the International Speech Communication Association . ISCA, 2020, pp. 926–930
2020
Cited alongside, same era.
A. Nagrani, J. S. Chung, W. Xie, , and A. Zisserman, “Voxceleb: Large-scale speaker verification in the wild,” Computer Speech and Language , vol. 60, 2020
2020
Cited alongside, same era.
Y. Fan, J. Kang, L. Li, K. Li, H. Chen, S. Cheng, P. Zhang, Z. Zhou, Y. Cai, and D. Wang, “CN-Celeb: a challenging chinese speaker recognition dataset,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7604–7608
2020
Cited alongside, same era.
Later among the works it cites.
C. Tan, Q. Chen, W. Wang, Q. Zhang, S. Zheng, and Z. Ling, “Ponet: Pooling network for efficient token mixing in long sequences,” in The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022
2022
Later among the works it cites.
T. Liu, R. K. Das, K. Aik Lee, and H. Li, “MFA: TDNN with multi-scale frequency-channel attention for text-independent speaker verification with short utterances,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 7517–7521
2022
Later among the works it cites.
L. Li, R. Liu, J. Kang, Y. Fan, H. Cui, Y. Cai, R. Vipperla, T. F. Zheng, and D. Wang, “CN-Celeb: multi-genre speaker recognition,” Speech Communication , 2022
2022
Later among the works it cites.