Fetching the paper…
Reading the bibliography…
One of the most popular speaker embeddings is x-vectors, which are obtained from an architecture that gradually builds a larger temporal context with layers.
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, “Phoneme recognition using time-delay neural networks,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 37, no. 3, pp. 328–339, 1989
1989
Earlier work this paper cites.
S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” in Proc. of Computer Vision and Pattern Recognition (CVPR) . IEEE, 2005, pp. 539–546
2005
Earlier work this paper cites.
S. Ioffe, “Probabilistic linear discriminant analysis,” in Proc. of European Conference on Computer Vision . Springer, 2006, pp. 531–542
2006
Earlier work this paper cites.
R. Hadsell, S. Chopra, and Y. LeCun, “Dimensionality reduction by learning an invariant mapping,” in Proc. of Computer Vision and Pattern Recognition (CVPR) . IEEE, 2006, pp. 1735–1742
2006
Earlier work this paper cites.
2008
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 4, pp. 788–798, 2011
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The Kaldi speech recognition toolkit,” in Workshop on Automatic Speech Recognition and Understanding . IEEE Signal Processing Society, 2011
2011
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” in Proc. of Empirical Methods in Natural Language Processing , 2015, pp. 1412–1421
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. of Computer Vision and Pattern Recognition (CVPR) . IEEE, 2016, pp. 770–778
2016
Earlier work this paper cites.
R. Arandjelovic, P. Gronat, A. Torii, T. Pajdla, and J. Sivic, “NetVLAD: CNN architecture for weakly supervised place recognition,” in Proc. of Computer Vision and Pattern Recognition (CVPR) . IEEE, 2016, pp. 5297–5307
2016
Earlier work this paper cites.
D. Snyder, D. Garcia Romero, D. Povey, and S. Khudanpur, “Deep neural network embeddings for text-independent speaker verification.” in Proc. of Interspeech , 2017, pp. 999–1003
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. of Advances in Neural Information Processing Systems (NIPS) , 2017, pp. 5998–6008
2017
Cited alongside, same era.
A. Nagrani, J. S. Chung, and A. Zisserman, “VoxCeleb: A large-scale speaker identification dataset,” in Proc. of Interspeech , 2017, pp. 2616–2620
2017
Cited alongside, same era.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 5220–5224
2017
Cited alongside, same era.
G. Sun, C. Zhang, and P. C. Woodland, “Speaker diarisation using 2D self-attentive combination of embeddings,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5801–5805
2019
Later among the works it cites.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural speech synthesis with transformer network,” in Proc. of AAAI Conference on Artificial Intelligence , 2019, pp. 6706–6713
2019
Later among the works it cites.
A. Nagrani, J. S. Chung, W. Xie, and A. Zisserman, “VoxCeleb: Large-scale speaker verification in the wild,” Computer Speech & Language , vol. 60, p. 101027, 2020
2020
Closest in time.
Y. Zhang, H. Yu, and Z. Ma, “Speaker verification system based on deformable CNN and time-frequency attention,” in Proc. of Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) . IEEE, 2020, pp. 1689–1692
2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Snyder, D. Garcia Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5329–5333
2018
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep speaker recognition,” in Proc. of Interspeech , 2018, pp. 1086–1090
2018
Cited alongside, same era.
F. Wang, J. Cheng, W. Liu, and H. Liu, “Additive margin softmax for face verification,” IEEE Signal Processing Letters , vol. 25, no. 7, pp. 926–930, 2018
2018
Cited alongside, same era.
K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statistics pooling for deep speaker embedding,” in Proc. of Interspeech , 2018, pp. 2252–2256
2018
Cited alongside, same era.
Y. Zhu, T. Ko, D. Snyder, B. Mak, and D. Povey, “Self-attentive speaker embeddings for text-independent speaker verification.” in Proc. of Interspeech , vol. 2018, 2018, pp. 3573–3577
2018
Cited alongside, same era.
Y. Zhong, R. Arandjelović, and A. Zisserman, “GhostVLAD for set-based face recognition,” in Proc. of Asian Conference on Computer Vision . Springer, 2018, pp. 35–50
2018
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5884–5888
2018
Cited alongside, same era.
2018
Cited alongside, same era.
V. M. Shetty and N. J. Metilda Sagaya Mary, “Improving the performance of transformer based low resource speech recognition for Indian languages,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 8279–8283
2020
Closest in time.
N. J. Metilda Sagaya Mary, V. M. Shetty, and S. Umesh, “Investigation of methods to improve the recognition performance of Tamil-English code-switched data in transformer framework,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7889–7893
2020
Closest in time.
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,” in Proc. of Interspeech , 2020, pp. 3830–3834
2020
Closest in time.
Y. Tong, W. Xue, S. Huang, L. Fan, C. Zhang, G. Ding, and X. He, “The JD AI speaker verification system for the FFSVC 2020 challenge.” in Proc. of Interspeech , 2020, pp. 3476–3480
2020
Closest in time.
2020
Closest in time.
W. W. Lin and M. W. Mak, “Wav2Spk: A simple DNN architecture for learning speaker embeddings from waveforms.” in Proc. of Interspeech , 2020, pp. 3211–3215
2020
Closest in time.
C. Wang, J. Yi, J. Tao, Y. Bai, and Z. Tian, “Hierarchically attending time-frequency and channel features for improving speaker verification,” in Proc. of International Symposium on Chinese Spoken Language Processing (ISCSLP) . IEEE, 2021, pp. 1–5
2021
Closest in time.
A. T. Liu, S. W. Li, and H. y. Lee, “TERA: Self-supervised learning of transformer encoder representation for speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2351–2366, 2021
2021
Closest in time.