Fetching the paper…
Reading the bibliography…
The speech representations learned from large-scale unlabeled data have shown better generalizability than those from supervised learning and thus attract a lot of interest to be applied for various downstream tasks.
Z. N. Karam, W. M. Campbell, and N. Dehak, “Towards reduced false-alarms using cohorts,” in
2011
Earlier work this paper cites.
S. Cumani, P. D. Batzu, D. Colibro, C. Vair, P. Laface, and V. Vasilakakis, “Comparison of speaker recognition approaches for real applications.” in
2011
Earlier work this paper cites.
Y. Liu, Y. Qian, N. Chen, T. Fu, Y. Zhang, and K. Yu, “Deep feature for text-dependent speaker verification,”
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Todisco, H. Delgado, and N. W. Evans, “Articulation rate filtering of cqcc features for automatic speaker verification.” in
2016
Earlier work this paper cites.
C. Zhang and K. Koishida, “End-to-end text-independent speaker verification with triplet loss on short utterances.” in
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: a large-scale speaker identification dataset,”
2017
Earlier work this paper cites.
M. Todisco, H. Delgado, and N. Evans, “Constant q cepstral coefficients: A spoofing countermeasure for automatic speaker verification,”
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in
2018
Earlier work this paper cites.
Z. Huang, S. Wang, and K. Yu, “Angular softmax for short-duration text-independent speaker verification.” in
2018
Earlier work this paper cites.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in
2018
Earlier work this paper cites.
K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statistics pooling for deep speaker embedding,”
2018
Earlier work this paper cites.
Y. Zhu, T. Ko, D. Snyder, B. Mak, and D. Povey, “Self-attentive speaker embeddings for text-independent speaker verification.” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” 2018
2018
Cited alongside, same era.
M. Ravanelli and Y. Bengio, “Speaker recognition from raw waveform with sincnet,” in
2018
Cited alongside, same era.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in
2018
Cited alongside, same era.
2020
Later among the works it cites.
2021
Closest in time.
2021
Closest in time.
H. Zhang, Y. Zou, and H. Wang, “Contrastive self-supervised learning for text-independent speaker verification,” in
2021
Closest in time.
W. Xia, C. Zhang, C. Weng, M. Yu, and D. Yu, “Self-supervised text-independent speaker verification using prototypical momentum contrastive learning,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2019
Cited alongside, same era.
X. Xiang, S. Wang, H. Huang, Y. Qian, and K. Yu, “Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition,” in
2019
Cited alongside, same era.
G. Jawahar, B. Sagot, and D. Seddah, “What does BERT learn about the structure of language?” in
2019
Cited alongside, same era.
2019
Cited alongside, same era.
S. Gao, M.-M. Cheng, K. Zhao, X.-Y. Zhang, M.-H. Yang, and P. H. Torr, “Res2net: A new multi-scale backbone architecture,”
2019
Cited alongside, same era.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in
2019
Cited alongside, same era.
2021
Closest in time.
D. Cai, W. Wang, and M. Li, “An iterative framework for self-supervised deep speaker representation learning,” in
2021
Closest in time.
2021
Closest in time.
Z. Fan, M. Li, S. Zhou, and B. Xu, “Exploring wav2vec 2.0 on Speaker Verification and Language Identification,” in
2021
Closest in time.
2021
Closest in time.
S. Chen, Y. Wu, C. Wang, Z. Chen, Z. Chen, S. Liu, J. Wu, Y. Qian, F. Wei, J. Li
2021
Closest in time.
L. Pepino, P. Riera, and L. Ferrer, “Emotion Recognition from Speech Using wav2vec 2.0 Embeddings,” in
2021
Closest in time.
J. Thienpondt, B. Desplanques, and K. Demuynck, “The idlab voxsrc-20 submission: Large margin fine-tuning and quality-aware score calibration in dnn based speaker verification,” in
2021
Closest in time.