Fetching the paper…
Reading the bibliography…
In this paper, we propose a simple but powerful unsupervised learning method for speaker recognition, namely Contrastive Equilibrium Learning (CEL), which increases the uncertainty on nuisance factors latent in the embeddings by employing the uniformity loss.
“A geometric characterization of maximum rényi entropy distributions,”
C. Vignat, A. Hero, and J. Costa, · 2006
Earlier work this paper cites.
“Universally optimal distribution of points on spheres,”
H. Cohn and A. Kumar, · 2007
Earlier work this paper cites.
“Musan: A music, speech, and noise corpus,”
D. Snyder, G. Chen, and D. Povey, · 2015
Earlier work this paper cites.
“ VoxCeleb
A. Nagrani, J. S. Chung, and A. Zisserman, · 2017
Earlier work this paper cites.
“ VoxCeleb2
J. S. Chung, A. Nagrani, and A. Zisserman, · 2018
Earlier work this paper cites.
“X-vectors: Robust dnn embeddings for speaker recognition,”
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, · 2018
Earlier work this paper cites.
“Attentive statistics pooling for deep speaker embedding,”
K. Okabe, T. Koshinaka, and K. Shinoda, · 2018
Earlier work this paper cites.
“Generalized end-to-end loss for speaker verification,”
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, · 2018
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
A. v. d. Oord, Y. Li, and O. Vinyals, · 2018
Earlier work this paper cites.
“Cosface: Large margin cosine loss for deep face recognition,”
H. Wang et al., · 2018
Earlier work this paper cites.
“Additive margin softmax for face verification,”
F. Wang, J. Cheng, W. Liu, and H. Liu, · 2018
Earlier work this paper cites.
“The voices from a distance challenge 2019.,”
M. K. Nandwana, J. v. Hout, C. Richey, M. McLaren, M. A. Barrios, and A. Lawson, · 2019
Cited alongside, same era.
“Utterance-level aggregation for speaker recognition in the wild,”
W. Xie, A. Nagrani, J. S. Chung, and A. Zisserman, · 2019
Cited alongside, same era.
“Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition,”
X. Xiang, S. Wang, H. Huang, Y. Qian, and K. Yu, · 2019
Cited alongside, same era.
“An unsupervised autoregressive model for speech representation learning,”
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, · 2019
Cited alongside, same era.
“Learning speaker representations with mutual information,”
M. Ravanelli and Y. Bengio, · 2019
Cited alongside, same era.
“Neural predictive coding using convolutional neural networks toward unsupervised learning of speaker characteristics,”
N. Inoue and K. Goto, · 2020
Closest in time.
“Augmentation adversarial training for unsupervised speaker recognition,”
J. Huh, H. S. Heo, J. Kang, S. Watanabe, and J. S. Chung, · 2020
Closest in time.
“Disentangled speaker and nuisance attribute embedding for robust speaker verification,”
W. H. Kang, S. H. Mun, M. H. Han, and N. S. Kim, · 2020
Closest in time.
“Understanding contrastive representation learning through alignment and uniformity on the hypersphere,”
T. Wang and P. Isola, · 2020
Closest in time.
“Momentum contrast for unsupervised visual representation learning,”
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Jati and P. Georgiou, · 2019
Cited alongside, same era.
Discrete energy on rectifiable sets
S. Borodachov, D. Hardin, and E. Saff, · 2019
Cited alongside, same era.
“Arcface: Additive angular margin loss for deep face recognition,”
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, · 2019
Cited alongside, same era.
“Adacos: Adaptively scaling cosine logits for effectively learning deep face representations,”
X. Zhang, R. Zhao, Y. Qiao, X. Wang, and H. Li, · 2019
Cited alongside, same era.
“In defence of metric learning for speaker recognition,”
J. S. Chung et al., · 2020
Cited alongside, same era.
“Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,”
A. T. Liu, S. Yang, P.-H. Chi, P. Hsu, and H. Lee, · 2020
Cited alongside, same era.
“Disentangled speech embeddings using cross-modal self-supervision,”
A. Nagrani, J. S. Chung, S. Albanie, and A. Zisserman, · 2020
Closest in time.
“Seeing voices and hearing voices: learning discriminative embeddings using cross-modal self-supervision,”
S.-W. Chung, H. G. Kang, and J. S. Chung, · 2020
Closest in time.
“Information preservation pooling for speaker embedding,”
M. H. Han, W. H. Kang, S. H. Mun, and N. S. Kim, · 2020
Closest in time.
“An end-to-end approach for the verification problem: learning the right distance,”
J. Monteiro, I. Albuquerque, J. Alam, R. D. Hjelm, and T. Falk, · 2020
Closest in time.
“Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning,”
Y. Zhang et al., · 2084
Closest in time.