Fetching the paper…
Reading the bibliography…
We held the second installment of the VoxCeleb Speaker Recognition Challenge in conjunction with Interspeech 2020.
“Unleashing the killer corpus: experiences in creating the multi-everything ami meeting corpus,”
J. Carletta, · 2007
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Musan: A music, speech, and noise corpus,”
D. Snyder, G. Chen, and D. Povey, · 2015
Earlier work this paper cites.
“Out of time: automated lip sync in the wild,”
J. S. Chung and A. Zisserman, · 2016
Earlier work this paper cites.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Earlier work this paper cites.
“VoxCeleb: a large-scale speaker identification dataset,”
A. Nagrani, J. S. Chung, and A. Zisserman, · 2017
Earlier work this paper cites.
“A study on data augmentation of reverberant speech for robust speech recognition,”
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, · 2017
Earlier work this paper cites.
“Aggregated residual transformations for deep neural networks,”
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, · 2017
Earlier work this paper cites.
“Dual path networks,”
Y. Chen, J. Li, H. Xiao, X. Jin, S. Yan, and J. Feng, · 2017
Earlier work this paper cites.
“The 2016 nist speaker recognition evaluation.,”
S. O. Sadjadi, T. Kheyrkhah, A. Tong, C. S. Greenberg, D. A. Reynolds, E. Singer, L. P. Mason, and J. Hernandez-Cordero, · 2017
Earlier work this paper cites.
“Voxceleb2: Deep speaker recognition,”
J. S. Chung, A. Nagrani, and A. Zisserman, · 2018
Earlier work this paper cites.
“The conversation: Deep audio-visual speech enhancement,”
T. Afouras, J. S. Chung, and A. Zisserman, · 2018
Earlier work this paper cites.
https://www.nist.gov/system/files/documents/2018/08/17/sre18_eval_plan_2018-05-31_v6.pdf
NIST 2018 Speaker Recognition Evaluation Plan · 2018
Earlier work this paper cites.
“Diarization is hard: Some experiences and lessons learned for the jhu team in the inaugural dihard challenge.,”
G. Sell, D. Snyder, A. McCree, D. Garcia-Romero, J. Villalba, M. Maciejewski, V. Manohar, N. Dehak, D. Povey, S. Watanabe, et al., · 2018
Earlier work this paper cites.
“Deepmine speech processing database: Text-dependent and independent speaker verification and speech recognition in persian and english.,”
H. Zeinali, H. Sameti, and T. Stafylakis, · 2018
Earlier work this paper cites.
“Additive margin softmax for face verification,”
F. Wang, J. Cheng, W. Liu, and H. Liu, · 2018
Earlier work this paper cites.
“First dihard challenge evaluation plan,”
N. Ryant, K. Church, C. Cieri, A. Cristia, J. Du, S. Ganapathy, and M. Liberman, · 2018
Earlier work this paper cites.
“Voxsrc 2019: The first voxceleb speaker recognition challenge,”
J. S. Chung, A. Nagrani, E. Coto, W. Xie, M. McLaren, D. A. Reynolds, and A. Zisserman, · 2019
Earlier work this paper cites.
“Self-supervised speaker embeddings,”
T. Stafylakis, J. Rohdin, O. Plchot, P. Mizera, and L. Burget, · 2019
Cited alongside, same era.
“The second dihard diarization challenge: Dataset, task, and baselines,”
N. Ryant, K. Church, C. Cieri, A. Cristia, J. Du, S. Ganapathy, and M. Liberman, · 2019
Cited alongside, same era.
“Res2net: A new multi-scale backbone architecture,”
S. Gao, M.-M. Cheng, K. Zhao, X.-Y. Zhang, M.-H. Yang, and P. H. Torr, · 2019
Cited alongside, same era.
“Dover: A method for combining diarization outputs,”
A. Stolcke and T. Yoshioka, · 2019
Cited alongside, same era.
“The voices from a distance challenge 2019 evaluation plan,”
M. K. Nandwana, J. Van Hout, M. McLaren, C. Richey, A. Lawson, and M. A. Barrios, · 2019
Cited alongside, same era.
U. Khan and J. Hernando, · 2020
Closest in time.
“Analysis of the but diarization system for voxconverse challenge,”
F. Landini, O. Glembek, P. Matějka, J. Rohdin, L. Burget, M. Diez, and A. Silnova, · 2020
Closest in time.
“Microsoft speaker diarization system for the voxceleb speaker recognition challenge 2020,”
X. Xiao, N. Kanda, Z. Chen, T. Zhou, T. Yoshioka, Y. Zhao, G. Liu, J. Wu, J. Li, and Y. Gong, · 2020
Closest in time.
“In defence of metric learning for speaker recognition,”
J. S. Chung, J. Huh, S. Mun, M. Lee, H. S. Heo, S. Choe, C. Ham, S. Jung, B.-J. Lee, and I. Han, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“BUT system description to voxceleb speaker recognition challenge 2019,”
H. Zeinali, S. Wang, A. Silnova, P. Matějka, and O. Plchot, · 2019
Cited alongside, same era.
“Spot the conversation: speaker diarisation in the wild,”
J. S. Chung, J. Huh, A. Nagrani, T. Afouras, and A. Zisserman, · 2020
Cited alongside, same era.
“Playing a part: Speaker verification at the movies,”
A. Brown, J. Huh, A. Nagrani, J. S. Chung, and A. Zisserman, · 2020
Cited alongside, same era.
“Condensed movies: Story based retrieval with contextual embeddings,”
M. Bain, A. Nagrani, A. Brown, and A. Zisserman, · 2020
Cited alongside, same era.
“Momentum contrast speaker representation learning,”
J. Lee, J. Koh, and S. Yoon, · 2020
Cited alongside, same era.
“Disentangled speech embeddings using cross-modal self-supervision,”
A. Nagrani, J. S. Chung, S. Albanie, and A. Zisserman, · 2020
Cited alongside, same era.
“Augmentation adversarial training for unsupervised speaker recognition,”
J. Huh, H. S. Heo, J. Kang, S. Watanabe, and J. S. Chung, · 2020
Cited alongside, same era.
N. Inoue and K. Goto, · 2020
Closest in time.
“Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,”
B. Desplanques, J. Thienpondt, and K. Demuynck, · 2020
Closest in time.
“Sub-center arcface: Boosting face recognition by large-scale noisy web faces,”
J. Deng, J. Guo, T. Liu, M. Gong, and S. Zafeiriou, · 2020
Closest in time.
“Momentum contrast for unsupervised visual representation learning,”
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, · 2020
Closest in time.
“Sub-center arcface: Boosting face recognition by large-scale noisy web faces,”
J. Deng, J. Guo, T. Liu, M. Gong, and S. Zafeiriou, · 2020
Closest in time.
“A simple framework for contrastive learning of visual representations,”
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, · 2020
Closest in time.
“A framework for contrastive self-supervised learning and designing a new approach,”
W. Falcon and K. Cho, · 2020
Closest in time.
“Conformer: Convolution-augmented transformer for speech recognition,”
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, et al., · 2020
Closest in time.
“Pyannote. audio: neural building blocks for speaker diarization,”
H. Bredin, R. Yin, J. M. Coria, G. Gelly, P. Korshunov, M. Lavechin, D. Fustes, H. Titeux, W. Bouaziz, and M.-P. Gill, · 2020
Closest in time.
“The 2019 nist speaker recognition evaluation cts challenge,”
S. O. Sadjadi, C. Greenberg, E. Singer, D. Reynolds, L. Mason, and J. Hernandez-Cordero, · 2020
Closest in time.
“The interspeech 2020 far-field speaker verification challenge,”
X. Qin, M. Li, H. Bu, W. Rao, R. K. Das, S. Narayanan, and H. Li, · 2020
Closest in time.
“Third dihard challenge evaluation plan,”
N. Ryant, K. Church, C. Cieri, J. Du, S. Ganapathy, and M. Liberman, · 2020
Closest in time.
“Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,”
S. Watanabe, M. Mandel, J. Barker, and E. Vincent, · 2020
Closest in time.