Fetching the paper…
Reading the bibliography…
The goal of this work is to train robust speaker recognition models without speaker labels.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Bayesian speaker verification with heavy-tailed priors
Patrick Kenny · 2010
Earlier work this paper cites.
Discriminatively trained probabilistic linear discriminant analysis for speaker verification
Lukáš Burget, Oldřich Plchot, Sandro Cumani, Ondřej Glembek, Pavel Matějka, and Niko Brümmer · 2011
Earlier work this paper cites.
Front-end factor analysis for speaker verification
Najim Dehak, Patrick J Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet · 2011
Earlier work this paper cites.
Full-covariance ubm and heavy-tailed plda in i-vector speaker verification
Pavel Matějka, Ondřej Glembek, Fabio Castaldo, Md Jahangir Alam, Oldřich Plchot, Patrick Kenny, Lukáš Burget, and Jan Černocky · 2011
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al · 2011
Earlier work this paper cites.
Applications of speaker recognition
Nilu Singh, RA Khan, and Raj Shree · 2012
Earlier work this paper cites.
Speaker recognition anti-spoofing
Nicholas Evans, Tomi Kinnunen, Junichi Yamagishi, Zhizheng Wu, Federico Alegre, and Phillip De Leon · 2014
Earlier work this paper cites.
Musan: A music, speech, and noise corpus
David Snyder, Guoguo Chen, and Daniel Povey · 2015
Earlier work this paper cites.
Simultaneous deep transfer across domains and tasks
Eric Tzeng, Judy Hoffman, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky · 2016
Earlier work this paper cites.
The speakers in the wild (SITW) speaker recognition database
Mitchell McLaren, Luciana Ferrer, Diego Castan, and Aaron Lawson · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros · 2016
Earlier work this paper cites.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2016
Earlier work this paper cites.
Colorful image colorization
Richard Zhang, Phillip Isola, and Alexei A Efros · 2016
Cited alongside, same era.
Look, listen and learn
Relja Arandjelović. and Andrew Zisserman · 2017
Cited alongside, same era.
VoxCeleb: a large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman · 2017
Cited alongside, same era.
https://www.nist.gov/system/files/documents/2018/08/17/sre18_eval_plan_2018-05-31_v6.pdf , See Section 3.1
NIST 2018 Speaker Recognition Evaluation Plan · 2018
Cited alongside, same era.
VoxCeleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Cited alongside, same era.
Unified hypersphere embedding for speaker recognition
Mahdi Hajibabaei and Dengxin Dai · 2018
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Learning speaker representations with mutual information
Mirco Ravanelli and Yoshua Bengio · 2019
Later among the works it cites.
Speaker verification using end-to-end adversarial language adaptation
Johan Rohdin, Themos Stafylakis, Anna Silnova, Hossein Zeinali, Lukáš Burget, and Oldřich Plchot · 2019
Later among the works it cites.
Self-supervised audio-visual co-segmentation
Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh McDermott, and Antonio Torralba · 2019
Later among the works it cites.
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Angular softmax loss for end-to-end speaker verification
Yutian Li, Feng Gao, Zhijian Ou, and Jiasong Sun · 2018
Cited alongside, same era.
Voices obscured in complex environmental settings (voices) corpus
Colleen Richey, Maria A Barrios, Zeb Armstrong, Chris Bartels, Horacio Franco, Martin Graciarena, Aaron Lawson, Mahesh Kumar Nandwana, Allen Stauffer, Julien van Hout, et al · 2018
Cited alongside, same era.
X-vectors: Robust dnn embeddings for speaker recognition
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur · 2018
Cited alongside, same era.
Generalized end-to-end loss for speaker verification
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno · 2018
Cited alongside, same era.
Unsupervised domain adaptation via domain adversarial training for speaker recognition
Qing Wang, Wei Rao, Sining Sun, Leib Xie, Eng Siong Chng, and Haizhou Li · 2018
Cited alongside, same era.
Few shot speaker recognition using deep neural networks
Prashant Anand, Ajeet Kumar Singh, Siddharth Srivastava, and Brejesh Lall · 2019
Cited alongside, same era.
Jixuan Wang, Kuan-Chieh Wang, Marc T Law, Frank Rudzicz, and Michael Brudno · 2019
Later among the works it cites.
Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition
Xu Xiang, Shuai Wang, Houjun Huang, Yanmin Qian, and Kai Yu · 2019
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Closest in time.
Delving into VoxCeleb: environment invariant speaker recognition
Joon Son Chung, Jaesung Huh, and Seongkyu Mun · 2020
Closest in time.
In defence of metric learning for speaker recognition
Joon Son Chung, Jaesung Huh, Seongkyu Mun, Minjae Lee, Hee Soo Heo, Soyeon Choe, Chiheon Ham, Sunghwan Jung, Bong-Jin Lee, and Icksang Han · 2020
Closest in time.
Seeing voices and hearing voices: learning discriminative embeddings using cross-modal self-supervision
Soo-Whan Chung, Hong Goo Kang, and Joon Son Chung · 2020
Closest in time.
Nakamasa Inoue and Keita Goto · 2020
Closest in time.
Channel adversarial training for speaker verification and diarization
Chau Luu, Peter Bell, and Steve Renals · 2020
Closest in time.
Self-supervised learning of pretext-invariant representations
Ishan Misra and Laurens van der Maaten · 2020
Closest in time.
Disentangled speech embeddings using cross-modal self-supervision
Arsha Nagrani, Joon Son Chung, Samuel Albanie, and Andrew Zisserman · 2020
Closest in time.