Fetching the paper…
Reading the bibliography…
Recently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary.
“Ther DARPA speech recognition research database: specifications and status,”
William M Fisher, · 1986
Earlier work this paper cites.
“SWITCHBOARD: Telephone speech corpus for research and development,”
John J Godfrey, Edward C Holliman, and Jane McDaniel, · 1992
Earlier work this paper cites.
“Probabilistic linear discriminant analysis,”
Sergey Ioffe, · 2006
Earlier work this paper cites.
“Joint factor analysis versus eigenchannels in speaker recognition,”
Patrick Kenny, Gilles Boulianne, Pierre Ouellet, and Pierre Dumouchel, · 2007
Earlier work this paper cites.
“Visual object tracking using adaptive correlation filters,”
David S Bolme, J Ross Beveridge, Bruce A Draper, and Yui Man Lui, · 2010
Earlier work this paper cites.
Finding difficult speakers in automatic speaker recognition
Lara Lynn Stoll, · 2011
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
Najim Dehak, Patrick J Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet, · 2011
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition,”
Geoffrey Hinton, Li Deng, Dong Yu, George Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Brian Kingsbury, et al., · 2012
Cited alongside, same era.
“Imagenet classification with deep convolutional neural networks,”
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, · 2012
Cited alongside, same era.
“Very deep convolutional networks for large-scale image recognition,”
Karen Simonyan and Andrew Zisserman, · 2014
Cited alongside, same era.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Cited alongside, same era.
“Out of time: automated lip sync in the wild,”
Joon Son Chung and Andrew Zisserman, · 2016
Cited alongside, same era.
“The Speakers in the Wild (SITW) speaker recognition database.,”
“The 2016 NIST speaker recognition evaluation.,”
Seyed Omid Sadjadi, Timothée Kheyrkhah, Audrey Tong, Craig S Greenberg, Douglas A Reynolds, Elliot Singer, Lisa P Mason, and Jaime Hernandez-Cordero, · 2017
Later among the works it cites.
“Voxceleb: a large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Later among the works it cites.
“X-vectors: Robust DNN embeddings for speaker recognition,”
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, · 2018
Later among the works it cites.
“Voxceleb2: Deep speaker recognition,”
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2018
Later among the works it cites.
“Retinaface: Single-stage dense face localisation in the wild,”
Jiankang Deng, Jia Guo, Yuxiang Zhou, Jinke Yu, Irene Kotsia, and Stefanos Zafeiriou, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mitchell McLaren, Luciana Ferrer, Diego Castan, and Aaron Lawson, · 2016
Cited alongside, same era.
Robustness-related issues in speaker recognition
Thomas Fang Zheng and Lantian Li, · 2017
Cited alongside, same era.
“Deep speaker feature learning for text-independent speaker verification,”
Lantian Li, Yixiang Chen, Ying Shi, Zhiyuan Tang, and Dong Wang, · 2017
Cited alongside, same era.
Closest in time.
“Arcface: Additive angular margin loss for deep face recognition,”
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou, · 2019
Closest in time.
“Utterance-level aggregation for speaker recognition in the wild,”
Weidi Xie, Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2019
Closest in time.