Fetching the paper…
Reading the bibliography…
Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size.
W. M. Fisher, G. R. Doddington, and K. M. Goudie-Marshall, “The DARPA speech recognition research database: specifications and status,” in
1986
Earlier work this paper cites.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “Switchboard: Telephone speech corpus for research and development,” in
1992
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1,”
1993
Earlier work this paper cites.
J. B. Millar, J. P. Vonwiller, J. M. Harrington, and P. J. Dermody, “The Australian national database of spoken language,” in
1994
Earlier work this paper cites.
D. A. Reynolds and R. C. Rose, “Robust text-independent speaker identification using gaussian mixture speaker models,”
1995
Earlier work this paper cites.
D. A. Reynolds, T. F. Quatieri, and R. B. Dunn, “Speaker verification using adapted gaussian mixture models,”
2000
Earlier work this paper cites.
J. Hennebert, H. Melin, D. Petrovska, and D. Genoud, “POLYCOST: a telephone-speech database for speaker recognition,”
2000
Earlier work this paper cites.
J. H. Hansen, R. Sarikaya, U. H. Yapanel, and B. L. Pellom, “Robust speech recognition in noise: an evaluation using the spine corpus.,” in
2001
Earlier work this paper cites.
U. H. Yapanel, X. Zhang, and J. H. Hansen, “High performance digit recognition in real car environments.,” in
2002
Earlier work this paper cites.
A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stolcke,
2003
Earlier work this paper cites.
P. Kenny, “Joint factor analysis of speaker and session variability: Theory and algorithms,”
2005
Earlier work this paper cites.
I. McCowan, J. Carletta, W. Kraaij, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos,
2005
Earlier work this paper cites.
L. Feng and L. K. Hansen, “A new database for speaker recognition,” tech. rep., 2005
2005
Earlier work this paper cites.
S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” in
2005
Earlier work this paper cites.
R. Woo, A. Park, and T. J. Hazen, “The MIT Mobile Device Speaker Verification Corpus: Data collection and preliminary experiments,”
2006
Cited alongside, same era.
S. Ioffe, “Probabilistic linear discriminant analysis,” in
2006
Cited alongside, same era.
C. McCool and S. Marcel, “Mobio database for the ICPR 2010 face and speech competition,” tech. rep., IDIAP, 2009
2009
Cited alongside, same era.
D. E. King, “Dlib-ml: A machine learning toolkit,”
2009
Cited alongside, same era.
M. Everingham, J. Sivic, and A. Zisserman, “Taking the bite out of automatic naming of characters in TV video,”
2009
Cited alongside, same era.
L. L. Stoll, “Finding difficult speakers in automatic speaker recognition,”
2011
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,”
2015
Later among the works it cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, S. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. Berg, and F. Li, “Imagenet large scale visual recognition challenge,”
2015
Later among the works it cites.
T. N. Sainath, R. J. Weiss, A. W. Senior, K. W. Wilson, and O. Vinyals, “Learning the speech front-end with raw waveform CLDNNs,” in
2015
Later among the works it cites.
G. Morrison, C. Zhang, E. Enzinger, F. Ochoa, D. Bleach, M. Johnson, B. Folkes, S. De Souza, N. Cummins, and D. Chow, “Forensic database of voice recordings of 500+ Australian English speakers,”
2015
Later among the works it cites.
P. Bell, M. J. Gales, T. Hain, J. Kilgour, P. Lanchantin, X. Liu, A. McParland, S. Renals, O. Saz, M. Wester,
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in
2012
Cited alongside, same era.
C. S. Greenberg, “The NIST year 2012 speaker recognition evaluation plan,”
2012
Cited alongside, same era.
D. van der Vloed, J. Bouten, and D. A. van Leeuwen, “NFI-FRITS: a forensic speaker recognition database and some first experiments,” in
2014
Cited alongside, same era.
V. Kazemi and J. Sullivan, “One millisecond face alignment with an ensemble of regression trees,” in
2014
Cited alongside, same era.
K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman, “Return of the devil in the details: Delving deep into convolutional nets,” in
2014
Cited alongside, same era.
2015
Later among the works it cites.
O. M. Parkhi, A. Vedaldi, and A. Zisserman, “Deep face recognition,” in
2015
Later among the works it cites.
2015
Later among the works it cites.
M. McLaren, L. Ferrer, D. Castan, and A. Lawson, “The speakers in the wild (SITW) speaker recognition database,”
2016
Later among the works it cites.
2016
Later among the works it cites.
Y. Lukic, C. Vogt, O. Dürr, and T. Stadelmann, “Speaker identification and clustering using convolutional neural networks,” in
2016
Later among the works it cites.
2016
Later among the works it cites.
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in
2016
Later among the works it cites.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in
2016
Later among the works it cites.