Fetching the paper…
Reading the bibliography…
The objective of this paper is speaker recognition under noisy and unconstrained conditions.
W. M. Fisher, G. R. Doddington, and K. M. Goudie-Marshall, “The DARPA speech recognition research database: specifications and status,” in
1986
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1,”
1993
Earlier work this paper cites.
J. B. Millar, J. P. Vonwiller, J. M. Harrington, and P. J. Dermody, “The Australian national database of spoken language,” in
1994
Earlier work this paper cites.
K.-K. Sung, “Learning and example selection for object and pattern detection,” Ph.D. dissertation, 1996
1996
Earlier work this paper cites.
J. Hennebert, H. Melin, D. Petrovska, and D. Genoud, “POLYCOST: a telephone-speech database for speaker recognition,”
2000
Earlier work this paper cites.
S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” in
2005
Earlier work this paper cites.
R. Hadsell, S. Chopra, and Y. LeCun, “Dimensionality reduction by learning an invariant mapping,” in
2006
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
P. Matějka, O. Glembek, F. Castaldo, M. J. Alam, O. Plchot, P. Kenny, L. Burget, and J. Černocky, “Full-covariance ubm and heavy-tailed plda in i-vector speaker verification,” in
2011
Earlier work this paper cites.
D. Chen, S. Tsai, V. Chandrasekhar, G. Takacs, H. Chen, R. Vedantham, R. Grzeszczuk, and B. Girod, “Residual enhanced visual vectors for on-device image matching,” in
2011
Earlier work this paper cites.
C. S. Greenberg, “The NIST year 2012 speaker recognition evaluation plan,”
2012
Earlier work this paper cites.
S. Cumani, O. Plchot, and P. Laface, “Probabilistic linear discriminant analysis of i-vector posterior distributions,” in
2013
Earlier work this paper cites.
Y. Taigman, M. Yang, M. Ranzato, and L. Wolf, “Deepface: Closing the gap to human-level performance in face verification,” in
2014
Earlier work this paper cites.
K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman, “Return of the devil in the details: Delving deep into convolutional nets,” in
2014
Earlier work this paper cites.
E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez-Dominguez, “Deep neural networks for small footprint text-dependent speaker verification,” in
2014
Earlier work this paper cites.
Y. Lei, N. Scheffer, L. Ferrer, and M. McLaren, “A novel scheme for speaker recognition using a phonetically-aware deep neural network,” in
2014
Cited alongside, same era.
S. H. Yella, A. Stolcke, and M. Slaney, “Artificial neural network features for speaker diarization,” in
2014
Cited alongside, same era.
D. van der Vloed, J. Bouten, and D. A. van Leeuwen, “NFI-FRITS: a forensic speaker recognition database and some first experiments,” in
2014
Cited alongside, same era.
A. Vedaldi and K. Lenc, “Matconvnet – convolutional neural networks for matlab,”
2014
Cited alongside, same era.
F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in
2015
Cited alongside, same era.
2017
Later among the works it cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “VoxCeleb: a large-scale speaker identification dataset,” in
2017
Later among the works it cites.
J. S. Chung, A. Jamaludin, and A. Zisserman, “You said that?” in
2017
Later among the works it cites.
T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen, “Audio-driven facial animation by joint end-to-end learning of pose and emotion,”
2017
Later among the works it cites.
D. Snyder, D. Garcia-Romero, D. Povey, and S. Khudanpur, “Deep neural network embeddings for text-independent speaker verification,”
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
O. M. Parkhi, A. Vedaldi, and A. Zisserman, “Deep face recognition,” in
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,”
2015
Cited alongside, same era.
J. H. Hansen and T. Hasan, “Speaker recognition by machines and humans: A tutorial review,”
2015
Cited alongside, same era.
S. H. Ghalehjegh and R. C. Rose, “Deep bottleneck features for i-vector based text-independent speaker verification,” in
2015
Cited alongside, same era.
I. Kemelmacher-Shlizerman, S. M. Seitz, D. Miller, and E. Brossard, “The megaface benchmark: 1 million faces for recognition at scale,” in
2016
Cited alongside, same era.
Y. Guo, L. Zhang, Y. Hu, X. He, and J. Gao, “MS-Celeb-1M: A dataset and benchmark for large-scale face recognition,” in
2016
Cited alongside, same era.
M. McLaren, L. Ferrer, D. Castan, and A. Lawson, “The speakers in the wild (SITW) speaker recognition database,” in
2016
Cited alongside, same era.
2017
Later among the works it cites.
J. S. Chung and A. Zisserman, “Lip reading in profile,” in
2017
Later among the works it cites.
A. Hermans, L. Beyer, and B. Leibe, “In defense of the triplet loss for person re-identification,”
2017
Later among the works it cites.
2018
Closest in time.
2018
Closest in time.
A. Nagrani, S. Albanie, and A. Zisserman, “Seeing voices and hearing faces: Cross-modal biometric matching,” in
2018
Closest in time.
2018
Closest in time.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,”
2018
Closest in time.
J. S. Chung and A. Zisserman, “Learning to lip read words by watching videos,”
2018
Closest in time.