Fetching the paper…
Reading the bibliography…
Deep learning is progressively gaining popularity as a viable alternative to i-vectors for speaker recognition.
“DARPA TIMIT Acoustic Phonetic Continuous Speech Corpus CDROM,” 1993
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, and N. L. Dahlgren, · 1993
Earlier work this paper cites.
“Speaker verification using adapted Gaussian mixture models,”
D. A. Reynolds, T. F. Quatieri, and R. B. Dunn, · 2000
Earlier work this paper cites.
Digital Signal Processing
S. K. Mitra, · 2005
Earlier work this paper cites.
“Understanding the difficulty of training deep feedforward neural networks,”
X. Glorot and Y. Bengio, · 2010
Earlier work this paper cites.
Fundamentals of Speaker Recognition
H. Beigi, · 2011
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, · 2011
Earlier work this paper cites.
Theory and Applications of Digital Speech Processing
L. R. Rabiner and R. W. Schafer, · 2011
Earlier work this paper cites.
“The Kaldi Speech Recognition Toolkit,”
D. Povey et al., · 2011
Earlier work this paper cites.
“i-vector based speaker recognition on short utterances,”
A. Kanagasundaram, R. Vogt, D. Dean, S. Sridharan, and M. Mason, · 2011
Earlier work this paper cites.
“Context-dependent pre-trained deep neural networks for large vocabulary speech recognition,”
G. Dahl, D. Yu, L. Deng, and A. Acero, · 2012
Earlier work this paper cites.
“Bottleneck features for speaker recognition,”
S. Yaman, J. W. Pelecanos, and R. Sarikaya, · 2012
Earlier work this paper cites.
“Study of the effect of i-vector modeling on short and mismatch utterance duration for speaker verification,”
A. K. Sarkar, D Matrouf, P.M. Bousquet, and J.F. Bonastre, · 2012
Earlier work this paper cites.
“Learning filter banks within a deep neural network framework,”
T. N. Sainath, B. Kingsbury, A. R. Mohamed, and B. Ramabhadran, · 2013
Earlier work this paper cites.
“Rectifier nonlinearities improve neural network acoustic models,”
A. L. Maas, A. Y. Hannun, and A. Y. Ng, · 2013
Earlier work this paper cites.
“Deep neural networks for extracting baum-welch statistics for speaker recognition,”
P. Kenny, V. Gupta, T. Stafylakis, P. Ouellet, and J. Alam, · 2014
Earlier work this paper cites.
“Deep neural networks for small footprint text-dependent speaker verification,”
E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez-Dominguez, · 2014
Earlier work this paper cites.
“Acoustic modeling with deep neural networks using raw time signal for LVCSR,”
Z. Tüske, P. Golik, R. Schlüter, and H. Ney, · 2014
Earlier work this paper cites.
“Empirical evaluation of gated recurrent neural networks on sequence modeling,”
J. Chung, Ç. Gülçehre, K. Cho, and Y. Bengio, · 2014
Earlier work this paper cites.
“Modified-prior i-Vector Estimation for Language Identification of Short Duration Utterances,”
R. Travadi, M. Van Segbroeck, and S. Narayanan, · 2014
Cited alongside, same era.
Automatic Speech Recognition - A Deep Learning Approach
D. Yu and L. Deng, · 2015
Cited alongside, same era.
“Advances in deep neural network approaches to speaker recognition,”
M. McLaren, Y. Lei, and L. Ferrer, · 2015
Cited alongside, same era.
“Deep neural network approaches to speaker and language recognition,”
F. Richardson, D. Reynolds, and N. Dehak, · 2015
Cited alongside, same era.
“A unified deep neural network for speaker and language recognition,”
F. Richardson, D. A. Reynolds, and N. Dehak, · 2015
Cited alongside, same era.
“Analysis of CNN-based speech recognition system using raw speech as input,”
D. Palaz, M. Magimai-Doss, and R. Collobert, · 2015
J. Ba, R. Kiros, and G. E. Hinton, · 2016
Later among the works it cites.
“An extensible speaker identification sidekit in python,”
A. Larcher, K. A. Lee, and S. Meignier, · 2016
Later among the works it cites.
Deep learning for Distant Speech Recognition
M. Ravanelli, · 2017
Later among the works it cites.
“A network of deep neural networks for distant speech recognition,”
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio, · 2017
Later among the works it cites.
“Deep neural network embeddings for text-independent speaker verification,”
D. Snyder, D. Garcia-Romero, D. Povey, and S. Khudanpur, · 2017
Later among the works it cites.
“Deep speaker embeddings for short-duration speaker verification,”
G. Bhattacharya, J. Alam, and P. Kenny, · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Learning the speech front-end with raw waveform CLDNNs,”
T. N. Sainath, R. J. Weiss, A. W. Senior, K. W. Wilson, and O. Vinyals, · 2015
Cited alongside, same era.
“Speech acoustic modeling from raw multichannel waveforms,”
Y. Hoshen, R. Weiss, and K. W. Wilson, · 2015
Cited alongside, same era.
“Speaker localization and microphone spacing invariant acoustic modeling from raw multichannel waveforms,”
T. N. Sainath, R. J. Weiss, K. W. Wilson, A. Narayanan, M. Bacchiani, and A. Senior, · 2015
Cited alongside, same era.
“Librispeech: An ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Cited alongside, same era.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
S. Ioffe and C. Szegedy, · 2015
Cited alongside, same era.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville, · 2016
Cited alongside, same era.
Later among the works it cites.
“Voxceleb: a large-scale speaker identification dataset,”
A. Nagrani, J. S. Chung, and A. Zisserman, · 2017
Later among the works it cites.
“End-to-end spoofing detection with raw waveform CLDNNS,”
H. Dinkel, N. Chen, Y. Qian, and K. Yu, · 2017
Later among the works it cites.
“Improving speech recognition by revising gated recurrent units,”
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio, · 2017
Later among the works it cites.
“DNN Filter Bank Cepstral Coefficients for Spoofing Detection,”
H. Yu, Z. H. Tan, Y. Zhang, Z. Ma, and J. Guo, · 2017
Later among the works it cites.
“A deep neural network integrated with filterbank learning for speech recognition,”
H. Seki, K. Yamamoto, and S. Nakagawa, · 2017
Later among the works it cites.
“X-vectors: Robust dnn embeddings for speaker recognition,”
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, · 2018
Closest in time.
“Text-independent speaker verification based on triplet convolutional neural network embeddings,”
C. Zhang, K. Koishida, and J. Hansen, · 2018
Closest in time.
“Towards directly modeling raw speech signal for speaker verification using CNNs,”
H. Muckenhirn, M. Magimai-Doss, and S. Marcel, · 2018
Closest in time.
“A complete end-to-end speaker verification system using deep neural networks: From raw signals to verification result,”
J.-W. Jung, H.-S. Heo, I.-H. Yang, H.-J. Shim, , and H.-J. Yu, · 2018
Closest in time.
“Avoiding Speaker Overfitting in End-to-End DNNs using Raw Waveform for Text-Independent Speaker Verification,”
J.-W. Jung, H.-S. Heo, I.-H. Yang, H.-J. Shim, and H.-J. Yu, · 2018
Closest in time.
“Light gated recurrent units for speech recognition,”
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio, · 2018
Closest in time.
“Twin regularization for online speech recognition,”
M. Ravanelli, D. Serdyuk, and Y. Bengio, · 2018
Closest in time.