Fetching the paper…
Reading the bibliography…
Learning speaker-specific features is vital in many applications like speaker recognition, diarization and speech recognition.
S. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,” IEEE transactions on acoustics, speech, and signal processing , vol. 28, no. 4, pp. 357–366, 1980
1980
Earlier work this paper cites.
D. O’Shaughnessy, “Linear predictive coding,” IEEE potentials , vol. 7, no. 1, pp. 29–32, 1988
1988
Earlier work this paper cites.
H. Hermansky, “Perceptual linear predictive (plp) analysis of speech,” the Journal of the Acoustical Society of America , vol. 87, no. 4, pp. 1738–1752, 1990
1990
Earlier work this paper cites.
L. R. Rabiner and B.-H. Juang, “Fundamentals of speech recognition,” 1993
1993
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, N. L. Dahlgren, and V. Zue, “Timit acoustic-phonetic continuous speech corpus ldc93s1,” 1993. [Online]. Available: https://catalog.ldc.upenn.edu/LDC93S1
1993
Earlier work this paper cites.
J. Bromley, I. Guyon, Y. LeCun, E. Säckinger, and R. Shah, “Signature verification using a” siamese” time delay neural network,” in Advances in Neural Information Processing Systems , 1994, pp. 737–744
1994
Earlier work this paper cites.
J. P. Campbell, “Speaker recognition: A tutorial,” Proceedings of the IEEE , vol. 85, no. 9, pp. 1437–1462, 1997
1997
Earlier work this paper cites.
D. A. Reynolds, T. F. Quatieri, and R. B. Dunn, “Speaker verification using adapted gaussian mixture models,” Digital signal processing , vol. 10, no. 1-3, pp. 19–41, 2000
2000
Earlier work this paper cites.
X. Huang, A. Acero, and H.-W. Hon, Spoken Language Processing: A Guide to Theory, Algorithm, and System Development , 1st ed. Upper Saddle River, NJ, USA: Prentice Hall PTR, 2001
2001
Earlier work this paper cites.
I. Lapidot, H. Guterman, and A. Cohen, “Unsupervised speaker recognition based on competition between self-organizing maps,” IEEE Transactions on Neural Networks , vol. 13, no. 4, pp. 877–887, 2002
2002
Earlier work this paper cites.
K. Chen, “Towards better making a decision in speaker verification,” Pattern Recognition , vol. 36, no. 2, pp. 329–346, 2003
2003
Earlier work this paper cites.
P. Kenny, M. Mihoubi, and P. Dumouchel, “New map estimators for speaker recognition.” in INTERSPEECH , 2003
2003
Earlier work this paper cites.
P. Kenny and P. Dumouchel, “Disentangling speaker and channel effects in speaker verification,” in Acoustics, Speech, and Signal Processing, 2004. Proceedings.(ICASSP’04). IEEE International Conference on , vol. 1. IEEE, 2004, pp. I–37
2004
Earlier work this paper cites.
E. Shriberg, L. Ferrer, S. Kajarekar, A. Venkataraman, and A. Stolcke, “Modeling prosodic feature sequences for speaker recognition,” Speech Communication , vol. 46, no. 3, pp. 455–472, 2005
2005
Earlier work this paper cites.
S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” in Computer Vision and Pattern Recognition, 2005. CVPR 2005. IEEE Computer Society Conference on , vol. 1. IEEE, 2005, pp. 539–546
2005
Earlier work this paper cites.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” science , vol. 313, no. 5786, pp. 504–507, 2006
2006
Earlier work this paper cites.
R. Hadsell, S. Chopra, and Y. LeCun, “Dimensionality reduction by learning an invariant mapping,” in Computer vision and pattern recognition, 2006 IEEE computer society conference on , vol. 2. IEEE, 2006, pp. 1735–1742
2006
Earlier work this paper cites.
P. Kenny, G. Boulianne, P. Ouellet, and P. Dumouchel, “Joint factor analysis versus eigenchannels in speaker recognition,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 4, pp. 1435–1447, 2007
2007
Earlier work this paper cites.
M. Kotti, V. Moschou, and C. Kotropoulos, “Speaker segmentation and clustering,” Signal processing , vol. 88, no. 5, pp. 1091–1124, 2008
2008
Earlier work this paper cites.
2008
Earlier work this paper cites.
N. Dehak, R. Dehak, P. Kenny, N. Brümmer, P. Ouellet, and P. Dumouchel, “Support vector machines versus fast scoring in the low-dimensional total variability space for speaker verification,” in Tenth Annual conference of the international speech communication association , 2009
2009
Earlier work this paper cites.
H. Lee, P. Pham, Y. Largman, and A. Y. Ng, “Unsupervised feature learning for audio classification using convolutional deep belief networks,” in Advances in neural information processing systems , 2009, pp. 1096–1104
2009
Cited alongside, same era.
S. J. Pan, Q. Yang et al. , “A survey on transfer learning,” IEEE Transactions on knowledge and data engineering , vol. 22, no. 10, pp. 1345–1359, 2010
2010
Cited alongside, same era.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 4, pp. 788–798, 2011
2011
Cited alongside, same era.
A. Kanagasundaram, R. Vogt, D. B. Dean, S. Sridharan, and M. W. Mason, “I-vector based speaker recognition on short utterances,” in Proceedings of the 12th Annual Conference of the International Speech Communication Association . International Speech Communication Association (ISCA), 2011, pp. 2341–2344
2011
M. McLaren, Y. Lei, N. Scheffer, and L. Ferrer, “Application of convolutional neural networks to speaker recognition in noisy conditions,” in Fifteenth Annual Conference of the International Speech Communication Association , 2014
2014
Later among the works it cites.
S.-Y. Chang and N. Morgan, “Robust cnn-based speech recognition with gabor filter kernels,” in Fifteenth Annual Conference of the International Speech Communication Association , 2014
2014
Later among the works it cites.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting.” Journal of Machine Learning Research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Later among the works it cites.
J. H. Hansen and T. Hasan, “Speaker recognition by machines and humans: A tutorial review,” IEEE Signal processing magazine , vol. 32, no. 6, pp. 74–99, 2015
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
K. Chen and A. Salman, “Learning speaker-specific characteristics with a deep neural architecture,” IEEE Transactions on Neural Networks , vol. 22, no. 11, pp. 1744–1756, 2011
2011
Cited alongside, same era.
——, “Extracting speaker-specific information with a regularized siamese deep network,” in Advances in Neural Information Processing Systems , 2011, pp. 298–306
2011
Cited alongside, same era.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The kaldi speech recognition toolkit,” in IEEE 2011 Workshop on Automatic Speech Recognition and Understanding . IEEE Signal Processing Society, Dec. 2011, iEEE Catalog No.: CFP11SRW-USB
2011
Cited alongside, same era.
X. Anguera, S. Bozonnet, N. Evans, C. Fredouille, G. Friedland, and O. Vinyals, “Speaker diarization: A review of recent research,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 2, pp. 356–370, 2012
2012
Cited alongside, same era.
G. Dupuy, M. Rouvier, S. Meignier, and Y. Esteve, “I-vectors and ILP clustering adapted to cross-show speaker diarization,” in Thirteenth Annual Conference of the International Speech Communication Association , 2012
2012
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Cited alongside, same era.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Cited alongside, same era.
A. Rousseau, P. Deléglise, and Y. Esteve, “Ted-lium: an automatic speech recognition dedicated corpus.” in LREC , 2012, pp. 125–129
2012
Cited alongside, same era.
S. H. Ghalehjegh and R. C. Rose, “Deep bottleneck features for i-vector based text-independent speaker verification,” in Automatic Speech Recognition and Understanding (ASRU), 2015 IEEE Workshop on . IEEE, 2015, pp. 555–560
2015
Later among the works it cites.
2015
Later among the works it cites.
G. Koch, R. Zemel, and R. Salakhutdinov, “Siamese neural networks for one-shot image recognition,” in ICML Deep Learning Workshop , vol. 2, 2015
2015
Later among the works it cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, F. Bach and D. Blei, Eds., vol. 37. Lille, France: PMLR, 07–09 Jul 2015, pp. 448–456. [Online]. Available: http://proceedings.mlr.press/v37/ioffe15.html
2015
Later among the works it cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . MIT Press, 2016, http://www.deeplearningbook.org
2016
Later among the works it cites.
D. Snyder, P. Ghahremani, D. Povey, D. Garcia-Romero, Y. Carmiel, and S. Khudanpur, “Deep neural network-based speaker embeddings for end-to-end speaker verification,” in Spoken Language Technology Workshop (SLT), 2016 IEEE . IEEE, 2016, pp. 165–170
2016
Later among the works it cites.
M. M. Saleem and J. H. Hansen, “A discriminative unsupervised method for speaker recognition using deep learning,” in Machine Learning for Signal Processing (MLSP), 2016 IEEE 26th International Workshop on . IEEE, 2016, pp. 1–5
2016
Later among the works it cites.
Y. Lukic, C. Vogt, O. Dürr, and T. Stadelmann, “Speaker identification and clustering using convolutional neural networks,” in Machine Learning for Signal Processing (MLSP), 2016 IEEE 26th International Workshop on . IEEE, 2016, pp. 1–6
2016
Later among the works it cites.
2016
Later among the works it cites.
2017
Later among the works it cites.
H. Li, B. Baucom, and P. Georgiou, “Unsupervised latent behavior manifold learning from acoustic features: Audio2behavior,” in Proceedings of IEEE International Conference on Audio, Speech and Signal Processing (ICASSP) , New Orleans, Louisiana, March 2017
2017
Later among the works it cites.
A. Jati and P. Georgiou, “Speaker2vec: Unsupervised learning and adaptation of a speaker manifold using deep neural networks with an evaluation on speaker segmentation,” in Proceedings of Interspeech , August 2017
2017
Later among the works it cites.
D. Snyder, D. Garcia-Romero, D. Povey, and S. Khudanpur, “Deep neural network embeddings for text-independent speaker verification,” Proc. Interspeech 2017 , pp. 999–1003, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” Submitted to ICASSP , 2018
2018
Closest in time.
M. Rouvier, P.-M. Bousquet, and B. Favre, “Speaker diarization through speaker embeddings,” in Signal Processing Conference (EUSIPCO), 2015 23rd European . IEEE, 2015, pp. 2082–2086
2086
Closest in time.