Fetching the paper…
Reading the bibliography…
Research on speech processing has traditionally considered the task of designing hand-engineered acoustic features (feature engineering) as a separate distinct problem from the task of designing efficient machine learning (ML) models to make prediction and classification decisions.
K. Pearson, “Liii. on lines and planes of closest fit to systems of points in space,” The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science , vol. 2, no. 11, pp. 559–572, 1901
1901
Earlier work this paper cites.
R. A. Fisher, “The use of multiple measurements in taxonomic problems,” Annals of eugenics , vol. 7, no. 2, pp. 179–188, 1936
1936
Earlier work this paper cites.
S. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,” IEEE transactions on acoustics, speech, and signal processing , vol. 28, no. 4, pp. 357–366, 1980
1980
Earlier work this paper cites.
D. H. Ackley, G. E. Hinton, and T. J. Sejnowski, “A learning algorithm for boltzmann machines,” Cognitive science , vol. 9, no. 1, pp. 147–169, 1985
1985
Earlier work this paper cites.
M. Fisher William, “The darpa speech recognition research database: Specifications and status/william m. fisher, george r. doddington, kathleen m. goudie-marshall,” in Proceedings of DARPA Workshop on Speech Recognition , 1986, pp. 93–99
1986
Earlier work this paper cites.
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, “Phoneme recognition using time-delay neural networks,” IEEE transactions on acoustics, speech, and signal processing , vol. 37, no. 3, pp. 328–339, 1989
1989
Earlier work this paper cites.
K.-F. Lee and S. Mahajan, “Corrective and reinforcement learning for speaker-independent continuous speech recognition,” Computer Speech & Language , vol. 4, no. 3, pp. 231–245, 1990
1990
Earlier work this paper cites.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “Switchboard: Telephone speech corpus for research and development,” in [Proceedings] ICASSP-92: 1992 IEEE International Conference on Acoustics, Speech, and Signal Processing , vol. 1. IEEE, 1992, pp. 517–520
1992
Earlier work this paper cites.
D. B. Paul and J. M. Baker, “The design for the wall street journal-based csr corpus,” in Proceedings of the workshop on Speech and Natural Language . Association for Computational Linguistics, 1992, pp. 357–362
1992
Earlier work this paper cites.
S. Furui, “Speaker-independent isolated word recognition based on emphasized spectral dynamics,” in ICASSP’86. IEEE International Conference on Acoustics, Speech, and Signal Processing , vol. 11. IEEE, 1986, pp. 1991–1994
1994
Earlier work this paper cites.
G. E. Hinton and R. S. Zemel, “Autoencoders, minimum description length and helmholtz free energy,” in Advances in neural information processing systems , 1994, pp. 3–10
1994
Earlier work this paper cites.
P. C. Woodland, J. J. Odell, V. Valtchev, and S. J. Young, “Large vocabulary continuous speech recognition using htk.” in ICASSP (2) , 1994, pp. 125–128
1994
Earlier work this paper cites.
Y. LeCun, Y. Bengio et al. , “Convolutional networks for images, speech, and time series,” The handbook of brain theory and neural networks , vol. 3361, no. 10, p. 1995, 1995
1995
Earlier work this paper cites.
D. Svozil, V. Kvasnicka, and J. Pospichal, “Introduction to multi-layer feed-forward neural networks,” Chemometrics and intelligent laboratory systems , vol. 39, no. 1, pp. 43–62, 1997
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural networks,” IEEE Transactions on Signal Processing , vol. 45, no. 11, pp. 2673–2681, 1997
1997
Earlier work this paper cites.
B. Schölkopf, A. Smola, and K.-R. Müller, “Nonlinear component analysis as a kernel eigenvalue problem,” Neural computation , vol. 10, no. 5, pp. 1299–1319, 1998
1998
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, P. Haffner et al. , “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
R. Caruana, “Learning to learn, chapter multitask learning,” 1998
1998
Earlier work this paper cites.
D. D. Lee and H. S. Seung, “Learning the parts of objects by non-negative matrix factorization,” Nature , vol. 401, no. 6755, p. 788, 1999
1999
Earlier work this paper cites.
D. T. Tran, “Fuzzy approaches to speech and speaker recognition,” Ph.D. dissertation, university of Canberra, 2000
2000
Earlier work this paper cites.
A. Hyvärinen and E. Oja, “Independent component analysis: algorithms and applications,” Neural networks , vol. 13, no. 4-5, pp. 411–430, 2000
2000
Earlier work this paper cites.
G. Baudat and F. Anouar, “Generalized discriminant analysis using a kernel approach,” Neural computation , vol. 12, no. 10, pp. 2385–2404, 2000
2000
Earlier work this paper cites.
S. T. Roweis and L. K. Saul, “Nonlinear dimensionality reduction by locally linear embedding,” science , vol. 290, no. 5500, pp. 2323–2326, 2000
2000
Earlier work this paper cites.
J. B. Tenenbaum, V. De Silva, and J. C. Langford, “A global geometric framework for nonlinear dimensionality reduction,” science , vol. 290, no. 5500, pp. 2319–2323, 2000
2000
Earlier work this paper cites.
H. Purwins, B. Blankertz, and K. Obermayer, “A new method for tracking modulations in tonal music in audio data format,” in Proceedings of the IEEE-INNS-ENNS International Joint Conference on Neural Networks. IJCNN 2000. Neural Computing: New Challenges and Perspectives for the New Millennium , vol. 6. IEEE, 2000, pp. 270–275
2000
Earlier work this paper cites.
P. Dayan, L. F. Abbott, and L. Abbott, “Theoretical neuroscience: computational and mathematical modeling of neural systems,” 2001
2001
Earlier work this paper cites.
A. Lee, T. Kawahara, and K. Shikano, “Julius—an open source real-time large vocabulary recognition engine,” 2001
2001
Earlier work this paper cites.
I. Borg and P. Groenen, “Modern multidimensional scaling: Theory and applications,” Journal of Educational Measurement , vol. 40, no. 3, pp. 277–280, 2003
2003
Earlier work this paper cites.
G. E. Hinton and S. T. Roweis, “Stochastic neighbor embedding,” in Advances in neural information processing systems , 2003, pp. 857–864
2003
Earlier work this paper cites.
M. Dong and Z. Sun, “On human machine cooperative learning control,” in Proceedings of the 2003 IEEE International Symposium on Intelligent Control . IEEE, 2003, pp. 81–86
2003
Earlier work this paper cites.
P. Lamere, P. Kwok, E. Gouvea, B. Raj, R. Singh, W. Walker, M. Warmuth, and P. Wolf, “The cmu sphinx-4 speech recognition system,” in IEEE Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP 2003), Hong Kong , vol. 1, 2003, pp. 2–5
2003
Earlier work this paper cites.
D. R. Hardoon, S. Szedmak, and J. Shawe-Taylor, “Canonical correlation analysis: An overview with application to learning methods,” Neural computation , vol. 16, no. 12, pp. 2639–2664, 2004
2004
Earlier work this paper cites.
F. Lee, R. Scherer, R. Leeb, A. Schlögl, H. Bischof, and G. Pfurtscheller, Feature mapping using PCA, locally linear embedding and isometric feature mapping for EEG-based brain computer interface . Citeseer, 2004
2004
Earlier work this paper cites.
A. Kocsor and L. Tóth, “Kernel-based feature extraction with a speech technology application,” IEEE Transactions on Signal Processing , vol. 52, no. 8, pp. 2250–2263, 2004
2004
Earlier work this paper cites.
F. Bimbot, J.-F. Bonastre, C. Fredouille, G. Gravier, I. Magrin-Chagnolleau, S. Meignier, T. Merlin, J. Ortega-García, D. Petrovska-Delacrétaz, and D. A. Reynolds, “A tutorial on text-independent speaker verification,” EURASIP Journal on Advances in Signal Processing , vol. 2004, no. 4, p. 101962, 2004
2004
Earlier work this paper cites.
Y. Lu, F. Lu, S. Sehgal, S. Gupta, J. Du, C. H. Tham, P. Green, and V. Wan, “Multitask learning in connectionist speech recognition,” in Proceedings of the Australian International Conference on Speech Science and Technology , 2004
2004
Earlier work this paper cites.
D. Hakkani-Tur, G. Tur, M. Rahim, and G. Riccardi, “Unsupervised and active learning in automatic speech recognition for call classification,” in 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing , vol. 1. IEEE, 2004, pp. I–429
2004
Earlier work this paper cites.
F. Burkhardt, A. Paeschke, M. Rolfes, W. F. Sendlmeier, and B. Weiss, “A database of german emotional speech,” in Ninth European Conference on Speech Communication and Technology , 2005
2005
Earlier work this paper cites.
R. Xu and D. C. Wunsch, “Survey of clustering algorithms,” 2005
2005
Earlier work this paper cites.
L. Cayton, “Algorithms for manifold learning,” Univ. of California at San Diego Tech. Rep , vol. 12, no. 1-17, p. 1, 2005
2005
Earlier work this paper cites.
X. J. Zhu, “Semi-supervised learning literature survey,” University of Wisconsin-Madison Department of Computer Sciences, Tech. Rep., 2005
2005
Earlier work this paper cites.
J. Stadermann, W. Koska, and G. Rigoll, “Multi-task learning strategies for a recurrent neural net in a hybrid tied-posteriors acoustic model,” in Ninth European Conference on Speech Communication and Technology , 2005
2005
Earlier work this paper cites.
G. Riccardi and D. Hakkani-Tur, “Active learning: Theory and applications to automatic speech recognition,” IEEE transactions on speech and audio processing , vol. 13, no. 4, pp. 504–511, 2005
2005
Earlier work this paper cites.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” science , vol. 313, no. 5786, pp. 504–507, 2006
2006
Earlier work this paper cites.
A. Errity and J. McKenna, “An investigation of manifold learning for speech analysis,” in Ninth International Conference on Spoken Language Processing , 2006
2006
Earlier work this paper cites.
G. E. Hinton, S. Osindero, and Y.-W. Teh, “A fast learning algorithm for deep belief nets,” Neural computation , vol. 18, no. 7, pp. 1527–1554, 2006
2006
Earlier work this paper cites.
H. Cuayáhuitl, S. Renals, O. Lemon, and H. Shimodaira, “Reinforcement learning of dialogue strategies with hierarchical abstract machines,” in 2006 IEEE Spoken Language Technology Workshop . IEEE, 2006, pp. 182–185
2006
Earlier work this paper cites.
T. Takiguchi and Y. Ariki, “Pca-based speech enhancement for distorted speech recognition.” Journal of multimedia , vol. 2, no. 5, 2007
2007
Earlier work this paper cites.
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle, “Greedy layer-wise training of deep networks,” in Advances in neural information processing systems , 2007, pp. 153–160
2007
Earlier work this paper cites.
C. Poultney, S. Chopra, Y. L. Cun et al. , “Efficient learning of sparse representations with an energy-based model,” in Advances in neural information processing systems , 2007, pp. 1137–1144
2007
Earlier work this paper cites.
T. Hain, L. Burget, J. Dines, G. Garau, V. Wan, M. Karafi, J. Vepa, and M. Lincoln, “The ami system for the transcription of speech in meetings,” in 2007 IEEE International Conference on Acoustics, Speech and Signal Processing-ICASSP’07 , vol. 4. IEEE, 2007, pp. IV–357
2007
Earlier work this paper cites.
J. J. DiCarlo and D. D. Cox, “Untangling invariant object recognition,” Trends in cognitive sciences , vol. 11, no. 8, pp. 333–341, 2007
2007
Earlier work this paper cites.
R. Raina, A. Battle, H. Lee, B. Packer, and A. Y. Ng, “Self-taught learning: transfer learning from unlabeled data,” in Proceedings of the 24th international conference on Machine learning . ACM, 2007, pp. 759–766
2007
Earlier work this paper cites.
S. Asakawa, N. Minematsu, and K. Hirose, “Automatic recognition of connected vowels only using speaker-invariant representation of speech dynamics,” in Eighth Annual Conference of the International Speech Communication Association , 2007
2007
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing data using t-sne,” Journal of machine learning research , vol. 9, no. Nov, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
M. Gales, S. Young et al. , “The application of hidden markov models in speech recognition,” Foundations and Trends® in Signal Processing , vol. 1, no. 3, pp. 195–304, 2008
2008
Earlier work this paper cites.
M. Wöllmer, F. Eyben, S. Reiter, B. Schuller, C. Cox, E. Douglas-Cowie, and R. Cowie, “Abandoning emotion classes-towards continuous emotion recognition with modelling of long-range dependencies,” in Proc. 9th Interspeech 2008 incorp. 12th Australasian Int. Conf. on Speech Science and Technology SST 2008, Brisbane, Australia , 2008, pp. 597–600
2008
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, no. 4, p. 335, 2008
2008
Earlier work this paper cites.
A. Cutler, “The abstract representations in speech processing,” The Quarterly Journal of Experimental Psychology , vol. 61, no. 11, pp. 1601–1619, 2008
2008
Earlier work this paper cites.
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol, “Extracting and composing robust features with denoising autoencoders,” in Proceedings of the 25th international conference on Machine learning . ACM, 2008, pp. 1096–1103
2008
Earlier work this paper cites.
R. Collobert and J. Weston, “A unified architecture for natural language processing: Deep neural networks with multitask learning,” in Proceedings of the 25th international conference on Machine learning . ACM, 2008, pp. 160–167
2008
Earlier work this paper cites.
B. Schuller, S. Steidl, and A. Batliner, “The interspeech 2009 emotion challenge,” in Tenth Annual Conference of the International Speech Communication Association , 2009
2009
Earlier work this paper cites.
Y. Bengio et al. , “Learning deep architectures for ai,” Foundations and trends® in Machine Learning , vol. 2, no. 1, pp. 1–127, 2009
2009
Earlier work this paper cites.
H. Lee, P. Pham, Y. Largman, and A. Y. Ng, “Unsupervised feature learning for audio classification using convolutional deep belief networks,” in Advances in neural information processing systems , 2009, pp. 1096–1104
2009
Earlier work this paper cites.
A.-r. Mohamed, G. Dahl, and G. Hinton, “Deep belief networks for phone recognition,” in Nips workshop on deep learning for speech recognition and related applications , vol. 1, no. 9. Vancouver, Canada, 2009, p. 39
2009
Earlier work this paper cites.
F. Eyben, M. Wöllmer, and B. Schuller, “Openear—introducing the munich open-source emotion and affect recognition toolkit,” in 2009 3rd international conference on affective computing and intelligent interaction and workshops . IEEE, 2009, pp. 1–6
2009
Earlier work this paper cites.
M. Wöllmer, Y. Sun, F. Eyben, and B. Schuller, “Long short-term memory networks for noise robust speech recognition,” in Proc. INTERSPEECH 2010, Makuhari, Japan , 2010, pp. 2966–2969
2010
Earlier work this paper cites.
G. Dahl, A.-r. Mohamed, G. E. Hinton et al. , “Phone recognition with the mean-covariance restricted boltzmann machine,” in Advances in neural information processing systems , 2010, pp. 469–477
2010
Earlier work this paper cites.
A.-r. Mohamed, D. Yu, and L. Deng, “Investigation of full-sequence training of deep belief networks for speech recognition,” in Eleventh Annual Conference of the International Speech Communication Association , 2010
2010
Earlier work this paper cites.
D. Yu, L. Deng, and G. Dahl, “Roles of pre-training and fine-tuning in context-dependent dbn-hmms for real-world speech recognition,” in Proc. NIPS Workshop on Deep Learning and Unsupervised Feature Learning , 2010
2010
Earlier work this paper cites.
F. Eyben, M. Wöllmer, and B. Schuller, “OpenSMILE – the Munich versatile and fast open-source audio feature extractor,” in Proceedings of the 18th ACM international conference on Multimedia . ACM, 2010, pp. 1459–1462
2010
Earlier work this paper cites.
N. Jaitly and G. Hinton, “Learning a better representation of speech soundwaves using restricted boltzmann machines,” in 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2011, pp. 5884–5887
2011
Earlier work this paper cites.
B. Schuller, A. Batliner, S. Steidl, and D. Seppi, “Recognising Realistic Emotions and Affect in Speech: State of the Art and Lessons Learnt from the First Challenge,” Speech Communication , vol. 53, no. 9/10, pp. 1062–1087, November/December 2011
2011
Earlier work this paper cites.
Y. Ma and Y. Fu, Manifold learning theory and applications . CRC press, 2011
2011
Earlier work this paper cites.
A. Ng et al. , “Sparse autoencoder,” CS294A Lecture notes , vol. 72, no. 2011, pp. 1–19, 2011
2011
Earlier work this paper cites.
S. Rifai, P. Vincent, X. Muller, X. Glorot, and Y. Bengio, “Contractive auto-encoders: Explicit invariance during feature extraction,” in Proceedings of the 28th International Conference on International Conference on Machine Learning . Omnipress, 2011, pp. 833–840
2011
Earlier work this paper cites.
A.-r. Mohamed, T. N. Sainath, G. E. Dahl, B. Ramabhadran, G. E. Hinton, M. A. Picheny et al. , “Deep belief networks using discriminative features for phone recognition.” in ICASSP , 2011, pp. 5060–5063
2011
Earlier work this paper cites.
D. Yu and M. L. Seltzer, “Improved bottleneck features using pretrained deep neural networks,” in Twelfth annual conference of the international speech communication association , 2011
2011
Earlier work this paper cites.
D. Hau and K. Chen, “Exploring hierarchical speech representations with a deep convolutional neural network,” UKCI 2011 Accepted Papers , p. 37, 2011
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The kaldi speech recognition toolkit,” in IEEE 2011 workshop on automatic speech recognition and understanding , no. CONF. IEEE Signal Processing Society, 2011
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal processing magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
T. Bänziger, M. Mortillaro, and K. R. Scherer, “Introducing the geneva multimodal expression corpus for experimental research on emotion perception.” Emotion , vol. 12, no. 5, p. 1161, 2012
2012
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Esteve, “Ted-lium: an automatic speech recognition dedicated corpus.” in LREC , 2012, pp. 125–129
2012
Earlier work this paper cites.
G. McKeown, M. Valstar, R. Cowie, M. Pantic, and M. Schroder, “The semaine database: Annotated multimodal records of emotionally colored conversations between a person and a limited agent,” IEEE Transactions on Affective Computing , vol. 3, no. 1, pp. 5–17, 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
S. Yaman, J. Pelecanos, and R. Sarikaya, “Bottleneck features for speaker recognition,” in Odyssey 2012-The Speaker and Language Recognition Workshop , 2012
2012
Earlier work this paper cites.
Y. Bengio, “Deep learning of representations for unsupervised and transfer learning,” in Proceedings of ICML workshop on unsupervised and transfer learning , 2012, pp. 17–36
2012
Earlier work this paper cites.
S. Thrun and L. Pratt, Learning to learn . Springer Science & Business Media, 2012
2012
Earlier work this paper cites.
D. Yu, F. Seide, G. Li, and L. Deng, “Exploiting sparseness in deep neural networks for large vocabulary speech recognition,” in 2012 IEEE International conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2012, pp. 4409–4412
2012
Earlier work this paper cites.
P. Swietojanski, A. Ghoshal, and S. Renals, “Unsupervised cross-lingual knowledge transfer in DNN-based LVCSR,” in 2012 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2012, pp. 246–251
2012
Earlier work this paper cites.
F. Eyben, M. Wöllmer, and B. Schuller, “A multitask approach to continuous five-dimensional affect sensing in natural speech,” ACM Transactions on Interactive Intelligent Systems (TiiS) , vol. 2, no. 1, p. 6, 2012
2012
Earlier work this paper cites.
Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,” IEEE transactions on pattern analysis and machine intelligence , vol. 35, no. 8, pp. 1798–1828, 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
F. Ringeval, A. Sonderegger, J. Sauer, and D. Lalanne, “Introducing the recola multimodal corpus of remote collaborative and affective interactions,” in 2013 10th IEEE international conference and workshops on automatic face and gesture recognition (FG) . IEEE, 2013, pp. 1–8
2013
Earlier work this paper cites.
A. Makhzani and B. Frey, “K-sparse autoencoders,” arXiv preprint arXiv:1312.5663 , 2013
2013
Earlier work this paper cites.
J. Deng, Z. Zhang, E. Marchi, and B. Schuller, “Sparse autoencoder-based feature transfer learning for speech emotion recognition,” in 2013 Humaine Association Conference on Affective Computing and Intelligent Interaction . IEEE, 2013, pp. 511–516
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
R. Xia and Y. Liu, “Using denoising autoencoder for emotion recognition.” in Interspeech , 2013, pp. 2886–2889
2013
Earlier work this paper cites.
S. Thomas, M. L. Seltzer, K. Church, and H. Hermansky, “Deep neural network features and semi-supervised training for low resource speech recognition,” in 2013 IEEE international conference on acoustics, speech and signal processing . IEEE, 2013, pp. 6704–6708
2013
Earlier work this paper cites.
J.-T. Huang, J. Li, D. Yu, L. Deng, and Y. Gong, “Cross-language knowledge transfer using multilingual deep neural network with shared hidden layers,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2013, pp. 7304–7308
2013
Earlier work this paper cites.
L. Deng, G. Hinton, and B. Kingsbury, “New types of deep neural network learning for speech recognition and related applications: An overview,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2013, pp. 8599–8603
2013
Earlier work this paper cites.
B. Ons, N. Tessema, J. Van De Loo, J. Gemmeke, G. De Pauw, W. Daelemans, and H. Van Hamme, “A self learning vocal interface for speech-impaired users,” in Proceedings of the Fourth Workshop on Speech and Language Processing for Assistive Technologies , 2013, pp. 73–81
2013
Earlier work this paper cites.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research , vol. 32, no. 11, pp. 1238–1274, 2013
2013
Earlier work this paper cites.
J. Huang and B. Kingsbury, “Audio-visual deep learning for noise robust speech recognition,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2013, pp. 7596–7599
2013
Earlier work this paper cites.
V. Vasilakakis, S. Cumani, P. Laface, and P. Torino, “Speaker recognition by means of deep belief networks,” Proc. Biometric Technologies in Forensic Science , pp. 52–57, 2013
2013
Earlier work this paper cites.
T. Yamada, L. Wang, and A. Kai, “Improvement of distant-talking speaker identification using bottleneck features of dnn.” in Interspeech , 2013, pp. 3661–3664
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
H. Lu, S. King, and O. Watts, “Combining a vector space representation of linguistic context with a deep neural network for text-to-speech synthesis,” in Eighth ISCA Workshop on Speech Synthesis , 2013
2013
Earlier work this paper cites.
T. Ishii, H. Komiyama, T. Shinozaki, Y. Horiuchi, and S. Kuroiwa, “Reverberant speech recognition based on denoising autoencoder.” in Interspeech , 2013, pp. 3512–3516
2013
Earlier work this paper cites.
B. Xia and C. Bao, “Speech enhancement with weighted denoising auto-encoder.” in INTERSPEECH , 2013, pp. 3444–3448
2013
Earlier work this paper cites.
N. E. Cibau, E. M. Albornoz, and H. L. Rufiner, “Speech emotion recognition using a deep autoencoder,” Anales de la XV Reunion de Procesamiento de la Informacion y Control , vol. 16, pp. 934–939, 2013
2013
Earlier work this paper cites.
M. L. Seltzer and J. Droppo, “Multi-task learning in deep neural networks for improved phoneme recognition,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2013, pp. 6965–6969
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
A. Larcher, J.-F. Bonastre, B. Fauve, K. Lee, C. Lévy, H. Li, J. Mason, and J.-Y. Parfait, “Alize 3.0-open source toolkit for state-of-the-art speaker recognition,” 2013
2013
Earlier work this paper cites.
M. A. Pathak, B. Raj, S. D. Rane, and P. Smaragdis, “Privacy-preserving speech processing: cryptographic and string-matching frameworks show promise,” IEEE signal processing magazine , vol. 30, no. 2, pp. 62–74, 2013
2013
Earlier work this paper cites.
X. Huang, J. Baker, and R. Reddy, “A historical perspective of speech recognition,” Commun. ACM , vol. 57, no. 1, pp. 94–103, 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in neural information processing systems , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
M. Längkvist, L. Karlsson, and A. Loutfi, “A review of unsupervised feature learning and deep learning for time-series modeling,” Pattern Recognition Letters , vol. 42, pp. 11–24, 2014
2014
Earlier work this paper cites.
E. Golchin and K. Maghooli, “Overview of manifold learning and its application in medical data set,” International journal of biomedical engineering and science (IJBES) , vol. 1, no. 2, pp. 23–33, 2014
2014
Earlier work this paper cites.
L. Deng, “A tutorial survey of architectures, algorithms, and applications for deep learning,” APSIPA Transactions on Signal and Information Processing , vol. 3, 2014
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Advances in neural information processing systems , 2014, pp. 3104–3112
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. Bengio, E. Laufer, G. Alain, and J. Yosinski, “Deep generative stochastic networks trainable by backprop,” in International Conference on Machine Learning , 2014, pp. 226–234
2014
Earlier work this paper cites.
X. Feng, Y. Zhang, and J. Glass, “Speech feature denoising and dereverberation via deep autoencoders for noisy reverberant speech recognition,” in 2014 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2014, pp. 1759–1763
2014
Earlier work this paper cites.
F. Weninger, S. Watanabe, Y. Tachioka, and B. Schuller, “Deep recurrent de-noising auto-encoder and blind de-reverberation for reverberated speech recognition,” in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2014, pp. 4623–4627
2014
Earlier work this paper cites.
R. Xia, J. Deng, B. Schuller, and Y. Liu, “Modeling gender information for emotion recognition using denoising autoencoder,” in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2014, pp. 990–994
2014
Earlier work this paper cites.
Z. Huang, M. Dong, Q. Mao, and Y. Zhan, “Speech emotion recognition using cnn,” in Proceedings of the 22Nd ACM International Conference on Multimedia , 2014
2014
Earlier work this paper cites.
Y. Tu, J. Du, Y. Xu, L. Dai, and C.-H. Lee, “Speech separation based on improved deep neural networks with dual outputs of speech features for both target and interfering speakers,” in The 9th International Symposium on Chinese Spoken Language Processing . IEEE, 2014, pp. 250–254
2014
Earlier work this paper cites.
K. M. Knill, M. J. Gales, A. Ragni, and S. P. Rath, “Language independent and unsupervised acoustic models for speech recognition and keyword spotting,” 2014
2014
Cited alongside, same era.
J. Deng, Z. Zhang, F. Eyben, and B. Schuller, “Autoencoder-based unsupervised domain adaptation for speech emotion recognition,” IEEE Signal Processing Letters , vol. 21, no. 9, pp. 1068–1072, 2014
2014
Cited alongside, same era.
R. Price, K.-i. Iso, and K. Shinoda, “Speaker adaptation of deep neural networks using a hierarchy of output layers,” in 2014 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2014, pp. 153–158
2014
Cited alongside, same era.
A. Narayanan and D. Wang, “Investigation of speech separation as a front-end for noise robust speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 22, no. 4, pp. 826–835, 2014
2014
Cited alongside, same era.
W. Dai, C. Dai, S. Qu, J. Li, and S. Das, “Very deep convolutional neural networks for raw waveforms,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 421–425
2017
Later among the works it cites.
A. van den Oord, O. Vinyals et al. , “Neural discrete representation learning,” in Advances in Neural Information Processing Systems , 2017, pp. 6306–6315
2017
Later among the works it cites.
J. Deng, S. Frühholz, Z. Zhang, and B. Schuller, “Recognizing emotions from whispered speech based on acoustic feature transfer learning,” IEEE Access , vol. 5, pp. 5235–5246, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Liu and K. Kirchhoff, “Graph-based semi-supervised acoustic modeling in dnn-based speech recognition,” in 2014 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2014, pp. 177–182
2014
Cited alongside, same era.
M. Gholamipoor and B. Nasersharif, “Feature mapping using deep belief networks for robust speech recognition,” The Modares Journal of Electrical Engineering , vol. 14, no. 3, pp. 24–30, 2014
2014
Cited alongside, same era.
O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, “Convolutional neural networks for speech recognition,” IEEE/ACM Transactions on audio, speech, and language processing , vol. 22, no. 10, pp. 1533–1545, 2014
2014
Cited alongside, same era.
W. M. Campbell, “Using deep belief networks for vector-based speaker recognition,” in Fifteenth Annual Conference of the International Speech Communication Association , 2014
2014
Cited alongside, same era.
O. Ghahabi and J. Hernando, “I-vector modeling with deep belief networks for multi-session speaker recognition,” network , vol. 20, p. 13, 2014
2014
Cited alongside, same era.
M. McLaren, Y. Lei, N. Scheffer, and L. Ferrer, “Application of convolutional neural networks to speaker recognition in noisy conditions,” in Fifteenth Annual Conference of the International Speech Communication Association , 2014
2014
Cited alongside, same era.
M. N. Stolar, M. Lech, and I. S. Burnett, “Optimized multi-channel deep neural network with 2D graphical representation of acoustic speech features for emotion recognition,” in 2014 8th International Conference on Signal Processing and Communication Systems (ICSPCS) . IEEE, 2014, pp. 1–6
2014
Cited alongside, same era.
Q. Mao, M. Dong, Z. Huang, and Y. Zhan, “Learning salient features for speech emotion recognition using convolutional neural networks,” IEEE transactions on multimedia , vol. 16, no. 8, pp. 2203–2213, 2014
2014
Cited alongside, same era.
Z. Tang, L. Li, D. Wang, R. Vipperla, Z. Tang, L. Li, D. Wang, and R. Vipperla, “Collaborative joint training with multitask recurrent model for speech and speaker recognition,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP) , vol. 25, no. 3, pp. 493–504, 2017
2017
Later among the works it cites.
Y. Zhang, Y. Liu, F. Weninger, and B. Schuller, “Multi-task deep neural network with shared hidden layers: Breaking down the wall between emotion representations,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2017, pp. 4990–4994
2017
Later among the works it cites.
D. Le, Z. Aldeneh, and E. M. Provost, “Discretized continuous speech emotion recognition with multi-task deep recurrent neural network,” Interspeech, 2017 (to apear) , 2017
2017
Later among the works it cites.
K. Jia, D. Tao, S. Gao, and X. Xu, “Improving training of deep neural networks via singular value bounding,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 4344–4352
2017
Later among the works it cites.
H. Li, B. Baucom, and P. Georgiou, “Unsupervised latent behavior manifold learning from acoustic features: Audio2behavior,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 5620–5624
2017
Later among the works it cites.
M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein gan,” arXiv preprint arXiv:1701.07875 , 2017
2017
Later among the works it cites.
M. Arjovsky and L. Bottou, “Towards principled methods for training generative adversarial networks. arxiv,” 2017
2017
Later among the works it cites.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers et al. , “In-datacenter performance analysis of a tensor processing unit,” in 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2017, pp. 1–12
2017
Later among the works it cites.
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2223–2232
2017
Later among the works it cites.
Z. Zhang, J. Geiger, J. Pohjalainen, A. E.-D. Mousa, W. Jin, and B. Schuller, “Deep learning for environmentally robust speech recognition: An overview of recent developments,” ACM Transactions on Intelligent Systems and Technology (TIST) , vol. 9, no. 5, p. 49, 2018
2018
Later among the works it cites.
M. Swain, A. Routray, and P. Kabisatpathy, “Databases, features and classifiers for speech emotion recognition: a review,” International Journal of Speech Technology , vol. 21, no. 1, pp. 93–120, 2018
2018
Later among the works it cites.
S. Latif, M. Usman, R. Rana, and J. Qadir, “Phonocardiographic sensing using deep learning for abnormal heartbeat detection,” IEEE Sensors Journal , vol. 18, no. 22, pp. 9393–9400, 2018
2018
Later among the works it cites.
A. Qayyum, S. Latif, and J. Qadir, “Quran reciter identification: A deep learning approach,” in 2018 7th International Conference on Computer and Communication Engineering (ICCCE) . IEEE, 2018, pp. 492–497
2018
Later among the works it cites.
2018
Later among the works it cites.
B. Milde and A. Köhn, “Open source automatic speech recognition for german,” in Speech Communication; 13th ITG-Symposium . VDE, 2018, pp. 1–5
2018
Later among the works it cites.
B. Mitra, N. Craswell et al. , “An introduction to neural information retrieval,” Foundations and Trends® in Information Retrieval , vol. 13, no. 1, pp. 1–126, 2018
2018
Later among the works it cites.
J. Kim, J. Urbano, C. C. Liem, and A. Hanjalic, “One deep music representation to rule them all? a comparative analysis of different representation learning strategies,” Neural Computing and Applications , pp. 1–27, 2018
2018
Later among the works it cites.
X. Liu, M. Cheng, H. Zhang, and C.-J. Hsieh, “Towards robust neural networks via random self-ensemble,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 369–385
2018
Later among the works it cites.
2018
Later among the works it cites.
E. Min, X. Guo, Q. Liu, G. Zhang, J. Cui, and J. Long, “A survey of clustering with deep learning: From the perspective of network architecture,” IEEE Access , vol. 6, pp. 39 501–39 514, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
F. Deng, J. Ren, and F. Chen, “Abstraction learning,” arXiv preprint arXiv:1809.03956 , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
D. Rethage, J. Pons, and X. Serra, “A wavenet for speech denoising,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5069–5073
2018
Later among the works it cites.
H. Ali, S. N. Tran, E. Benetos, and A. S. d. Garcez, “Speaker recognition with hybrid features from a deep belief network,” Neural Computing and Applications , vol. 29, no. 6, pp. 13–19, 2018
2018
Later among the works it cites.
H. Muckenhirn, M. M. Doss, and S. Marcell, “Towards directly modeling raw speech signal for speaker verification using cnns,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4884–4888
2018
Later among the works it cites.
P. Tzirakis, J. Zhang, and B. W. Schuller, “End-to-end speech emotion recognition using deep neural networks,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5089–5093
2018
Later among the works it cites.
M. Sarma, P. Ghahremani, D. Povey, N. K. Goel, K. K. Sarma, and N. Dehak, “Emotion identification from raw speech signals using dnns.” in Interspeech , 2018, pp. 3097–3101
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
S. Latif, A. Qayyum, M. Usman, and J. Qadir, “Cross lingual speech emotion recognition: Urdu vs. western languages,” in 2018 International Conference on Frontiers of Information Technology (FIT) . IEEE, 2018, pp. 88–93
2018
Later among the works it cites.
J. Huang, Y. Li, J. Tao, Z. Lian, M. Niu, and J. Yi, “Speech emotion recognition using semi-supervised learning with ladder networks,” in 2018 First Asian Conference on Affective Computing and Intelligent Interaction (ACII Asia) . IEEE, 2018, pp. 1–5
2018
Later among the works it cites.
——, “Ladder networks for emotion recognition: Using unsupervised auxiliary tasks to improve predictions of emotional attributes,” Proc. Interspeech 2018 , pp. 3698–3702, 2018
2018
Later among the works it cites.
J. Deng, X. Xu, Z. Zhang, S. Frühholz, and B. Schuller, “Semisupervised autoencoders for speech emotion recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 1, pp. 31–43, 2018
2018
Later among the works it cites.
S. Karita, S. Watanabe, T. Iwata, A. Ogawa, and M. Delcroix, “Semi-supervised end-to-end speech recognition.” in Interspeech , 2018, pp. 2–6
2018
Later among the works it cites.
W.-N. Hsu and J. Glass, “Extracting domain invariant features by unsupervised learning for robust automatic speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5614–5618
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
C.-X. Qin, D. Qu, and L.-H. Zhang, “Towards end-to-end speech recognition with transfer learning,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2018, no. 1, p. 18, 2018
2018
Later among the works it cites.
P. Denisov, N. T. Vu, and M. F. Font, “Unsupervised domain adaptation by adversarial learning for robust speech recognition,” in Speech Communication; 13th ITG-Symposium . VDE, 2018, pp. 1–5
2018
Later among the works it cites.
S. Sun, C.-F. Yeh, M.-Y. Hwang, M. Ostendorf, and L. Xie, “Domain adversarial training for accented speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4854–4858
2018
Later among the works it cites.
2018
Later among the works it cites.
Z. Meng, J. Li, Y. Gong, and B.-H. Juang, “Adversarial teacher-student learning for unsupervised domain adaptation,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5949–5953
2018
Later among the works it cites.
Z. Meng, J. Li, Z. Chen, Y. Zhao, V. Mazalov, Y. Gang, and B.-H. Juang, “Speaker-invariant training via adversarial learning,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5969–5973
2018
Later among the works it cites.
A. Tripathi, A. Mohan, S. Anand, and M. Singh, “Adversarial learning of raw speech features for domain invariant speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5959–5963
2018
Later among the works it cites.
Q. Wang, W. Rao, S. Sun, L. Xie, E. S. Chng, and H. Li, “Unsupervised domain adaptation via domain adversarial training for speaker recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4889–4893
2018
Later among the works it cites.
2018
Later among the works it cites.
R. Lotfian and C. Busso, “Predicting categorical emotions by jointly learning primary and secondary emotions through multitask learning,” in Proc. Interspeech 2018 , 2018, pp. 951–955. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2018-2464
2018
Later among the works it cites.
F. Tao and G. Liu, “Advanced LSTM: A study about better time dependency modeling in emotion recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 2906–2910
2018
Later among the works it cites.
2018
Later among the works it cites.
T. Zhang, M. Huang, and L. Zhao, “Learning structured representation for text classification via reinforcement learning,” in Thirty-Second AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
E. Lakomkin, M. A. Zamani, C. Weber, S. Magg, and S. Wermter, “Emorl: continuous acoustic emotion classification using deep reinforcement learning,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 1–6
2018
Later among the works it cites.
2018
Later among the works it cites.
Z. Zhang, J. Han, J. Deng, X. Xu, F. Ringeval, and B. Schuller, “Leveraging unlabeled data for emotion recognition with enhanced collaborative semi-supervised learning,” IEEE Access , vol. 6, pp. 22 196–22 209, 2018
2018
Later among the works it cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5329–5333
2018
Later among the works it cites.
S. Shon, H. Tang, and J. Glass, “Frame-level speaker embeddings for text-independent speaker recognition and analysis of end-to-end model,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 1007–1013
2018
Later among the works it cites.
E. Marchi, S. Shum, K. Hwang, S. Kajarekar, S. Sigtia, H. Richards, R. Haynes, Y. Kim, and J. Bridle, “Generalised discriminative transform via curriculum learning for speaker recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5324–5328
2018
Later among the works it cites.
J. Lorenzo-Trueba, G. E. Henter, S. Takaki, J. Yamagishi, Y. Morino, and Y. Ochiai, “Investigating different representations for modeling and controlling multiple emotions in dnn-based speech synthesis,” Speech Communication , vol. 99, pp. 135–143, 2018
2018
Later among the works it cites.
P. Li, Y. Song, I. V. McLoughlin, W. Guo, and L. Dai, “An attention pooling based representation learning method for speech emotion recognition.” in Interspeech , 2018, pp. 3087–3091
2018
Later among the works it cites.
Z. Zhao, Y. Zheng, Z. Zhang, H. Wang, Y. Zhao, and C. Li, “Exploring spatio-temporal representations by integrating attention-based bidirectional-lstm-rnns and fcns for speech emotion recognition,” 2018
2018
Later among the works it cites.
D. Luo, Y. Zou, and D. Huang, “Investigation on joint representation learning for robust feature extraction in speech emotion recognition.” in Interspeech , 2018, pp. 152–156
2018
Later among the works it cites.
S. H. Kabil, H. Muckenhirn, and M. Magimai-Doss, “On learning to identify genders from raw speech signal using cnns.” in Interspeech , 2018, pp. 287–291
2018
Later among the works it cites.
X. Liu, “Deep convolutional and LSTM neural networks for acoustic modelling in automatic speech recognition,” 2018
2018
Later among the works it cites.
N. Zeghidour, N. Usunier, I. Kokkinos, T. Schaiz, G. Synnaeve, and E. Dupoux, “Learning filterbanks from raw speech for phone recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5509–5513
2018
Later among the works it cites.
M. Ravanelli and Y. Bengio, “Speaker recognition from raw waveform with sincnet,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 1021–1028
2018
Later among the works it cites.
J.-W. Jung, H.-S. Heo, I.-H. Yang, H.-J. Shim, and H.-J. Yu, “Avoiding speaker overfitting in end-to-end dnns using raw waveform for text-independent speaker verification,” extraction , vol. 8, no. 12, pp. 23–24, 2018
2018
Later among the works it cites.
M. Ravanelli and Y. Bengio, “Speaker recognition from raw waveform with SincNet,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 1021–1028
2018
Later among the works it cites.
J.-W. Jung, H.-S. Heo, I.-H. Yang, H.-J. Shim, and H.-J. Yu, “A complete end-to-end speaker verification system using deep neural networks: From raw signals to verification result,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5349–5353
2018
Later among the works it cites.
2018
Later among the works it cites.
S. E. Eskimez, Z. Duan, and W. Heinzelman, “Unsupervised learning approach to feature analysis for automatic speech emotion recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5099–5103
2018
Later among the works it cites.
Y. Zhao, S. Takaki, H.-T. Luong, J. Yamagishi, D. Saito, and N. Minematsu, “Wasserstein GAN and waveform loss-based acoustic model training for multi-speaker text-to-speech synthesis systems using a wavenet vocoder,” IEEE Access , vol. 6, pp. 60 478–60 488, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
M. Neumann et al. , “Cross-lingual and multilingual speech emotion recognition on english and french,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5769–5773
2018
Later among the works it cites.
M. Abdelwahab and C. Busso, “Domain adversarial for acoustic emotion recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 12, pp. 2423–2435, 2018
2018
Later among the works it cites.
S. Yadav and A. Rai, “Learning discriminative features for speaker identification and verification.” in Interspeech , 2018, pp. 2237–2241
2018
Later among the works it cites.
F. Ma, W. Gu, W. Zhang, S. Ni, S.-L. Huang, and L. Zhang, “Speech emotion recognition via attention-based DNN from multi-task learning,” in Proceedings of the 16th ACM Conference on Embedded Networked Sensor Systems . ACM, 2018, pp. 363–364
2018
Later among the works it cites.
2018
Later among the works it cites.
N. Carlini and D. Wagner, “Audio adversarial examples: Targeted attacks on speech-to-text,” in 2018 IEEE Security and Privacy Workshops (SPW) . IEEE, 2018, pp. 1–7
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
N. Akhtar and A. Mian, “Threat of adversarial attacks on deep learning in computer vision: A survey,” IEEE Access , vol. 6, pp. 14 410–14 430, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
R. Chesney and D. K. Citron, “Deep fakes: a looming challenge for privacy, democracy, and national security,” 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
A. B. Nassif, I. Shahin, I. Attili, M. Azzeh, and K. Shaalan, “Speech recognition using deep neural networks: A systematic review,” IEEE Access , vol. 7, pp. 19 143–19 165, 2019
2019
Later among the works it cites.
J. Gómez-García, L. Moro-Velázquez, and J. I. Godino-Llorente, “On the design of automatic voice condition analysis systems. part ii: Review of speaker recognition techniques and study on the effects of different variability factors,” Biomedical Signal Processing and Control , vol. 48, pp. 128–143, 2019
2019
Later among the works it cites.
G. Zhong, X. Ling, and L.-N. Wang, “From shallow feature learning to deep learning: Benefits from the width and depth of deep architectures,” Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery , vol. 9, no. 1, p. e1255, 2019
2019
Later among the works it cites.
H. Purwins, B. Li, T. Virtanen, J. Schlüter, S.-Y. Chang, and T. Sainath, “Deep learning for audio signal processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 13, no. 2, pp. 206–219, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Liu, G. Keren, and B. W. Schuller, “N-HANS: Introducing the Augsburg Neuro-Holistic Audio-eNhancement System,” arxiv.org , no. 1911.07062, November 2019, 5 pages
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
M. Neumann and N. T. Vu, “Improving speech emotion recognition with unsupervised representation learning on unlabeled speech,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 7390–7394
2019
Later among the works it cites.
W.-N. Hsu, Y. Zhang, R. J. Weiss, Y.-A. Chung, Y. Wang, Y. Wu, and J. Glass, “Disentangling correlated speaker and noise for speech synthesis via data augmentation and adversarial factorization,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5901–5905
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
Z. Meng, J. Li, and Y. Gong, “Adversarial speaker adaptation,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5721–5725
2019
Later among the works it cites.
G. Bhattacharya, J. Monteiro, J. Alam, and P. Kenny, “Generative adversarial speaker embedding networks for domain robust end-to-end speaker verification,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6226–6230
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Han, Z. Zhang, Z. Ren, and B. W. Schuller, “Emobed: Strengthening monomodal emotion recognition via training with crossmodal emotion embeddings,” IEEE Transactions on Affective Computing , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
M. Sarma, P. Ghahremani, D. Povey, N. K. Goel, K. K. Sarma, and N. Dehak, “Improving emotion identification using phone posteriors in raw speech waveform based dnn,” Proc. Interspeech 2019 , pp. 3925–3929, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
Y.-A. Chung, Y. Wang, W.-N. Hsu, Y. Zhang, and R. Skerry-Ryan, “Semi-supervised training for improving data efficiency in end-to-end speech synthesis,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6940–6944
2019
Later among the works it cites.
2019
Later among the works it cites.
Z. Zhang, B. Wu, and B. Schuller, “Attention-augmented end-to-end multi-task learning for emotion prediction from speech,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6705–6709
2019
Later among the works it cites.
2019
Later among the works it cites.
F. Arute, K. Arya, R. Babbush, D. Bacon, J. C. Bardin, R. Barends, R. Biswas, S. Boixo, F. G. Brandao, D. A. Buell et al. , “Quantum supremacy using a programmable superconducting processor,” Nature , vol. 574, no. 7779, pp. 505–510, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
F. Bao, M. Neumann, and N. T. Vu, “Cyclegan-based emotion style transfer as data augmentation for speech emotion recognition,” Manuscript submitted for publication , pp. 35–37, 2019
2019
Later among the works it cites.
B. M. L. Srivastava, A. Bellet, M. Tommasi, and E. Vincent, “Privacy-preserving adversarial representation learning in asr: Reality or illusion?” Proc. INTERPSPEECH , pp. 3700–3704, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
K. Roth, A. Lucchi, S. Nowozin, and T. Hofmann, “Stabilizing training of generative adversarial networks through regularization,” in Advances in neural information processing systems , 2017, pp. 2018–2028
2028
Closest in time.
Y. Z. Işik, H. Erdogan, and R. Sarikaya, “S-vector: A discriminative representation derived from i-vector for speaker verification,” in 2015 23rd European Signal Processing Conference (EUSIPCO) . IEEE, 2015, pp. 2097–2101
2097
Closest in time.