Fetching the paper…
Reading the bibliography…
Inspite the emerging importance of Speech Emotion Recognition (SER), the state-of-the-art accuracy is quite low and needs improvement to make commercial applications of SER viable.
1907
Earlier work this paper cites.
Y. LeCun, B. E. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. E. Hubbard, and L. D. Jackel, “Handwritten digit recognition with a back-propagation network,” in Advances in neural information processing systems , 1990, pp. 396–404
1990
Earlier work this paper cites.
P. Ekman, “An argument for basic emotions,” Cognition & emotion , vol. 6, no. 3-4, pp. 169–200, 1992
1992
Earlier work this paper cites.
I. Engberg and A. Hansen, “Documentation of the danish emotional speech database des,” Aalborg , 1996
1996
Earlier work this paper cites.
R. Caruana, “Multitask learning,” Machine learning , vol. 28, no. 1, pp. 41–75, 1997
1997
Earlier work this paper cites.
S. Lawrence, C. L. Giles, A. C. Tsoi, and A. D. Back, “Face recognition: A convolutional neural-network approach,” IEEE transactions on neural networks , vol. 8, no. 1, pp. 98–113, 1997
1997
Earlier work this paper cites.
J. Baxter, “A model of inductive bias learning,” Journal of Artificial Intelligence Research , vol. 12, pp. 149–198, 2000
2000
Earlier work this paper cites.
S. Ben-David and R. Schuller, “Exploiting task relatedness for multiple task learning,” in Learning Theory and Kernel Machines . Springer, 2003, pp. 567–580
2003
Earlier work this paper cites.
J. Posner, J. A. Russell, and B. S. Peterson, “The circumplex model of affect: An integrative approach to affective neuroscience, cognitive development, and psychopathology,” Development and psychopathology , vol. 17, no. 3, pp. 715–734, 2005
2005
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal et al. , “The ami meeting corpus: A pre-announcement,” in International workshop on machine learning for multimodal interaction . Springer, 2005, pp. 28–39
2005
Earlier work this paper cites.
F. Burkhardt, A. Paeschke, M. Rolfes, W. F. Sendlmeier, and B. Weiss, “A database of german emotional speech.” in Interspeech , vol. 5, 2005, pp. 1517–1520
2005
Earlier work this paper cites.
F. Burkhardt, J. Ajmera, R. Englert, J. Stegmann, and W. Burleson, “Detecting anger in automated voice portal dialogs,” in Ninth International Conference on Spoken Language Processing , 2006
2006
Earlier work this paper cites.
T. Vogt and E. André, “Improving automatic emotion recognition from speech via gender differentiation,” in Proc. Language Resources and Evaluation Conference (LREC 2006), Genoa , 2006
2006
Earlier work this paper cites.
R. Collobert and J. Weston, “A unified architecture for natural language processing: Deep neural networks with multitask learning,” in Proceedings of the 25th international conference on Machine learning . ACM, 2008, pp. 160–167
2008
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, no. 4, p. 335, 2008
2008
Earlier work this paper cites.
L. Fu, X. Mao, and L. Chen, “Speaker independent emotion recognition based on svm/hmms fusion system,” in 2008 International Conference on Audio, Language and Image Processing . IEEE, 2008, pp. 61–65
2008
Earlier work this paper cites.
A. Mill, J. Allik, A. Realo, and R. Valk, “Age-related differences in emotion recognition ability: A cross-sectional study.” Emotion , vol. 9, no. 5, p. 619, 2009
2009
Earlier work this paper cites.
C.-C. Lee, C. Busso, S. Lee, and S. S. Narayanan, “Modeling mutual influence of interlocutor emotion states in dyadic spoken interactions,” in Tenth Annual Conference of the International Speech Communication Association , 2009
2009
Earlier work this paper cites.
M. Wöllmer, A. Metallinou, F. Eyben, B. Schuller, and S. Narayanan, “Context-sensitive multimodal emotion recognition from speech and facial expression using bidirectional lstm modeling,” in Proc. INTERSPEECH 2010, Makuhari, Japan , 2010, pp. 2362–2365
2010
Earlier work this paper cites.
X. Li, Y.-Y. Wang, and G. Tur, “Multi-task learning for spoken language understanding with shared slots,” in Twelfth Annual Conference of the International Speech Communication Association , 2011
2011
Earlier work this paper cites.
L. S. Roberts, “A forensic phonetic study of the vocal responses of individuals in distress,” Ph.D. dissertation, University of York, 2012
2012
Earlier work this paper cites.
N. Ding, V. Sethu, J. Epps, and E. Ambikairajah, “Speaker variability in emotion recognition-an adaptation based approach,” in 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2012, pp. 5101–5104
2012
Earlier work this paper cites.
M. Wöllmer, A. Metallinou, N. Katsamanis, B. Schuller, and S. Narayanan, “Analyzing the memory of blstm neural networks for enhanced emotion classification in dyadic spoken interactions,” in 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2012, pp. 4157–4160
2012
Earlier work this paper cites.
X. Zhu and D. Ramanan, “Face detection, pose estimation, and landmark localization in the wild,” in Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on . IEEE, 2012, pp. 2879–2886
2012
Earlier work this paper cites.
F. Eyben, M. Wöllmer, and B. Schuller, “A multitask approach to continuous five-dimensional affect sensing in natural speech,” ACM Transactions on Interactive Intelligent Systems (TiiS) , vol. 2, no. 1, p. 6, 2012
2012
Earlier work this paper cites.
H. Larochelle, M. Mandel, R. Pascanu, and Y. Bengio, “Learning algorithms for the classification restricted boltzmann machine,” Journal of Machine Learning Research , vol. 13, no. Mar, pp. 643–669, 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
A. Metallinou, M. Wollmer, A. Katsamanis, F. Eyben, B. Schuller, and S. Narayanan, “Context-sensitive learning for enhanced audiovisual emotion classification,” IEEE Transactions on Affective Computing , vol. 3, no. 2, pp. 184–198, 2012
2012
Earlier work this paper cites.
P. Trower, B. Bryant, and M. Argyle, Social skills and mental health (psychology revivals) . Routledge, 2013
2013
Earlier work this paper cites.
S. E. Kahou, C. Pal, X. Bouthillier, P. Froumenty, Ç. Gülçehre, R. Memisevic, P. Vincent, A. Courville, Y. Bengio, R. C. Ferrari et al. , “Combining modality specific deep neural networks for emotion recognition in video,” in Proceedings of the 15th ACM on International conference on multimodal interaction . ACM, 2013, pp. 543–550
2013
Earlier work this paper cites.
C. Farabet, C. Couprie, L. Najman, and Y. LeCun, “Learning hierarchical features for scene labeling,” IEEE transactions on pattern analysis and machine intelligence , vol. 35, no. 8, pp. 1915–1929, 2013
2013
Earlier work this paper cites.
J. Deng, Z. Zhang, E. Marchi, and B. Schuller, “Sparse autoencoder-based feature transfer learning for speech emotion recognition,” in 2013 Humaine Association Conference on Affective Computing and Intelligent Interaction . IEEE, 2013, pp. 511–516
2013
Earlier work this paper cites.
S. Mariooryad and C. Busso, “Exploring cross-modality affective reactions for audiovisual emotion recognition,” IEEE Transactions on affective computing , vol. 4, no. 2, pp. 183–196, 2013
2013
Earlier work this paper cites.
P. Laukka, D. Neiberg, and H. A. Elfenbein, “Evidence for cultural dialects in vocal emotion expression: Acoustic classification within and across five nations.” Emotion , vol. 14, no. 3, p. 445, 2014
2014
Earlier work this paper cites.
Q. Mao, M. Dong, Z. Huang, and Y. Zhan, “Learning salient features for speech emotion recognition using convolutional neural networks,” IEEE transactions on multimedia , vol. 16, no. 8, pp. 2203–2213, 2014
2014
Earlier work this paper cites.
S. Li, Z.-Q. Liu, and A. B. Chan, “Heterogeneous multi-task learning for human pose estimation with deep convolutional neural network,” in Proceedings of the IEEE conference on computer vision and pattern recognition workshops , 2014, pp. 482–489
2014
Earlier work this paper cites.
Z. Zhang, P. Luo, C. C. Loy, and X. Tang, “Facial landmark detection by deep multi-task learning,” in European Conference on Computer Vision . Springer, 2014, pp. 94–108
2014
Earlier work this paper cites.
D. Chen, S. Ren, Y. Wei, X. Cao, and J. Sun, “Joint cascade face detection and alignment,” in European Conference on Computer Vision . Springer, 2014, pp. 109–122
2014
Cited alongside, same era.
Z. Huang, M. Dong, Q. Mao, and Y. Zhan, “Speech emotion recognition using cnn,” in Proceedings of the 22Nd ACM International Conference on Multimedia , 2014
2014
Cited alongside, same era.
P. Jackson and S. Haq, “Surrey audio-visual expressed emotion(savee) database,” University of Surrey: Guildford, UK , 2014
2014
Cited alongside, same era.
J. Deng, Z. Zhang, and B. Schuller, “Linked source and target domain subspace feature transfer learning–exemplified by speech emotion recognition,” in 2014 22nd International Conference on Pattern Recognition . IEEE, 2014, pp. 761–766
2014
Cited alongside, same era.
S. Latif, R. Rana, S. Younis, J. Qadir, and J. Epps, “Transfer learning for improving speech emotion classification accuracy,” in Proc. Interspeech 2018 , 2018, pp. 257–261. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2018-1625
2018
Later among the works it cites.
S. Latif, R. Rana, J. Qadir, and J. Epps, “Variational autoencoders for learning latent representations of speech emotion: A preliminary study,” in Proc. Interspeech 2018 , 2018, pp. 3107–3111. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2018-1568
2018
Later among the works it cites.
R. Lotfian and C. Busso, “Predicting categorical emotions by jointly learning primary and secondary emotions through multitask learning,” in Proc. Interspeech 2018 , 2018, pp. 951–955. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2018-2464
2018
Later among the works it cites.
X. Liu, J. van de Weijer, and A. D. Bagdanov, “Leveraging unlabeled data for crowd counting by learning to rank,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 7661–7669
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez-Dominguez, “Deep neural networks for small footprint text-dependent speaker verification,” in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2014, pp. 4052–4056
2014
Cited alongside, same era.
A. Rousseau, P. Deléglise, and Y. Esteve, “Enhancing the ted-lium corpus with selected data for language modeling and more ted talks.” in LREC , 2014, pp. 3935–3939
2014
Cited alongside, same era.
A. Makhzani, J. Shlens, N. Jaitly, I. Goodfellow, and B. Frey, “Adversarial autoencoders,” ICLR 2016 Workshop , 2015
2015
Cited alongside, same era.
X. Zhang, J. Zhao, and Y. LeCun, “Character-level convolutional networks for text classification,” in Advances in neural information processing systems , 2015, pp. 649–657
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Cited alongside, same era.
2015
Cited alongside, same era.
G. Trigeorgis, F. Ringeval, R. Brueckner, E. Marchi, M. A. Nicolaou, B. Schuller, and S. Zafeiriou, “Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network,” in 2016 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2016, pp. 5200–5204
2016
Cited alongside, same era.
2018
Later among the works it cites.
2018
Later among the works it cites.
F. Ma, W. Gu, W. Zhang, S. Ni, S.-L. Huang, and L. Zhang, “Speech emotion recognition via attention-based dnn from multi-task learning,” in Proceedings of the 16th ACM Conference on Embedded Networked Sensor Systems . ACM, 2018, pp. 363–364
2018
Later among the works it cites.
F. Tao and G. Liu, “Advanced lstm: A study about better time dependency modeling in emotion recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 2906–2910
2018
Later among the works it cites.
Z. Zhang, J. Han, J. Deng, X. Xu, F. Ringeval, and B. Schuller, “Leveraging unlabeled data for emotion recognition with enhanced collaborative semi-supervised learning,” IEEE Access , vol. 6, pp. 22 196–22 209, 2018
2018
Later among the works it cites.
J. Deng, X. Xu, Z. Zhang, S. Frühholz, and B. Schuller, “Semisupervised autoencoders for speech emotion recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 1, pp. 31–43, 2018
2018
Later among the works it cites.
J. Huang, Y. Li, J. Tao, Z. Lian, M. Niu, and J. Yi, “Speech emotion recognition using semi-supervised learning with ladder networks,” in 2018 First Asian Conference on Affective Computing and Intelligent Interaction (ACII Asia) . IEEE, 2018, pp. 1–5
2018
Later among the works it cites.
S. Parthasarathy and C. Busso, “Ladder networks for emotion recognition: Using unsupervised auxiliary tasks to improve predictions of emotional attributes,” Proc. Interspeech 2018 , pp. 3698–3702, 2018
2018
Later among the works it cites.
A. Makhzani, “Unsupervised representation learning with autoencoders,” Ph.D. dissertation, 2018
2018
Later among the works it cites.
W. Xiong, L. Wu, F. Alleva, J. Droppo, X. Huang, and A. Stolcke, “The microsoft 2017 conversational speech recognition system,” in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2018, pp. 5934–5938
2018
Later among the works it cites.
J. Liu, W. Han, H. Ruan, X. Chen, D. Jiang, and H. Li, “Learning salient features for speech emotion recognition using cnn,” in 2018 First Asian Conference on Affective Computing and Intelligent Interaction (ACII Asia) . IEEE, 2018, pp. 1–5
2018
Later among the works it cites.
C. Etienne, G. Fidanza, A. Petrovskii, L. Devillers, and B. Schmauch, “Cnn+ lstm architecture for speech emotion recognition with data augmentation,” in Proc. Workshop on Speech, Music and Mind 2018 , 2018, pp. 21–25
2018
Later among the works it cites.
S. Latif, R. Rana, J. Qadir, and J. Epps, “Variational autoencoders for learning latent representations of speech emotion: A preliminary study,” Proc. Interspeech 2018 , pp. 3107–3111, 2018
2018
Later among the works it cites.
A. Bérard, L. Besacier, A. C. Kocabiyikoglu, and O. Pietquin, “End-to-end automatic speech translation of audiobooks,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 6224–6228
2018
Later among the works it cites.
2018
Later among the works it cites.
P. Yenigalla, A. Kumar, S. Tripathi, C. Singh, S. Kar, and J. Vepa, “Speech emotion recognition using spectrogram & phoneme embedding.” in Interspeech , 2018, pp. 3688–3692
2018
Later among the works it cites.
M. Ravanelli and Y. Bengio, “Speaker recognition from raw waveform with sincnet,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 1021–1028
2018
Later among the works it cites.
S. E. Eskimez, Z. Duan, and W. Heinzelman, “Unsupervised learning approach to feature analysis for automatic speech emotion recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5099–5103
2018
Later among the works it cites.
X. Ma, Z. Wu, J. Jia, M. Xu, H. Meng, and L. Cai, “Emotion recognition from variable-length speech segments using deep learning on spectrograms,” Proc. Interspeech 2018 , pp. 3683–3687, 2018
2018
Later among the works it cites.
S. Sahu, R. Gupta, and C. Espy-Wilson, “On enhancing speech emotion recognition using generative adversarial networks,” Proc. Interspeech 2018 , pp. 3693–3697, 2018
2018
Later among the works it cites.
Z. Huang, J. Epps, and D. Joachim, “Speech landmark bigrams for depression detection from naturalistic smartphone speech,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5856–5860
2019
Closest in time.
R. Rana, S. Latif, R. Gururajan, A. Gray, G. Mackenzie, G. Humphris, and J. Dunn, “Automated screening for distress: A perspective for the future,” European journal of cancer care , p. e13033, 2019
2019
Closest in time.
2019
Closest in time.
Z. Zhang, B. Wu, and B. Schuller, “Attention-augmented end-to-end multi-task learning for emotion prediction from speech,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6705–6709
2019
Closest in time.
X. Liu, J. Van De Weijer, and A. D. Bagdanov, “Exploiting unlabeled data in cnns by self-supervised learning to rank,” IEEE transactions on pattern analysis and machine intelligence , 2019
2019
Closest in time.
R. Ranjan, V. M. Patel, and R. Chellappa, “Hyperface: A deep multi-task learning framework for face detection, landmark localization, pose estimation, and gender recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 41, no. 1, pp. 121–135, 2019
2019
Closest in time.
J.-H. Tao, J. Huang, Y. Li, Z. Lian, and M.-Y. Niu, “Semi-supervised ladder networks for speech emotion recognition,” International Journal of Automation and Computing , 2019
2019
Closest in time.
2019
Closest in time.
J. Zhao, X. Mao, and L. Chen, “Speech emotion recognition using deep 1d & 2d cnn lstm networks,” Biomedical Signal Processing and Control , vol. 47, pp. 312–323, 2019
2019
Closest in time.
H. Dubey, A. Sangwan, and J. H. Hansen, “Transfer learning using raw waveform sincnet for robust speaker diarization,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6296–6300
2019
Closest in time.
S. Latif, J. Qadir, and M. Bilal, “Unsupervised adversarial domain adaptation for cross-lingual speech emotion recognition,” in 2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII) . IEEE, 2019, pp. 732–737
2019
Closest in time.
S. Parthasarathy, V. Rozgic, M. Sun, and C. Wang, “Improving emotion classification through variational inference of latent variables,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 7410–7414
2019
Closest in time.
M. Neumann and N. T. Vu, “Improving speech emotion recognition with unsupervised representation learning on unlabeled speech,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 7390–7394
2019
Closest in time.
F. Bao, M. Neumann, and N. T. Vu, “Cyclegan-based emotion style transfer as data augmentation for speech emotion recognition,” Manuscript submitted for publication , pp. 35–37, 2019
2019
Closest in time.