Fetching the paper…
Reading the bibliography…
Speech emotion recognition (SER) systems find applications in various fields such as healthcare, education, and security and defense.
R. Caruana, “Multitask learning,” Machine Learning , vol. 28, no. 1, pp. 41–75, July 1997
1997
Earlier work this paper cites.
A. Blum and T. Mitchell, “Combining labeled and unlabeled data with co-training,” in Proceedings of the eleventh annual conference on Computational learning theory (COLT 1998) , Madison, WI, USA, July 1998, pp. 92–100
1998
Earlier work this paper cites.
I. Cohen, N. Sebe, F. Cozman, and T. Huang, “Semi-supervised learning for facial expression recognition,” in ACM SIGMM international workshop on Multimedia information retrieval (MIR 2003) , Berkeley, CA, USA, November 2003, pp. 17–22
2003
Earlier work this paper cites.
I. Cohen, N. Sebe, F. G. Gozman, M. C. Cirelo, and T. S. Huang, “Learning Bayesian network classifiers for facial expression recognition both labeled and unlabeled data,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2003) , Madison, WI, USA, June 2003, pp. 1–7
2003
Earlier work this paper cites.
D. Litman and K. Forbes-Riley, “Predicting student emotions in computer-human tutoring dialogues,” in ACM Association for Computational Linguistics (ACL 2004) , Barcelona, Spain, July 2004, pp. 1–8
2004
Earlier work this paper cites.
J. Liu, C. Chen, J. Bu, M. You, and J. Tao, “Speech emotion recognition using an enhanced co-training algorithm,” in IEEE International Conference on Multimedia and Expo (ICME 2007) , Beijing, China, July 2007, pp. 999–1002
2007
Earlier work this paper cites.
C. Clavel, I. Vasilescu, L. Devillers, G. Richard, and T. Ehrette, “Fear-type emotion recognition for future audio-based surveillance systems,” Speech Communication , vol. 50, no. 6, pp. 487–503, June 2008
2008
Earlier work this paper cites.
C. Busso and S. Narayanan, “Recording audio-visual emotional databases from actors: a closer look,” in Second International Workshop on Emotion: Corpora for Research on Emotion and Affect, International conference on Language Resources and Evaluation (LREC 2008) , Marrakech, Morocco, May 2008, pp. 17–22
2008
Earlier work this paper cites.
C. Busso, M. Bulut, C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. Chang, S. Lee, and S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” Journal of Language Resources and Evaluation , vol. 42, no. 4, pp. 335–359, December 2008
2008
Earlier work this paper cites.
B. Schuller, S. Steidl, and A. Batliner, “The INTERSPEECH 2009 emotion challenge,” in Interspeech 2009 - Eurospeech , Brighton, UK, September 2009, pp. 312–315
2009
Earlier work this paper cites.
A. Mahdhaoui and M. Chetouani, “Emotional speech classification based on multi view characterization,” in International Conference on Pattern Recognition (ICPR 2010) , Istanbul, Turkey, August 2010, pp. 4488–4491
2010
Earlier work this paper cites.
M. Wollmer, B. Schuller, F. Eyben, and G. Rigoll, “Combining long short-term memory and dynamic Bayesian networks for incremental emotion-sensitive artificial listening,” IEEE Journal of Selected Topics in Signal Processing , vol. 4, no. 5, pp. 867–881, October 2010
2010
Earlier work this paper cites.
Z. Zhang, F. Weninger, M. Wollmer, and B. Schuller, “Unsupervised learning in cross-corpus acoustic emotion recognition,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU 2011) , Waikoloa, HI, USA, December 2011, pp. 523–528
2011
Earlier work this paper cites.
B. Schuller, S. Steidl, A. Batliner, A. Vinciarelli, K. Scherer, F. Ringeval, M. Chetouani, F. Weninger, F. Eyben, E. Marchi, M. Mortillaro, H. Salamin, A. Polychroniou, F. Valente, and S. Kim, “The INTERSPEECH 2013 computational paralinguistics challenge: Social signals, conflict, emotion, autism,” in Interspeech 2013 , Lyon, France, August 2013, pp. 148–152
2013
Earlier work this paper cites.
Z. Huang, M. Dong, Q. Mao, and Y. Zhan, “Speech emotion recognition using CNN,” in ACM international conference on Multimedia (MM 2014) , Orlando, FL, USA, November 2014, pp. 801–804
2014
Earlier work this paper cites.
Q. Mao, M. Dong, Z. Huang, and Y. Zhan, “Learning salient features for speech emotion recognition using convolutional neural networks,” IEEE Transactions on Multimedia , vol. 16, no. 8, pp. 2203–2213, December 2014
2014
Earlier work this paper cites.
S. Mariooryad, R. Lotfian, and C. Busso, “Building a naturalistic emotional speech corpus by retrieving expressive behaviors from existing speech corpora,” in Interspeech 2014 , Singapore, September 2014, pp. 238–242
2014
Earlier work this paper cites.
N. Cummins, S. Scherer, J. Krajewski, S. Schnieder, J. Epps, and T. Quatieri, “A review of depression and suicide risk assessment using speech analysis,” Speech Communication , vol. 71, pp. 10–49, July 2015
2015
Earlier work this paper cites.
Z. Zhang, E. Coutinho, J. Deng, and B. Schuller, “Cooperative learning and its application to emotion recognition from speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 1, pp. 115–126, January 2015
2015
Cited alongside, same era.
H. Valpola, “From neural PCA to deep unsupervised learning,” in Advances in Independent Component Analysis and Learning Machines , E. Bingham, S. Kaski, J. Laaksonen, and J. Lampinen, Eds. London, UK: Academic Press, May 2015, pp. 143–171
2015
Cited alongside, same era.
A. Rasmusi, M. Berglund, M. Honkala, H. Valpola, and T. Raiko, “Semi-supervised learning with ladder networks,” in Advances in neural information processing systems (NIPS 2015) , Montreal, Canada, December 2015, pp. 3546–3554
2015
Cited alongside, same era.
J. Kim, G. Englebienne, K. Truong, and V. Evers, “Towards speech emotion recognition “in the Wild” using aggregated corpora and deep multi-task learning,” in Interspeech 2017 , Stockholm, Sweden, August 2017, pp. 1113–1117
2017
Later among the works it cites.
R. Lotfian and C. Busso, “Formulating emotion perception as a probabilistic model with application to categorical emotion classification,” in International Conference on Affective Computing and Intelligent Interaction (ACII 2017) , San Antonio, TX, USA, October 2017, pp. 415–420
2017
Later among the works it cites.
C. Busso, S. Parthasarathy, A. Burmania, M. AbdelWahab, N. Sadoughi, and E. Mower Provost, “MSP-IMPROV: An acted corpus of dyadic interactions to study emotion perception,” IEEE Transactions on Affective Computing , vol. 8, no. 1, pp. 67–80, January-March 2017
2017
Later among the works it cites.
N. Cummins, S. Amiriparian, G. Hagerer, A. Batliner, S. Steidl, and B. Schuller, “An image-based deep spectrum feature representation for the recognition of emotional speech,” in ACM international conference on Multimedia (MM 2017) , Mountain View, CA, USA, October 2017, pp. 478–484
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
F. Ringeval, F. Eyben, E. Kroupi, A. Yuce, J.-P. Thiran, T. Ebrahimi, D. Lalanne, and B. Schuller, “Prediction of asynchronous dimensional emotion ratings from audiovisual and physiological data,” Pattern Recognition Letters , vol. 66, pp. 22–30, November 2015
2015
Cited alongside, same era.
T. Dozat, “Incorporating Nesterov momentum into Adam,” in Workshop track at International Conference on Learning Representations (ICLR 2015) , San Juan, Puerto Rico, May 2015, pp. 1–4
2015
Cited alongside, same era.
Z. Zhang, F. Ringeval, B. Dong, E. Coutinho, E. Marchi, and B. Schuller, “Enhanced semi-supervised learning for multimodal emotion recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2016) , Shanghai, China, March 2016, pp. 5185–5189
2016
Cited alongside, same era.
M. Pezeshki, L. Fan, P. Brakel, A. Courville, and Y. Bengio, “Deconstructing the ladder network architecture,” in International Conference on Machine Learning (ICML 2016) , New York, NY, USA, June 2016, pp. 2368–2376
2016
Cited alongside, same era.
A. Burmania, S. Parthasarathy, and C. Busso, “Increasing the reliability of crowdsourcing evaluations using online quality assessment,” IEEE Transactions on Affective Computing , vol. 7, no. 4, pp. 374–388, October-December 2016
2016
Cited alongside, same era.
G. Trigeorgis, F. Ringeval, R. Brueckner, E. Marchi, M. Nicolaou, B. Schuller, and S. Zafeiriou, “Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2016) , Shanghai, China, March 2016, pp. 5200–5204
2016
Cited alongside, same era.
Y. Zhang, Y. Liu, F. Weninger, and B. Schuller, “Multi-task deep neural network with shared hidden layers: Breaking down the wall between emotion representations,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2017) , New Orleans, LA, USA, March 2017, pp. 4490–4494
2017
Cited alongside, same era.
S. Parthasarathy and C. Busso, “Jointly predicting arousal, valence and dominance with multi-task learning,” in Interspeech 2017 , Stockholm, Sweden, August 2017, pp. 1103–1107
2017
Cited alongside, same era.
2017
Later among the works it cites.
M. Neumann and N. Vu, “Attentive convolutional neural network based speech emotion recognition: A study on the impact of input features, signal length, and acted speech,” in Interspeech 2017 , Stockholm, Sweden, August 2017, pp. 1263–1267
2017
Later among the works it cites.
Z. Aldeneh and E. Mower Provost, “Using regional saliency for speech emotion recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2017) , New Orleans, LA, USA, March 2017, pp. 2741–2745
2017
Later among the works it cites.
R. Lotfian and C. Busso, “Predicting categorical emotions by jointly learning primary and secondary emotions through multitask learning,” in Interspeech 2018 , Hyderabad, India, September 2018, pp. 951–955
2018
Later among the works it cites.
F. Tao and G. Liu, “Advanced LSTM: a study about better time dependency modeling in emotion recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2018) , Calgary, AB, Canada, April 2018, pp. 2906–2910
2018
Later among the works it cites.
S. Parthasarathy and C. Busso, “Ladder networks for emotion recognition: Using unsupervised auxiliary tasks to improve predictions of emotional attributes,” in Interspeech 2018 , Hyderabad, India, September 2018, pp. 3698–3702
2018
Later among the works it cites.
Z. Zhang, J. Han, J. Deng, X. Xu, F. Ringeval, and B. Schuller, “Leveraging unlabeled data for emotion recognition with enhanced collaborative semi-supervised learning,” IEEE Access , vol. 6, pp. 22 196–22 209, April 2018
2018
Later among the works it cites.
J. Deng, X. Xu, Z. Zhang, S. Frühholz, and B. Schuller, “Semisupervised autoencoders for speech emotion recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 1, pp. 31–43, January 2018
2018
Later among the works it cites.
M. Abdelwahab and C. Busso, “Domain adversarial for acoustic emotion recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 12, pp. 2423–2435, December 2018
2018
Later among the works it cites.
J. Huang, Y. Li, J. Tao, Z. Lian, M. Niu, and J. Yi, “Speech emotion recognition using semi-supervised learning with ladder networks,” in Asian Conference on Affective Computing and Intelligent Interaction (ACII Asia 2018) , Beijing, China, May 2018, pp. 1–5
2018
Later among the works it cites.
Z. Yang and J. Hirschberg, “Predicting arousal and valence from waveforms and spectrograms using deep neural networks,” in Interspeech 2018 , Hyderabad, India, September 2018, pp. 3092–3096
2018
Later among the works it cites.
K. Sridhar, S. Parthasarathy, and C. Busso, “Role of regularization in the prediction of valence from speech,” in Interspeech 2018 , Hyderabad, India, September 2018, pp. 941–945
2018
Later among the works it cites.
R. Lotfian and C. Busso, “Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings,” IEEE Transactions on Affective Computing , vol. To appear, 2019
2019
Closest in time.
C. Busso and T. Rahman, “Unveiling the acoustic properties that describe the valence dimension,” in Interspeech 2012 , Portland, OR, USA, September 2012, pp. 1179–1182
2019
Closest in time.