Fetching the paper…
Reading the bibliography…
Natural human-computer interaction and audio-visual human behaviour sensing systems, which would achieve robust performance in-the-wild are more needed than ever as digital devices are increasingly becoming an indispensable part of our life.
W. Labov, Sociolinguistic patterns . University of Pennsylvania Press, 1972, no. 4
1972
Earlier work this paper cites.
W. E. Rinn, “The neuropsychology of facial expression: A review of the neurological and psychological mechanisms for producing facial expressions,” Psychological bulletin , vol. 95, no. 1, pp. 52–77, 01 1984
1984
Earlier work this paper cites.
P. Ekman, “Facial expression and emotion,” American Psychologist , vol. 28, no. 4, pp. 384–392, 1993
1993
Earlier work this paper cites.
K. R. Scherer and G. Ceschi, “Lost luggage: a field study of emotion–antecedent appraisal,” Motivation and emotion , vol. 21, no. 3, pp. 211–235, 1997
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, Nov 1997
1997
Earlier work this paper cites.
M. Pantic and L. J. Rothkrantz, “Expert system for automatic analysis of facial expressions,” Image and Vision Computing , vol. 18, no. 11, pp. 881–905, 2000
2000
Earlier work this paper cites.
R. Cowie, E. Douglas-Cowie, S. Savvidou*, E. McMahon, M. Sawey, and M. Schröder, “‘feeltrace’: An instrument for recording perceived emotion in real time,” in ISCA tutorial and research workshop (ITRW) on speech and emotion , 2000
2000
Earlier work this paper cites.
M. Kipp, “Anvil-a generic annotation tool for multimodal dialogue,” in Proc. 7th European Conference on Speech Communication and Technology , 2001
2001
Earlier work this paper cites.
F. Schiel, S. Steininger, and U. Türk, “The SmartKom Multimodal Corpus at BAS.” in LREC , 2002
2002
Earlier work this paper cites.
S. Brave and C. Nass, “The human-computer interaction handbook,” J. A. Jacko and A. Sears, Eds. Hillsdale, NJ, USA: L. Erlbaum Associates Inc., 2003, ch. Emotion in Human-computer Interaction, pp. 81–96
2003
Earlier work this paper cites.
M. Pantic and L. J. Rothkrantz, “Toward an affect-sensitive multimodal human-computer interaction,” Proceedings of the IEEE , vol. 91, no. 9, pp. 1370–1390, 2003
2003
Earlier work this paper cites.
J. Russell, B. J.A, and J. Fernandez-Dols, “Facial and vocal expressions of emotions,” Annu. Rev. Psychol. , vol. 54, pp. 329–349, 2003
2003
Earlier work this paper cites.
D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International Journal of Computer Vision (IJCV) , vol. 60, no. 2, pp. 91–110, 2004
2004
Earlier work this paper cites.
P. Wittenburg, H. Brugman, A. Russel, A. Klassmann, and H. Sloetjes, “Elan: a professional framework for multimodality research,” in Proceedings of LREC , vol. 2006, 2006, p. 5th
2006
Earlier work this paper cites.
P. Lang and M. M. Bradley, “The international affective picture system (iaps) in the study of emotion and attention,” Handbook of emotion elicitation and assessment , vol. 29, 2007
2007
Earlier work this paper cites.
2007
Earlier work this paper cites.
M. Grimm, K. Kroschel, and S. Narayanan, “The vera am mittag german audio-visual emotional speech database,” in ICME . IEEE, 2008, pp. 865–868
2008
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, no. 4, pp. 335–359, 2008
2008
Earlier work this paper cites.
J. Kim and E. André, “Emotion recognition based on physiological changes in music listening,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 30, no. 12, pp. 2067–2083, 2008
2008
Earlier work this paper cites.
R. Caruana, N. Karampatziakis, and A. Yessenalina, “An empirical evaluation of supervised learning in high dimensions,” in Proceedings of the 25th International Conference on Machine Learning . New York, NY, USA: ACM, 2008, pp. 96–103
2008
Earlier work this paper cites.
Z. Ambadar, J. F. Cohn, and L. I. Reed, “All smiles are not created equal: Morphology and timing of smiles perceived as amused, polite, and embarrassed/nervous,” Journal of nonverbal behavior , vol. 33, no. 1, p. 17—34, March 2009. [Online]. Available: http://europepmc.org/articles/PMC2701206
2009
Earlier work this paper cites.
Z. Zeng, M. Pantic, G. I. Roisman, and T. S. Huang, “A Survey of Affect Recognition Methods: Audio, Visual, and Spontaneous Expressions,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 31, no. 1, pp. 39–58, 2009
2009
Earlier work this paper cites.
R. Gross, I. Matthews, J. Cohn, T. Kanade, and S. Baker, “Multi-pie,” Image and Vision Computing , vol. 28, no. 5, pp. 807–813, 2010
2010
Earlier work this paper cites.
S. Scherer, I. Siegert, L. Bigalke, and S. Meudt, “Developing an expressive speech labelling tool incorporating the temporal characteristics of emotion.” in LREC , 2010
2010
Earlier work this paper cites.
P. Ekman and D. Cordaro, “What is meant by calling emotions basic,” Emotion Review , vol. 3, no. 4, pp. 364–370, 2011
2011
Earlier work this paper cites.
M. E. Ayadi, M. S. Kamel, and F. Karray, “Survey on speech emotion recognition: Features, classification schemes, and databases,” Pattern Recognition , vol. 44, no. 3, pp. 572–587, 2011
2011
Earlier work this paper cites.
B. Schuller, A. Batliner, S. Steidl, and D. Seppi, “Recognising Realistic Emotions and Affect in Speech: State of the Art and Lessons Learnt from the 1st Challenge,” Speech Communication, Special Issue on Sensing Emotion and Affect -– Facing Realism in Speech Process. , vol. 53, no. 9/10, pp. 1062–1087, Nov./Dec. 2011
2011
Earlier work this paper cites.
R. Böck, I. Siegert, M. Haase, J. Lange, and A. Wendemuth, “ikannotate–a tool for labelling, transcription, and annotation of emotionally coloured speech,” in International Conference on Affective Computing and Intelligent Interaction . Springer, 2011, pp. 25–34
2011
Earlier work this paper cites.
M. Wöllmer, E. Marchi, S. Squartini, and B. Schuller, “Robust multi-stream keyword and non-linguistic vocalization detection for computationally intelligent virtual agents,” in International Symposium on Neural Networks . Springer, 2011, pp. 496–505
2011
Earlier work this paper cites.
F. Weninger, J. Geiger, M. Wöllmer, B. Schuller, and G. Rigoll, “The munich 2011 chime challenge contribution: Nmf-blstm speech enhancement and recognition for reverberated multisource environments.”
2011
Cited alongside, same era.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research , vol. 12, pp. 2825–2830, 2011
2011
Cited alongside, same era.
G. McKeown, M. Valstar, R. Cowie, M. Pantic, and M. Schroder, “The SEMAINE database: Annotated multimodal records of emotionally colored conversations between a person and a limited agent,” IEEE Transactions on Affective Computing , vol. 3, no. 1, pp. 5–17, 2012
2012
Cited alongside, same era.
T. Bänziger, M. Mortillaro, and K. R. Scherer, “Introducing the Geneva Multimodal Expression Corpus for Experimental Research on Emotion Perception,” Emotion , vol. 12, no. 2, pp. 1161–1179, 2012
M. Valstar, B. Schuller, K. Smith, T. Almaev, F. Eyben, J. Krajewski, R. Cowie, and M. Pantic, “AVEC 2014: 3D Dimensional Affect and Depression Recognition Challenge,” Proceedings of the 4th ACM International Workshop on Audio/Visual Emotion Challenge (AVEC ’14) , pp. 3–10, 2014
2014
Later among the works it cites.
——, “Fast newton active appearance models,” in Proceedings of the IEEE Int’l Conf. on Image Processing (ICIP’14) , Paris, France, October 2014, pp. 1420–1424
2014
Later among the works it cites.
B. Schuller, S. Steidl, A. Batliner, J. Epps, F. Eyben, F. Ringeval, E. Marchi, and Y. Zhang, “The INTERSPEECH 2014 computational paralinguistics challenge: cognitive & physical load.” in Proc. of INTERSPEECH . Singapore, Singapore: ISCA, 2014, pp. 427–431
2014
Later among the works it cites.
S. K. D’mello and J. Kory, “A review and meta-analysis of multimodal affect detection systems,” ACM Computing Surveys (CSUR) , vol. 47, no. 3, p. 43, 2015
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2012
Cited alongside, same era.
S. Koelstra, C. Muhl, M. Soleymani, J.-S. Lee, A. Yazdani, T. Ebrahimi, T. Pun, A. Nijohlt, and I. Patras, “Deap: A database for emotion analysis using physiological signals,” IEEE Transactions on Affective Computing , vol. 3, no. 1, pp. 18–31, 2012
2012
Cited alongside, same era.
A. Dhall, R. Goecke, S. Lucey, and T. Gedeon, “Collecting Large, Richly Annotated Facial-Expression Databases from Movies,” IEEE Transactions on Multimedia , vol. 19, no. 3, pp. 34–41, 2012
2012
Cited alongside, same era.
S. Meudt, L. Bigalke, and F. Schwenker, “Atlas–an annotation tool for hci data utilizing machine learning methods,” Proc. of APD , vol. 12, pp. 5347–5352, 2012
2012
Cited alongside, same era.
I. Sneddon, M. McRorie, G. McKeown, and J. Hanratty, “The belfast induced natural emotion database,” IEEE Transactions on Affective Computing , vol. 3, no. 1, pp. 32–41, 2012
2012
Cited alongside, same era.
B. Schuller, M. Valster, F. Eyben, R. Cowie, and M. Pantic, “AVEC 2012: the continuous audio/visual emotion challenge,” Proc. 14th Int’l Conf. Multimodal Interaction Workshops , pp. 449–456, 2012
2012
Cited alongside, same era.
M. A. Nicolaou, H. Gunes, and M. Pantic, “Output-associative rvm regression for dimensional and continuous emotion prediction,” Image and Vision Computing , vol. 30, no. 3, pp. 186–196, 2012
2012
Cited alongside, same era.
A. Graves, Supervised sequence labelling with recurrent neural networks . Berlin/Heidelberg, Germany: Springer, 2012, vol. 385
2012
Cited alongside, same era.
S. M. Mavadati, M. H. Mahoor, K. Bartlett, P. Trinh, and J. F. Cohn, “Disfa: A spontaneous facial action intensity database,” IEEE Trans. Affect. Comput. , vol. 4, no. 2, pp. 151–160, Apr. 2013
2013
Cited alongside, same era.
S. Bilakhia, S. Petridis, A. Nijholt, and M. Pantic, “The mahnob mimicry database: A database of naturalistic human interactions,” Pattern recognition letters , vol. 66, pp. 52–61, 2015
2015
Later among the works it cites.
S. B.-C. Ofer Golan, Yana Sinai-Gavrilov, “The Cambridge mindreading face-voice battery for children (CAM-C): Complex emotion recognition in children with and without autism spectrum conditions,” Molecular Autism , vol. 22, no. 6, 2015
2015
Later among the works it cites.
E. Marchi, Y. Zhang, F. Eyben, F. Ringeval, and B. Schuller, “Autism and Speech, Language, and Emotion -– a Survey,” in Evaluating the Role of Speech Technology in Medical Case Management , H. Patil and M. Kulshreshtha, Eds. Berlin: De Gruyter, 2015
2015
Later among the works it cites.
E. Sariyanidi, H. Gunes, and A. Cavallaro, “Automatic analysis of facial affect: A survey of registration, representation, and recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 37, no. 6, pp. 1113–1133, 2015
2015
Later among the works it cites.
J. Vandeventer, A. J. Aubrey, P. L. Rosin, and D. Marshall, “4D Cardiff Conversation Database (4D CCDb): A 4D database of natural, dyadic conversations,” in Proceedings of the 1st Joint Conference on Facial Analysis, Animation and Auditory-Visual Speech Processing (FAAVSP 2015) , 2015
2015
Later among the works it cites.
J. Shen, S. Zafeiriou, G. G. Chrysos, J. Kossaifi, G. Tzimiropoulos, and M. Pantic, “The first facial landmark tracking in-the-wild challenge: Benchmark and results,” in Computer Vision Workshop (ICCVW), 2015 IEEE International Conference on . IEEE, 2015, pp. 1003–1011
2015
Later among the works it cites.
A. Asthana, S. Zafeiriou, G. Tzimiropoulos, S. Cheng, and M. Pantic, “From pixels to response maps: Discriminative image filtering for face alignment in the wild,” Pattern Analysis and Machine Intelligence, IEEE Transactions on , vol. 37, no. 6, pp. 1312–1320, 2015
2015
Later among the works it cites.
G. G. Chrysos, E. Antonakos, S. Zafeiriou, and P. Snape, “Offline deformable face tracking in arbitrary videos,” in Proceedings of the IEEE International Conference on Computer Vision Workshops , 2015, pp. 1–9
2015
Later among the works it cites.
R. Walecki, O. Rudovic, V. Pavlovic, and M. Pantic, “Variable-state latent conditional random fields for facial expression recognition and action unit detection,” in Proceedings of IEEE International Conference on Automatic Face and Gesture Recognition (FG’15) , Ljubljana, Slovenia, May 2015, pp. 1–8
2015
Later among the works it cites.
M. F. Valstar, T. Almaev, J. M. Girard, G. McKeown, M. Mehu, L. Yin, M. Pantic, and J. F. Cohn, “Fera 2015-second facial expression recognition and analysis challenge,” in Automatic Face and Gesture Recognition (FG), 2015 11th IEEE International Conference and Workshops on , vol. 6. IEEE, 2015, pp. 1–8
2015
Later among the works it cites.
H. Chen, J. Li, F. Zhang, Y. Li, and H. Wang, “3d model-based continuous emotion recognition,” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2015, pp. 1836–1845
2015
Later among the works it cites.
F. Weninger, J. Bergmann, and B. Schuller, “Introducing CURRENNT: the Munich Open-Source CUDA RecurREnt Neural Network Toolkit,” J. Machine Learning Research , vol. 16, pp. 547–551, 2015
2015
Later among the works it cites.
J. Kossaifi, G. Tzimiropoulos, and M. Pantic, “Fast and exact bi-directional fitting of active appearance models,” in Proceedings of the IEEE Int’l Conf. on Image Processing (ICIP’15) , Quebec City, QC, Canada, September 2015, pp. 1135–1139
2015
Later among the works it cites.
B. Schuller, S. Steidl, A. Batliner, S. Hantke, F. Hönig, J. R. Orozco-Arroyave, E. Nöth, Y. Zhang, and F. Weninger, “The interspeech 2015 computational paralinguistics challenge: Nativeness, parkinson & eating condition,” in Proc. of INTERSPEECH . Dresden, Germany: ISCA, 2015, pp. 478–482
2015
Later among the works it cites.
F. Ringeval, B. Schuller, M. Valstar, S. Jaiswal, E. Marchi, D. Lalanne, R. Cowie, and M. Pantic, “AV+EC 2015: The first affect recognition challenge bridging across audio, video, and physiological data,” in Proc. of the 5th International Workshop on Audio/Visual Emotion Challenge (AVEC) . Brisbane, Australia: ACM, 2015, pp. 3–8
2015
Later among the works it cites.
F. Zhou and F. De la Torre, “Generalized canonical time warping,” Transactions on Pattern Analysis and Machine Intelligence (PAMI) , vol. 38, no. 2, pp. 279–294, 2016
2016
Later among the works it cites.
Y. Panagakis, M. A. Nicolaou, S. Zafeiriou, and M. Pantic, “Robust correlated and individual component analysis,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Special Issue in Multimodal Pose Estimation and Behaviour Analysis , 2016
2016
Later among the works it cites.
F. Eyben, K. R. Scherer, B. W. Schuller, J. Sundberg, E. André, C. Busso, L. Y. Devillers, J. Epps, P. Laukka, S. S. Narayanan, and K. P. Truong, “The geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing,” IEEE Transactions on Affective Computing , vol. 7, no. 2, pp. 190–202, 2016
2016
Later among the works it cites.
M. Valstar, J. Gratch, B. Schuller, F. Ringeval, D. Lalanne, M. Torres Torres, S. Scherer, G. Stratou, R. Cowie, and M. Pantic, “Avec 2016: Depression, mood, and emotion recognition workshop and challenge,” in Proceedings of the 6th International Workshop on Audio/Visual Emotion Challenge , ser. AVEC ’16. New York, NY, USA: ACM, 2016, pp. 3–10
2016
Later among the works it cites.
Z. Zhang, F. Ringeval, J. Han, J. Deng, E. Marchi, and B. Schuller, “Facing realism in spontaneous emotion recognition from speech: Feature enhancement by autoencoder with LSTM neural networks,” in Proc. INTERSPEECH , San Francisco, CA, 2016, pp. 3593–3597
2016
Later among the works it cites.
J. Han, Z. Zhang, N. Cummins, F. Ringeval, and B. Schuller, “Strength modelling for real-worldautomatic continuous affect recognition from audiovisual signals,” Image and Vision Computing , 2016
2016
Later among the works it cites.
——, “Fast and exact newton and bidirectional fitting of active appearance models,” IEEE Transactions on Image Processing (TIP), accepted for publication , 2016
2016
Later among the works it cites.
F. Eyben, Real-time speech and music classification by large audio feature space extraction . Springer, 2016
2016
Later among the works it cites.
C. Georgakis, Y. Panagakis, S. Zafeiriou, and M. Pantic, “The conflict escalation resolution (confer) database,” Image and Vision Computing , 2017
2017
Later among the works it cites.
J. Kossaifi, G. Tzimiropoulos, S. Todorovic, and M. Pantic, “Afew-va database for valence and arousal estimation in-the-wild,” Image and Vision Computing , vol. 65, pp. 23 – 36, 2017, multimodal Sentiment Analysis and Mining in the Wild Image and Vision Computing
2017
Later among the works it cites.