Fetching the paper…
Reading the bibliography…
Obtaining large, human labelled speech datasets to train models for emotion recognition is a notoriously challenging task, hindered by annotation cost and label ambiguity.
The effect of position in utterance on speech segment duration in english
D. K. Oller · 1973
Earlier work this paper cites.
Visual following and pattern discrimination of face-like stimuli by newborn infants
C. C. Goren, M. Sarty, and P. Y. Wu · 1975
Earlier work this paper cites.
Unsupervised learning
H. B. Barlow · 1989
Earlier work this paper cites.
Newborns’ preferential tracking of face-like stimuli and its subsequent decline
M. H. Johnson, S. Dziurawiec, H. Ellis, and J. Morton · 1991
Earlier work this paper cites.
A new emotion database: considerations, sources and scope
E. Douglas-Cowie, R. Cowie, and M. Schröder · 2000
Earlier work this paper cites.
Influence of emotion and focus location on prosody in matched statements and questions
M. D. Pell · 2001
Earlier work this paper cites.
Ldc emotional prosody speech transcripts database
M. Liberman, K. Davis, M. Grossman, N. Martey, and J. Bell · 2002
Earlier work this paper cites.
You stupid tin box-children interacting with the aibo robot: A cross-linguistic emotional speech corpus
A. Batliner, C. Hacker, S. Steidl, E. Nøth, S. D’Arcy, M. J. Russell, and M. Wong · 2004
Earlier work this paper cites.
Analysis of emotion recognition using facial expressions, speech and multimodal information
C. Busso, Z. Deng, S. Yildirim, M. Bulut, C. M. Lee, A. Kazemzadeh, S. Lee, U. Neumann, and S. Narayanan · 2004
Earlier work this paper cites.
A database of german emotional speech
F. Burkhardt, A. Paeschke, M. Rolfes, W. F. Sendlmeier, and B. Weiss · 2005
Earlier work this paper cites.
Prosody–face interactions in emotional processing as revealed by the facial affect decision task
M. D. Pell · 2005
Earlier work this paper cites.
Model compression
C. Bucilua, R. Caruana, and A. Niculescu-Mizil · 2006
Earlier work this paper cites.
The enterface’05 audio-visual emotion database
O. Martin, I. Kotsia, B. Macq, and I. Pitas · 2006
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan · 2008
Earlier work this paper cites.
Facial expression and prosodic prominence: Effects of modality and facial area
M. Swerts and E. Krahmer · 2008
Earlier work this paper cites.
Not on the face alone: perception of contextualized face expressions in huntington’s disease
H. Aviezer, S. Bentin, R. R. Hassin, W. S. Meschino, J. Kennedy, S. Grewal, S. Esmail, S. Cohen, and M. Moscovitch · 2009
Earlier work this paper cites.
Prosody off the top of the head: Prosodic contrasts can be discriminated by head motion
E. Cvejic, J. Kim, and C. Davis · 2010
Earlier work this paper cites.
The development of emotion perception in face and voice during infancy
T. Grossmann · 2010
Earlier work this paper cites.
Cross-corpus acoustic emotion recognition: Variances and strategies
B. Schuller, B. Vlasenko, F. Eyben, M. Wollmer, A. Stuhlsatz, A. Wendemuth, and G. Rigoll · 2010
Earlier work this paper cites.
Unsupervised learning in cross-corpus acoustic emotion recognition
Z. Zhang, F. Weninger, M. Wöllmer, and B. Schuller · 2011
Earlier work this paper cites.
Collecting large, richly annotated facial-expression databases from movies
A. Dhall, R. Goecke, S. Lucey, T. Gedeon, et al · 2012
Earlier work this paper cites.
Challenges in representation learning: A report on three machine learning contests
I. J. Goodfellow, D. Erhan, P. L. Carrier, A. Courville, M. Mirza, B. Hamner, W. Cukierski, Y. Tang, D. Thaler, D.-H. Lee, et al · 2013
Earlier work this paper cites.
Inherently ambiguous: Facial expressions of emotions, in context
R. R. Hassin, H. Aviezer, and S. Bentin · 2013
Cited alongside, same era.
Analysis and compensation of the reaction lag of evaluators in continuous emotional annotations
S. Mariooryad and C. Busso · 2013
Cited alongside, same era.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Cited alongside, same era.
Return of the devil in the details: Delving deep into convolutional nets
K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman · 2014
Cited alongside, same era.
Autoencoder-based unsupervised domain adaptation for speech emotion recognition
J. Deng, Z. Zhang, F. Eyben, and B. Schuller · 2014
Cited alongside, same era.
Linked source and target domain subspace feature transfer learning–exemplified by speech emotion recognition
J. Deng, Z. Zhang, and B. Schuller · 2014
Using regional saliency for speech emotion recognition
Z. Aldeneh and E. M. Provost · 2017
Later among the works it cites.
Look, listen and learn
R. Arandjelovic and A. Zisserman · 2017
Later among the works it cites.
See, hear, and read: Deep aligned representations
Y. Aytar, C. Vondrick, and A. Torralba · 2017
Later among the works it cites.
Moonshine: Distilling with cheap convolutions
E. J. Crowley, G. Gray, and A. Storkey · 2017
Later among the works it cites.
An image-based deep spectrum feature representation for the recognition of emotional speech
N. Cummins, S. Amiriparian, G. Hagerer, A. Batliner, S. Steidl, and B. W. Schuller · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning small-size dnn with output-distribution-based criteria
J. Li, R. Zhao, J.-T. Huang, and Y. Gong · 2014
Cited alongside, same era.
Emotion in the voice influences the way we scan emotional faces
S. Rigoulot and M. D. Pell · 2014
Cited alongside, same era.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Cited alongside, same era.
Towards an intelligent framework for multimodal affective data analysis
S. Poria, E. Cambria, A. Hussain, and G.-B. Huang · 2015
Cited alongside, same era.
Image based static facial expression recognition with multiple deep network learning
Z. Yu and C. Zhang · 2015
Cited alongside, same era.
J. Han, Z. Zhang, M. Schmitt, M. Pantic, and B. Schuller · 2017
Later among the works it cites.
Learning supervised scoring ensemble for emotion recognition in the wild
P. Hu, D. Cai, S. Wang, A. Yao, and Y. Chen · 2017
Later among the works it cites.
Combining convolutional neural networks for emotion recognition
C. Huang · 2017
Later among the works it cites.
Deep temporal models using identity skip-connections for speech emotion recognition
J. Kim, G. Englebienne, K. P. Truong, and V. Evers · 2017
Later among the works it cites.
VoxCeleb: a large-scale speaker identification dataset
A. Nagrani, J. S. Chung, and A. Zisserman · 2017
Later among the works it cites.
Audio-visual emotion recognition in video clips
F. Noroozi, M. Marjanovic, A. Njegus, S. Escalera, and G. Anbarjafari · 2017
Later among the works it cites.
A study of speaker verification performance with expressive speech
S. Parthasarathy, C. Zhang, J. H. Hansen, and C. Busso · 2017
Later among the works it cites.
Deep learning is robust to massive label noise
D. Rolnick, A. Veit, S. Belongie, and N. Shavit · 2017
Later among the works it cites.
Learning visual emotion distributions via multi-modal features fusion
S. Zhao, G. Ding, Y. Gao, and J. Han · 2017
Later among the works it cites.
VGGFace2: A dataset for recognising faces across pose and age
Q. Cao, L. Shen, W. Xie, O. M. Parkhi, and A. Zisserman · 2018
Closest in time.
Squeeze-and-excitation networks
J. Hu, L. Shen, and G. Sun · 2018
Closest in time.
On learning associations of faces and voices
C. Kim, H. V. Shin, T.-H. Oh, A. Kaspar, M. Elgharib, and W. Matusik · 2018
Closest in time.
Transfer learning for improving speech emotion classification accuracy
S. Latif, R. Rana, S. Younis, J. Qadir, and J. Epps · 2018
Closest in time.
Exploring the limits of weakly supervised pretraining
D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. van der Maaten · 2018
Closest in time.
Learnable PINs: Cross-modal embeddings for person identity
A. Nagrani, S. Albanie, and A. Zisserman · 2018
Closest in time.
Seeing voices and hearing faces: Cross-modal biometric matching
A. Nagrani, S. Albanie, and A. Zisserman · 2018
Closest in time.
Self-supervised learning of a facial attribute embedding from video
O. Wiles, A. S. Koepke, and A. Zisserman · 2018
Closest in time.