Fetching the paper…
Reading the bibliography…
Audio-Visual Emotion Recognition (AVER) has garnered increasing attention in recent years for its critical role in creating emotion-ware intelligent machines.
A. MEHRABIAN, Communication without words, Psychol. Today 2 (4) (1968) 53–56
1968
Earlier work this paper cites.
M. Minsky, Society of mind, Simon and Schuster, 1988
1988
Earlier work this paper cites.
N. Schwarz, et al., Emotion, cognition, and decision making, Cognition & emotion 14 (4) (2000) 433–440
2000
Earlier work this paper cites.
N. Dalal, B. Triggs, Histograms of oriented gradients for human detection, in: 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), Vol. 1, Ieee, 2005, pp. 886–893
2005
Earlier work this paper cites.
G. Zhao, M. Pietikainen, Dynamic texture recognition using local binary patterns with an application to facial expressions, IEEE transactions on pattern analysis and machine intelligence 29 (6) (2007) 915–928
2007
Earlier work this paper cites.
Z. Zeng, M. Pantic, G. I. Roisman, T. S. Huang, A survey of affect recognition methods: Audio, visual, and spontaneous expressions, IEEE Transactions on Pattern Analysis and Machine Intelligence 31 (1) (2008) 39–58
2008
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, S. S. Narayanan, Iemocap: Interactive emotional dyadic motion capture database, Language resources and evaluation 42 (2008) 335–359
2008
Earlier work this paper cites.
L. Van der Maaten, G. Hinton, Visualizing data using t-sne., Journal of machine learning research 9 (11) (2008)
2008
Earlier work this paper cites.
B. Schuller, S. Steidl, A. Batliner, A. Vinciarelli, K. Scherer, F. Ringeval, M. Chetouani, F. Weninger, F. Eyben, E. Marchi, et al., The interspeech 2013 computational paralinguistics challenge: Social signals, conflict, emotion, autism, in: Proceedings INTERSPEECH 2013, 14th Annual Conference of the International Speech Communication Association, Lyon, France, 2013
2013
Earlier work this paper cites.
C.-H. Wu, J.-C. Lin, W.-L. Wei, Survey on audiovisual emotion recognition: databases, features, and data fusion strategies, APSIPA transactions on signal and information processing 3 (2014) e12
2014
Earlier work this paper cites.
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, R. Verma, Crema-d: Crowd-sourced emotional multimodal actors dataset, IEEE transactions on affective computing 5 (4) (2014) 377–390
2014
Earlier work this paper cites.
O. Ronneberger, P. Fischer, T. Brox, U-net: Convolutional networks for biomedical image segmentation, in: Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
F. Eyben, K. R. Scherer, B. W. Schuller, J. Sundberg, E. André, C. Busso, L. Y. Devillers, J. Epps, P. Laukka, S. S. Narayanan, et al., The geneva minimalistic acoustic parameter set (gemaps) for voice research and affective computing, IEEE transactions on affective computing 7 (2) (2015) 190–202
2015
Earlier work this paper cites.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, M. Paluri, Learning spatiotemporal features with 3d convolutional networks, in: Proceedings of the IEEE international conference on computer vision, 2015, pp. 4489–4497
2015
Earlier work this paper cites.
O. Parkhi, A. Vedaldi, A. Zisserman, Deep face recognition, in: BMVC 2015-Proceedings of the British Machine Vision Conference 2015, British Machine Vision Association, 2015
2015
Earlier work this paper cites.
Y. Fan, X. Lu, D. Li, Y. Liu, Video-based emotion recognition using cnn-rnn and c3d hybrid networks, in: Proceedings of the 18th ACM international conference on multimodal interaction, 2016, pp. 445–450
2016
Earlier work this paper cites.
G. Trigeorgis, F. Ringeval, R. Brueckner, E. Marchi, M. A. Nicolaou, B. Schuller, S. Zafeiriou, Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network, in: 2016 IEEE international conference on acoustics, speech and signal processing (ICASSP), IEEE, 2016, pp. 5200–5204
2016
Earlier work this paper cites.
D. Hendrycks, K. Gimpel, Gaussian error linear units (gelus), arXiv preprint arXiv:1606.08415 (2016)
2016
Earlier work this paper cites.
C. Busso, S. Parthasarathy, A. Burmania, M. AbdelWahab, N. Sadoughi, E. M. Provost, Msp-improv: An acted corpus of dyadic interactions to study emotion perception, IEEE Transactions on Affective Computing 8 (1) (2016) 67–80
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778
2016
Earlier work this paper cites.
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, M. Rohrbach, Multimodal compact bilinear pooling for visual question answering and visual grounding, in: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 2016, pp. 457–468
2016
Earlier work this paper cites.
S. Zhang, S. Zhang, T. Huang, W. Gao, Q. Tian, Learning affective features with a hybrid deep model for audio–visual emotion recognition, IEEE Transactions on Circuits and Systems for Video Technology 28 (10) (2017) 3030–3043
2017
Earlier work this paper cites.
P. Tzirakis, G. Trigeorgis, M. A. Nicolaou, B. W. Schuller, S. Zafeiriou, End-to-end multimodal emotion recognition using deep neural networks, IEEE Journal of selected topics in signal processing 11 (8) (2017) 1301–1309
2017
Earlier work this paper cites.
S. Chen, Q. Jin, J. Zhao, S. Wang, Multimodal multi-task learning for dimensional and continuous emotion recognition, in: Proceedings of the 7th Annual Workshop on Audio/Visual Emotion Challenge, 2017, pp. 19–26
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, et al., Cnn architectures for large-scale audio classification, in: 2017 ieee international conference on acoustics, speech and signal processing (icassp), IEEE, 2017, pp. 131–135
2017
Earlier work this paper cites.
R. Arandjelovic, A. Zisserman, Look, listen and learn, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 609–617
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, Advances in neural information processing systems 30 (2017)
2017
Earlier work this paper cites.
A. Zadeh, M. Chen, S. Poria, E. Cambria, L.-P. Morency, Tensor fusion network for multimodal sentiment analysis, in: Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, 2017, pp. 1103–1114
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, Grad-cam: Visual explanations from deep networks via gradient-based localization, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 618–626
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. S. Chung, A. Nagrani, A. Zisserman, Voxceleb2: Deep speaker recognition, Proc. Interspeech 2018 (2018) 1086–1090
2018
Earlier work this paper cites.
Q. Cao, L. Shen, W. Xie, O. M. Parkhi, A. Zisserman, Vggface2: A dataset for recognising faces across pose and age, in: 2018 13th IEEE international conference on automatic face & gesture recognition (FG 2018), IEEE, 2018, pp. 67–74
2018
Earlier work this paper cites.
J. Huang, Y. Li, J. Tao, Z. Lian, J. Yi, End-to-end continuous emotion recognition from video using 3d convlstm networks, in: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2018, pp. 6837–6841
2018
Earlier work this paper cites.
A. Owens, A. A. Efros, Audio-visual scene analysis with self-supervised multisensory features, in: Proceedings of the European conference on computer vision (ECCV), 2018, pp. 631–648
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. R. Livingstone, F. A. Russo, The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english, PloS one 13 (5) (2018) e0196391
2018
Earlier work this paper cites.
S. Yoon, S. Byun, K. Jung, Multimodal speech emotion recognition using audio and text, in: 2018 IEEE Spoken Language Technology Workshop (SLT), IEEE, 2018, pp. 112–118
2018
Earlier work this paper cites.
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, M. Paluri, A closer look at spatiotemporal convolutions for action recognition, in: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2018, pp. 6450–6459
2018
Earlier work this paper cites.
K. Hara, H. Kataoka, Y. Satoh, Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?, in: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2018, pp. 6546–6555
2018
Earlier work this paper cites.
J. Hu, L. Shen, G. Sun, Squeeze-and-excitation networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7132–7141
2018
Cited alongside, same era.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, R. Salakhutdinov, Multimodal transformer for unaligned multimodal language sequences, in: Proceedings of the conference. Association for Computational Linguistics. Meeting, Vol. 2019, NIH Public Access, 2019, p. 6558
2019
Cited alongside, same era.
2019
Cited alongside, same era.
E. Ghaleb, M. Popa, S. Asteriadis, Multimodal and temporal perception of audio-visual cues for emotion recognition, in: 2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII), IEEE, 2019, pp. 552–558
2019
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Wang, Y. Sun, W. Song, S. Gao, Y. Huang, Z. Chen, W. Ge, W. Zhang, Dpcnet: Dual path multi-excitation collaborative network for facial expression representation learning in videos, in: Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 101–110
2022
Later among the works it cites.
B. Zhang, H. Lv, P. Guo, Q. Shao, C. Yang, L. Xie, X. Xu, H. Bu, X. Chen, C. Zeng, et al., Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition, in: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2022, pp. 6182–6186
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
X. Jiang, Y. Zong, W. Zheng, C. Tang, W. Xia, C. Lu, J. Liu, Dfew: A large-scale database for recognizing dynamic facial expressions in the wild, in: Proceedings of the 28th ACM International Conference on Multimedia, 2020, pp. 2881–2889
2020
Cited alongside, same era.
L. Sun, Z. Lian, J. Tao, B. Liu, M. Niu, Multi-modal continuous dimensional emotion recognition using recurrent neural network and self-attention mechanism, in: Proceedings of the 1st international on multimodal sentiment analysis in real-life media challenge and workshop, 2020, pp. 27–34
2020
Cited alongside, same era.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, M. D. Plumbley, Panns: Large-scale pretrained audio neural networks for audio pattern recognition, IEEE/ACM Transactions on Audio, Speech, and Language Processing 28 (2020) 2880–2894
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, M. Auli, wav2vec 2.0: A framework for self-supervised learning of speech representations, Advances in neural information processing systems 33 (2020) 12449–12460
2020
Cited alongside, same era.
J. Huang, J. Tao, B. Liu, Z. Lian, M. Niu, Multimodal transformer fusion for continuous emotion recognition, in: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2020, pp. 3507–3511
2020
Cited alongside, same era.
2020
Cited alongside, same era.
S. Yoon, S. Dey, H. Lee, K. Jung, Attentive modality hopping mechanism for speech emotion recognition, in: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2020, pp. 3362–3366
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2022
Later among the works it cites.
L. Goncalves, C. Busso, Robust audiovisual emotion recognition: Aligning modalities, capturing temporal information, and handling missing features, IEEE Transactions on Affective Computing 13 (04) (2022) 2156–2170
2022
Later among the works it cites.
B. Shi, W.-N. Hsu, K. Lakhotia, A. Mohamed, Learning audio-visual speech representation by masked multimodal cluster prediction, in: International Conference on Learning Representations, 2022
2022
Later among the works it cites.
S. Verbitskiy, V. Berikov, V. Vyshegorodtsev, Eranns: Efficient residual audio neural networks for audio pattern recognition, Pattern Recognition Letters 161 (2022) 38–44
2022
Later among the works it cites.
K. Chumachenko, A. Iosifidis, M. Gabbouj, Self-attention fusion for audiovisual emotion recognition with incomplete data, in: 2022 26th International Conference on Pattern Recognition (ICPR), IEEE, 2022, pp. 2822–2828
2022
Later among the works it cites.
H. Mittal, P. Morgado, U. Jain, A. Gupta, Learning state-aware visual representations from audible interactions, Advances in Neural Information Processing Systems 35 (2022) 23765–23779
2022
Later among the works it cites.
X. Zhang, M. Li, S. Lin, H. Xu, G. Xiao, Transformer-based multimodal emotional perception for dynamic facial expression recognition in the wild, IEEE Transactions on Circuits and Systems for Video Technology (2023)
2023
Later among the works it cites.
S. Zhang, Y. Yang, C. Chen, X. Zhang, Q. Leng, X. Zhao, Deep learning-based multimodal emotion recognition from audio, visual, and text modalities: A systematic review of recent advancements and future prospects, Expert Systems with Applications (2023) 121692
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Li, L. Zhu, R. Mao, E. Cambria, Skier: A symbolic knowledge integrated model for conversational emotion recognition, in: Proceedings of the AAAI Conference on Artificial Intelligence, 2023
2023
Later among the works it cites.
Y. Gong, A. Rouditchenko, A. H. Liu, D. Harwath, L. Karlinsky, H. Kuehne, J. R. Glass, Contrastive audio-visual masked autoencoder, in: The Eleventh International Conference on Learning Representations, 2023
2023
Later among the works it cites.
P.-Y. Huang, V. Sharma, H. Xu, C. Ryali, H. Fan, Y. Li, S.-W. Li, G. Ghosh, J. Malik, C. Feichtenhofer, MAVil: Masked audio-video learners, in: Thirty-seventh Conference on Neural Information Processing Systems, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Wang, Y. Tang, Y. Wang, J. Guo, Z.-H. Deng, K. Han, Masked image modeling with local multi-scale reconstruction, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 2122–2131
2023
Later among the works it cites.
Y. Liu, S. Zhang, J. Chen, Z. Yu, K. Chen, D. Lin, Improving pixel-based mim by reducing wasted modeling capability, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 5361–5372
2023
Later among the works it cites.
K. Zhang, X. Wu, X. Xie, X. Zhang, H. Zhang, X. Chen, L. Sun, Werewolf-xl: A database for identifying spontaneous affect in large competitive group interactions, IEEE Transactions on Affective Computing 14 (02) (2023) 1201–1214
2023
Later among the works it cites.
Z. Lian, H. Sun, L. Sun, K. Chen, M. Xu, K. Wang, K. Xu, Y. He, Y. Li, J. Zhao, et al., Mer 2023: Multi-label learning, modality robustness, and semi-supervised learning, in: Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 9610–9614
2023
Later among the works it cites.
2023
Later among the works it cites.
L. Sun, Z. Lian, B. Liu, J. Tao, Mae-dfer: Efficient masked autoencoder for self-supervised dynamic facial expression recognition, in: Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 6110–6121
2023
Later among the works it cites.
J.-H. Hsu, C.-H. Wu, Applying segment-level attention on bi-modal transformer encoder for audio-visual emotion recognition, IEEE Transactions on Affective Computing (2023)
2023
Later among the works it cites.
L. Sun, Z. Lian, B. Liu, J. Tao, Efficient multimodal transformer with dual-level feature restoration for robust multimodal sentiment analysis, IEEE Transactions on Affective Computing (2023)
2023
Later among the works it cites.
M.-I. Georgescu, E. Fonseca, R. T. Ionescu, M. Lucic, C. Schmid, A. Arnab, Audiovisual masked autoencoders, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 16144–16154
2023
Later among the works it cites.
Z. Zhao, I. Patras, Prompting visual-language models for dynamic facial expression recognition, in: British Machine Vision Conference (BMVC), 2023, pp. 1–14
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Liu, W. Wang, C. Feng, H. Zhang, Z. Chen, Y. Zhan, Expression snippet transformer for robust video-based facial expression recognition, Pattern Recognition 138 (2023) 109368
2023
Later among the works it cites.
H. Li, H. Niu, Z. Zhu, F. Zhao, Intensity-aware loss for dynamic facial expression recognition in the wild, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, 2023, pp. 67–75
2023
Later among the works it cites.
H. Wang, B. Li, S. Wu, S. Shen, F. Liu, S. Ding, A. Zhou, Rethinking the learning paradigm for dynamic facial expression recognition, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 17958–17968
2023
Later among the works it cites.
M. Tran, Y. Kim, C.-C. Su, C.-H. Kuo, M. Soleymani, Saaml: A framework for semi-supervised affective adaptation via metric learning, in: Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 6004–6015
2023
Later among the works it cites.
L.-W. Chen, A. Rudnicky, Exploring wav2vec 2.0 fine tuning for improved speech emotion recognition, in: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Keesing, Y. S. Koh, V. Yogarajan, M. Witbrock, Emotion recognition toolkit (ertk): Standardising tools for emotion recognition research, in: Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 9693–9696
2023
Later among the works it cites.
Y. Lei, H. Cao, Audio-visual emotion recognition with preference learning based on intended and multi-modal perceived labels, IEEE Transactions on Affective Computing (2023) 1–16
2023
Later among the works it cites.
L. Goncalves, C. Busso, Learning cross-modal audiovisual representations with ladder networks for emotion recognition, in: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
G. Pei, H. Li, Y. Lu, Y. Wang, S. Hua, T. Li, Affective computing: Recent advances, challenges, and future trends, Intelligent Computing 3 (2024) 0076
2024
Closest in time.