Fetching the paper…
Reading the bibliography…
An important challenge in emotion recognition is to develop methods that can leverage unlabeled training data.
Emotion recognition for in-the-wild videos
Liu, H., Zeng, J., Shan, S., Chen, X., 2020 · 2002
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Chen, X., Fan, H., Girshick, R., He, K., 2020b · 2003
Earlier work this paper cites.
IEMOCAP: Interactive emotional dyadic motion capture database
Busso, C., Bulut, M., Lee, C.C., Kazemzadeh, A., Mower, E., Kim, S., Chang, J.N., Lee, S., Narayanan, S.S., 2008 · 2008
Earlier work this paper cites.
Nonnegative matrix factorization with the Itakura-Saito divergence: With application to music analysis
Févotte, C., Bertin, N., Durrieu, J.L., 2009 · 2009
Earlier work this paper cites.
Survey on speech emotion recognition: Features, classification schemes, and databases
El Ayadi, M., Kamel, M.S., Karray, F., 2011 · 2011
Earlier work this paper cites.
CREMA-D: Crowd-sourced emotional multimodal actors dataset
Cao, H., Cooper, D.G., Keutmann, M.K., Gur, R.C., Nenkova, A., Verma, R., 2014 · 2014
Earlier work this paper cites.
Influence of lips on the production of vowels based on finite element simulations and experiments
Arnela, M., Blandin, R., Dabbaghchian, S., Guasch, O., Alías, F., Pelorson, X., Van Hirtum, A., Engwall, O., 2016 · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles, in: European Conference on Computer Vision (ECCV), pp. 69–84
Noroozi, M., Favaro, P., 2016 · 2016
Earlier work this paper cites.
How far are we from solving the 2D & 3D Face Alignment problem? (and a dataset of 230,000 3D facial landmarks), in: IEEE/CVF International Conference on Computer Vision (ICCV), pp. 1021–1030
Bulat, A., Tzimiropoulos, G., 2017 · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., He, K., 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization, in: International Conference on Learning Representations (ICLR)
Loshchilov, I., Hutter, F., 2017 · 2017
Earlier work this paper cites.
Nonverbal communication
Mehrabian, A., 2017 · 2017
Earlier work this paper cites.
Deep multimodal learning: A survey on recent advances and trends
Ramachandram, D., Taylor, G.W., 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al., 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I., 2017 · 2017
Earlier work this paper cites.
VoxCeleb2: Deep speaker recognition, in: INTERSPEECH, pp. 1086–1090
Chung, J., Nagrani, A., Zisserman, A., 2018 · 2018
Earlier work this paper cites.
Unsupervised representation learning by predicting image rotations, in: International Conference on Learning Representations (ICLR)
Gidaris, S., Singh, P., Komodakis, N., 2018 · 2018
Earlier work this paper cites.
Aff-Wild2: Extending the Aff-Wild database for affect recognition
Kollias, D., Zafeiriou, S., 2018 · 2018
Earlier work this paper cites.
The Ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in north american english
Livingstone, S.R., Russo, F.A., 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding, in: Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2019 · 2019
Earlier work this paper cites.
Co-separating sounds of visual objects, in: IEEE/CVF International Conference on Computer Vision (ICCV), pp. 3879–3888
Gao, R., Grauman, K., 2019 · 2019
Earlier work this paper cites.
Multimodal and temporal perception of audio-visual cues for emotion recognition, in: International Conference on Affective Computing and Intelligent Interaction (ACII), pp. 552–558
Ghaleb, E., Popa, M., Asteriadis, S., 2019 · 2019
Earlier work this paper cites.
Multimodal transformer for unaligned multimodal language sequences, in: Association for Computational Linguistics. Meeting, NIH Public Access. p. 6558
Tsai, Y.H.H., Bai, S., Liang, P.P., Kolter, J.Z., Morency, L.P., Salakhutdinov, R., 2019 · 2019
Cited alongside, same era.
The sound of motions, in: IEEE/CVF International Conference on Computer Vision (ICCV), pp. 1735–1744
Zhao, H., Gan, C., Ma, W.C., Torralba, A., 2019 · 2019
Cited alongside, same era.
Self-supervised multimodal versatile networks
Alayrac, J.B., Recasens, A., Schneider, R., Arandjelović, R., Ramapuram, J., De Fauw, J., Smaira, L., Dieleman, S., Zisserman, A., 2020 · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al., 2020 · 2020
Cited alongside, same era.
Multimodal attention-mechanism for temporal emotion recognition, in: IEEE International Conference on Image Processing (ICIP), pp. 251–255
Ghaleb, E., Niehues, J., Asteriadis, S., 2020 · 2020
MultiMAE: Multi-modal multi-task masked autoencoders, in: European Conference on Computer Vision (ECCV), pp. 348–367
Bachmann, R., Mizrahi, D., Atanov, A., Zamir, A., 2022 · 2022
Later among the works it cites.
Data2vec: A general framework for self-supervised learning in speech, vision and language, in: International Conference on Machine Learning (ICML), pp. 1298–1312
Baevski, A., Hsu, W.N., Xu, Q., Babu, A., Gu, J., Auli, M., 2022 · 2022
Later among the works it cites.
Self-attention fusion for audiovisual emotion recognition with incomplete data, in: International Conference on Pattern Recognition (ICPR), pp. 2822–2828
Chumachenko, K., Iosifidis, A., Gabbouj, M., 2022 · 2022
Later among the works it cites.
Masked autoencoders as spatiotemporal learners
Feichtenhofer, C., Li, Y., He, K., et al., 2022 · 2022
Later among the works it cites.
Multimodal masked autoencoders learn transferable representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dfew: A large-scale database for recognizing dynamic facial expressions in the wild, in: ACM International Conference on Multimedia, pp. 2881–2889
Jiang, X., Zong, Y., Zheng, W., Tang, C., Xia, W., Lu, C., Liu, J., 2020 · 2020
Cited alongside, same era.
VATT: Transformers for multimodal self-supervised learning from raw video, audio and text
Akbari, H., Yuan, L., Qian, R., Chuang, W.H., Chang, S.F., Cui, Y., Gong, B., 2021 · 2021
Cited alongside, same era.
An audiovisual and contextual approach for categorical and continuous emotion recognition in-the-wild, in: IEEE/CVF International Conference on Computer Vision (ICCV), pp. 3645–3651
Antoniadis, P., Pikoulis, I., Filntisis, P.P., Maragos, P., 2021 · 2021
Cited alongside, same era.
ViViT: A video vision transformer, in: IEEE/CVF International Conference on Computer Vision (ICCV), pp. 6836–6846
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., Schmid, C., 2021 · 2021
Cited alongside, same era.
BEiT: BERT pre-training of image transformers, in: International Conference on Learning Representations (ICLR)
Bao, H., Dong, L., Piao, S., Wei, F., 2021 · 2021
Cited alongside, same era.
PeCo: Perceptual codebook for BERT pre-training of vision transformers, in: AAAI Conference on Artificial Intelligence, pp. 552–560
Dong, X., Bao, J., Zhang, T., Chen, D., Zhang, W., Yuan, L., Chen, D., Wen, F., Yu, N., 2021 · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis, in: IEEE/CVF conference on computer vision and pattern recognition (CVPR), pp. 12873–12883
Esser, P., Rombach, R., Ommer, B., 2021 · 2021
Cited alongside, same era.
Geng, X., Liu, H., Lee, L., Schuurams, D., Levine, S., Abbeel, P., 2022 · 2022
Later among the works it cites.
Robust audiovisual emotion recognition: Aligning modalities, capturing temporal information, and handling missing features
Goncalves, L., Busso, C., 2022 · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 16000–16009
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., Girshick, R., 2022 · 2022
Later among the works it cites.
Masked autoencoders that listen
Huang, P.Y., Xu, H., Li, J., Baevski, A., Auli, M., Galuba, W., Metze, F., Feichtenhofer, C., 2022 · 2022
Later among the works it cites.
Audio self-supervised learning: A survey
Liu, S., Mallol-Ragolta, A., Parada-Cabaleiro, E., Qian, K., Jing, X., Kathan, A., Hu, B., Schuller, B.W., 2022 · 2022
Later among the works it cites.
VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z., Song, Y., Wang, J., Wang, L., 2022 · 2022
Later among the works it cites.
A pre-trained audio-visual transformer for emotion recognition, in: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4698–4702
Tran, M., Soleymani, M., 2022 · 2022
Later among the works it cites.
Simmim: A simple framework for masked image modeling, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9653–9663
Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., Hu, H., 2022 · 2022
Later among the works it cites.
A survey on masked autoencoder for self-supervised learning in vision and beyond
Zhang, C., Zhang, C., Song, J., Yi, J.S.K., Zhang, K., Kweon, I.S., 2022 · 2022
Later among the works it cites.
Exploring Wav2vec 2.0 fine tuning for improved speech emotion recognition, in: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE. pp. 1–5
Chen, L.W., Rudnicky, A., 2023 · 2023
Closest in time.
S2f2: Self-supervised high fidelity face reconstruction from monocular image, in: IEEE International Conference on Automatic Face and Gesture Recognition (FG), pp. 1–8
Dib, A., Ahn, J., Thebault, C., Gosselin, P.H., Chevallier, L., 2023 · 2023
Closest in time.
Contrastive masked autoencoders are stronger vision learners
Huang, Z., Jin, X., Lu, C., Hou, Q., Cheng, M.M., Fu, D., Shen, X., Feng, J., 2023 · 2023
Closest in time.
SS-VAERR: Self-supervised apparent emotional reaction recognition from video, in: IEEE International Conference on Automatic Face and Gesture Recognition (FG), pp. 1–8
Jegorova, M., Petridis, S., Pantic, M., 2023 · 2023
Closest in time.
MAGE: Masked generative encoder to unify representation learning and image synthesis, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2142–2152
Li, T., Chang, H., Mishra, S., Zhang, H., Katabi, D., Krishnan, D., 2023 · 2023
Closest in time.
A vector quantized masked autoencoder for speech emotion recognition, in: IEEE ICASSP Workshop on Self-Supervision in Audio, Speech and Beyond (SASB), pp. 1–5
Sadok, S., Leglaive, S., Séguier, R., 2023 · 2023
Closest in time.
Transformer-based multimodal emotional perception for dynamic facial expression recognition in the wild
Zhang, X., Li, M., Lin, S., Xu, H., Xiao, G., 2023 · 2023
Closest in time.
A multimodal dynamical variational autoencoder for audiovisual speech representation learning
Sadok, S., Leglaive, S., Girin, L., Alameda-Pineda, X., Séguier, R., 2024 · 2024
Closest in time.
AnCoGen: Analysis, control and generation of speech with a masked autoencoder, in: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Sadok, S., Leglaive, S., Girin, L., Richard, G., Alameda-Pineda, X., 2025 · 2025
Closest in time.