Fetching the paper…
Reading the bibliography…
Video-based facial affect analysis has recently attracted increasing attention owing to its critical role in human-computer interaction.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
M. Pantic and L. J. M. Rothkrantz, “Automatic analysis of facial expressions: The state of the art,” IEEE Transactions on pattern analysis and machine intelligence , vol. 22, no. 12, pp. 1424–1445, 2000
2000
Earlier work this paper cites.
N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05) , vol. 1. Ieee, 2005, pp. 886–893
2005
Earlier work this paper cites.
O. Martin, I. Kotsia, B. Macq, and I. Pitas, “The enterface’05 audio-visual emotion database,” in 22nd International Conference on Data Engineering Workshops (ICDEW’06) . IEEE, 2006, pp. 8–8
2006
Earlier work this paper cites.
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol, “Extracting and composing robust features with denoising autoencoders,” in Proceedings of the 25th international conference on Machine learning , 2008, pp. 1096–1103
2008
Earlier work this paper cites.
M. Mansoorizadeh and N. Moghaddam Charkari, “Multimodal information fusion application to human emotion recognition from face and speech,” Multimedia Tools and Applications , vol. 49, no. 2, pp. 277–297, 2010
2010
Earlier work this paper cites.
E. Sariyanidi, H. Gunes, and A. Cavallaro, “Automatic analysis of facial affect: A survey of registration, representation, and recognition,” IEEE transactions on pattern analysis and machine intelligence , vol. 37, no. 6, pp. 1113–1133, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “Crema-d: Crowd-sourced emotional multimodal actors dataset,” IEEE transactions on affective computing , vol. 5, no. 4, pp. 377–390, 2014
2014
Earlier work this paper cites.
Y.-H. Byeon and K.-C. Kwak, “Facial expression recognition using 3d convolutional neural network,” International journal of advanced computer science and applications , vol. 5, no. 12, 2014
2014
Earlier work this paper cites.
S. Ebrahimi Kahou, V. Michalski, K. Konda, R. Memisevic, and C. Pal, “Recurrent neural networks for emotion recognition in video,” in Proceedings of the 2015 ACM on international conference on multimodal interaction , 2015, pp. 467–474
2015
Earlier work this paper cites.
L. Chao, J. Tao, M. Yang, Y. Li, and Z. Wen, “Long short term memory recurrent neural network based multimodal dimensional emotion recognition,” in Proceedings of the 5th international workshop on audio/visual emotion challenge , 2015, pp. 65–72
2015
Earlier work this paper cites.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 4489–4497
2015
Earlier work this paper cites.
Y. Fan, X. Lu, D. Li, and Y. Liu, “Video-based emotion recognition using cnn-rnn and c3d hybrid networks,” in Proceedings of the 18th ACM international conference on multimodal interaction , 2016, pp. 445–450
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Zhalehpour, O. Onder, Z. Akhtar, and C. E. Erdem, “Baum-1: A spontaneous audio-visual face database of affective and mental states,” IEEE Transactions on Affective Computing , vol. 8, no. 3, pp. 300–313, 2016
2016
Earlier work this paper cites.
Y. Güçlütürk, U. Güçlü, M. A. van Gerven, and R. van Lier, “Deep impression: Audiovisual deep residual networks for multimodal apparent personality trait recognition,” in European conference on computer vision . Springer, 2016, pp. 349–358
2016
Earlier work this paper cites.
A. Subramaniam, V. Patel, A. Mishra, P. Balasubramanian, and A. Mittal, “Bi-modal first impressions recognition using temporally ordered deep audio and stochastic visual features,” in European conference on computer vision . Springer, 2016, pp. 337–348
2016
Earlier work this paper cites.
V. Ponce-López, B. Chen, M. Oliu, C. Corneanu, A. Clapés, I. Guyon, X. Baró, H. J. Escalante, and S. Escalera, “Chalearn lap 2016: First round challenge on first impressions-dataset and results,” in European conference on computer vision . Springer, 2016, pp. 400–418
2016
Earlier work this paper cites.
R. Lotfian and C. Busso, “Formulating emotion perception as a probabilistic model with application to categorical emotion classification,” in 2017 Seventh International Conference on Affective Computing and Intelligent Interaction (ACII) . IEEE, 2017, pp. 415–420
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 6299–6308
2017
Earlier work this paper cites.
Z. Qiu, T. Yao, and T. Mei, “Learning spatio-temporal representation with pseudo-3d residual networks,” in proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 5533–5541
2017
Earlier work this paper cites.
H.-Y. Lee, J.-B. Huang, M. Singh, and M.-H. Yang, “Unsupervised representation learning by sorting sequences,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 667–676
2017
Earlier work this paper cites.
A. Zadeh, M. Chen, S. Poria, E. Cambria, and L.-P. Morency, “Tensor fusion network for multimodal sentiment analysis,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , 2017, pp. 1103–1114
2017
Earlier work this paper cites.
X.-S. Wei, C.-L. Zhang, H. Zhang, and J. Wu, “Deep bimodal regression of apparent personality traits from short video sequences,” IEEE Transactions on Affective Computing , vol. 9, no. 3, pp. 303–315, 2017
2017
Earlier work this paper cites.
C. Ventura, D. Masip, and A. Lapedriza, “Interpreting cnn models for apparent personality trait regression,” in Proceedings of the IEEE conference on computer vision and pattern recognition workshops , 2017, pp. 55–63
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
K. Hara, H. Kataoka, and Y. Satoh, “Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 2018, pp. 6546–6555
2018
Earlier work this paper cites.
N. Komodakis and S. Gidaris, “Unsupervised representation learning by predicting image rotations,” in International conference on learning representations (ICLR) , 2018
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” Proc. Interspeech 2018 , pp. 1086–1090, 2018
2018
Earlier work this paper cites.
S. R. Livingstone and F. A. Russo, “The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,” PloS one , vol. 13, no. 5, p. e0196391, 2018
2018
Earlier work this paper cites.
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 2018, pp. 6450–6459
2018
Cited alongside, same era.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 7132–7141
2018
Cited alongside, same era.
P. V. Rouast, M. T. Adam, and R. Chiong, “Deep learning for human affect recognition: Insights and new developments,” IEEE Transactions on Affective Computing , vol. 12, no. 2, pp. 524–543, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J.-R. Chang, Y.-S. Chen, and W.-C. Chiu, “Learning facial representations from the cycle-consistency of face,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 9680–9689
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International conference on machine learning . PMLR, 2021, pp. 10 347–10 357
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 10 012–10 022
2021
Later among the works it cites.
G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” in ICML , vol. 2, no. 3, 2021, p. 4
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Li, J. Zeng, S. Shan, and X. Chen, “Self-supervised representation learning from videos for facial action unit detection,” in Proceedings of the IEEE/CVF Conference on Computer vision and pattern recognition , 2019, pp. 10 924–10 933
2019
Cited alongside, same era.
E. Ghaleb, M. Popa, and S. Asteriadis, “Multimodal and temporal perception of audio-visual cues for emotion recognition,” in 2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII) . IEEE, 2019, pp. 552–558
2019
Cited alongside, same era.
D. Meng, X. Peng, K. Wang, and Y. Qiao, “Frame attention networks for facial expression recognition in videos,” in 2019 IEEE international conference on image processing (ICIP) . IEEE, 2019, pp. 3866–3870
2019
Cited alongside, same era.
X. Pan, G. Ying, G. Chen, H. Li, and W. Li, “A deep spatial and temporal aggregation framework for video-based facial expression recognition,” IEEE Access , vol. 7, pp. 48 807–48 815, 2019
2019
Cited alongside, same era.
X. Pan, W. Guo, X. Guo, W. Li, J. Xu, and J. Wu, “Deep temporal–spatial aggregation for video-based facial expression recognition,” Symmetry , vol. 11, no. 1, p. 52, 2019
2019
Cited alongside, same era.
R. Miyoshi, N. Nagata, and M. Hashimoto, “Facial-expression recognition from video using enhanced convolutional lstm,” in 2019 Digital Image Computing: Techniques and Applications (DICTA) . IEEE, 2019, pp. 1–6
2019
Cited alongside, same era.
L. Zhang, S. Peng, and S. Winkler, “Persemon: a deep network for joint analysis of apparent personality, emotion and their relationship,” IEEE Transactions on Affective Computing , 2019
2019
Cited alongside, same era.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 6202–6211
2019
Cited alongside, same era.
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer, “Multiscale vision transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 6824–6835
2021
Later among the works it cites.
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira, “Perceiver: General perception with iterative attention,” in International conference on machine learning . PMLR, 2021, pp. 4651–4664
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Liu, E. Sangineto, W. Bi, N. Sebe, B. Lepri, and M. Nadai, “Efficient training of visual transformers with small datasets,” Advances in Neural Information Processing Systems , vol. 34, pp. 23 818–23 830, 2021
2021
Later among the works it cites.
C. Feichtenhofer, H. Fan, B. Xiong, R. Girshick, and K. He, “A large-scale study on unsupervised spatiotemporal representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 3299–3309
2021
Later among the works it cites.
R. Miyoshi, N. Nagata, and M. Hashimoto, “Enhanced convolutional lstm with spatial and temporal skip connections and temporal gates for facial expression recognition from video,” Neural Computing and Applications , vol. 33, no. 13, pp. 7381–7392, 2021
2021
Later among the works it cites.
R. Zhao, T. Liu, Z. Huang, D. P. Lun, and K.-M. Lam, “Spatial-temporal graphs plus transformers for geometry-guided facial expression recognition,” IEEE Transactions on Affective Computing , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 000–16 009
2022
Later among the works it cites.
Z. Tong, Y. Song, J. Wang, and L. Wang, “VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training,” in Advances in Neural Information Processing Systems , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Wang, Y. Sun, Y. Huang, Z. Liu, S. Gao, W. Zhang, W. Ge, and W. Zhang, “Ferv39k: A large-scale multi-scene dataset for facial expression recognition in videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 20 922–20 931
2022
Later among the works it cites.
Y. Liu, W. Dai, C. Feng, W. Wang, G. Yin, J. Zeng, and S. Shan, “Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 24–32
2022
Later among the works it cites.
Y. Wang, Y. Sun, W. Song, S. Gao, Y. Huang, Z. Chen, W. Ge, and W. Zhang, “Dpcnet: Dual path multi-excitation collaborative network for facial expression representation learning in videos,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 101–110
2022
Later among the works it cites.
A. Bulat, S. Cheng, J. Yang, A. Garbett, E. Sanchez, and G. Tzimiropoulos, “Pre-training strategies and datasets for facial representation learning,” in European Conference on Computer Vision . Springer, 2022, pp. 107–125
2022
Later among the works it cites.
Y. Zheng, H. Yang, T. Zhang, J. Bao, D. Chen, Y. Huang, L. Yuan, D. Chen, M. Zeng, and F. Wen, “General facial representation learning in a visual-linguistic manner,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18 697–18 709
2022
Later among the works it cites.
K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu et al. , “A survey on vision transformer,” IEEE transactions on pattern analysis and machine intelligence , vol. 45, no. 1, pp. 87–110, 2022
2022
Later among the works it cites.
Y. Li, C.-Y. Wu, H. Fan, K. Mangalam, B. Xiong, J. Malik, and C. Feichtenhofer, “Mvitv2: Improved multiscale vision transformers for classification and detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4804–4814
2022
Later among the works it cites.
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 3202–3211
2022
Later among the works it cites.
K. Ranasinghe, M. Naseer, S. Khan, F. S. Khan, and M. S. Ryoo, “Self-supervised video transformer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 2874–2884
2022
Later among the works it cites.
Y. Liu, C. Feng, X. Yuan, L. Zhou, W. Wang, J. Qin, and Z. Luo, “Clip-aware expressive feature learning for video-based facial expression recognition,” Information Sciences , vol. 598, pp. 182–195, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
L. Goncalves and C. Busso, “Robust audiovisual emotion recognition: Aligning modalities, capturing temporal information, and handling missing features,” IEEE Transactions on Affective Computing , vol. 13, no. 04, pp. 2156–2170, 2022
2022
Later among the works it cites.
M. Tran and M. Soleymani, “A pre-trained audio-visual transformer for emotion recognition,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 4698–4702
2022
Later among the works it cites.
2022
Later among the works it cites.
C. Suman, S. Saha, A. Gupta, S. K. Pandey, and P. Bhattacharyya, “A multi-modal personality prediction system,” Knowledge-Based Systems , vol. 236, p. 107715, 2022
2022
Later among the works it cites.
H. Li, H. Niu, Z. Zhu, and F. Zhao, “Intensity-aware loss for dynamic facial expression recognition in the wild,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 1, 2023, pp. 67–75
2023
Closest in time.
H. Wang, B. Li, S. Wu, S. Shen, F. Liu, S. Ding, and A. Zhou, “Rethinking the learning paradigm for dynamic facial expression recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 17 958–17 968
2023
Closest in time.
Y. Lei and H. Cao, “Audio-visual emotion recognition with preference learning based on intended and multi-modal perceived labels,” IEEE Transactions on Affective Computing , pp. 1–16, 2023
2023
Closest in time.
K. Zhang, X. Wu, X. Xie, X. Zhang, H. Zhang, X. Chen, and L. Sun, “Werewolf-xl: A database for identifying spontaneous affect in large competitive group interactions,” IEEE Transactions on Affective Computing , vol. 14, no. 02, pp. 1201–1214, 2023
2023
Closest in time.