Fetching the paper…

A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition · Around