Fetching the paper…
Reading the bibliography…
Emotions lie on a broad continuum and treating emotions as a discrete number of classes limits the ability of a model to capture the nuances in the continuum.
“Acoustic concomitants of emotional dimensions: Judging affect from synthesized tone sequences.,”
Klaus R Scherer, · 1972
Earlier work this paper cites.
“Communicating emotion: The role of prosodic features.,”
Robert W Frick, · 1985
Earlier work this paper cites.
The emotions
Robert Plutchik, · 1991
Earlier work this paper cites.
Are there basic emotions?
Paul Ekman, · 1992
Earlier work this paper cites.
Emotions and multilingualism
Aneta Pavlenko, · 2005
Earlier work this paper cites.
“Iemocap: Interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan, · 2008
Earlier work this paper cites.
“Crema-d: Crowd-sourced emotional multimodal actors dataset,”
Houwei Cao, David G Cooper, Michael K Keutmann, Ruben C Gur, Ani Nenkova, and Ragini Verma, · 2014
Earlier work this paper cites.
“librosa: Audio and music signal analysis in python,”
Brian McFee, Colin Raffel, Dawen Liang, Daniel P Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto, · 2015
Cited alongside, same era.
“Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,”
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency, · 2018
Cited alongside, same era.
“Meld: A multimodal multi-party dataset for emotion recognition in conversations,”
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea, · 2018
Cited alongside, same era.
“The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,”
Steven R Livingstone and Frank A Russo, · 2018
Cited alongside, same era.
“AudioCaps: Generating Captions for Audios in The Wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley, · 2020
Later among the works it cites.
“What is the ground truth? reliability of multi-annotator data for audio tagging,”
Irene Martín-Morató and Annamaria Mesaros, · 2021
Later among the works it cites.
“A proposal for multimodal emotion recognition using aural transformers and action units on ravdess dataset,”
Cristina Luna-Jiménez, Ricardo Kleinlein, David Griol, Zoraida Callejas, Juan M Montero, and Fernando Fernández-Martínez, · 2021
Later among the works it cites.
“Clap: Learning audio concepts from natural language supervision,”
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang, · 2022
Closest in time.
“Audio retrieval with wavtext5k and clap training,”
Soham Deshmukh, Benjamin Elizalde, and Huaming Wang, · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Improving content-based audio retrieval by vocal imitation feedback,”
Bongjun Kim and Bryan Pardo, · 2019
Cited alongside, same era.
“Clotho: an audio captioning dataset,”
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen, · 2020
Cited alongside, same era.
[Online]
“Praat.,” https://www.fon.hum.uva.nl/praat
Cited in the paper.
Closest in time.
“Fsd50k: An open dataset of human-labeled sound events,”
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra, · 2022
Closest in time.
“The ability of self-supervised speech models for audio representations,”
Tung-Yu Wu, Chen-An Li, Tzu-Han Lin, Tsu-Yuan Hsu, and Hung-Yi Lee, · 2022
Closest in time.