Fetching the paper…
Reading the bibliography…
Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable.
“Acoustic concomitants of emotional dimensions: Judging affect from synthesized tone sequences.,”
Klaus R Scherer, · 1972
Earlier work this paper cites.
“Communicating emotion: The role of prosodic features.,”
Robert W Frick, · 1985
Earlier work this paper cites.
The emotions
Robert Plutchik, · 1991
Earlier work this paper cites.
Are there basic emotions?
Paul Ekman, · 1992
Earlier work this paper cites.
“The frequency range of the voice fundamental in the speech of male and female adults,”
Hartmut Traunmüller and Anders Eriksson, · 1995
Earlier work this paper cites.
Emotions and multilingualism
Aneta Pavlenko, · 2005
Earlier work this paper cites.
“Iemocap: Interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan, · 2008
Earlier work this paper cites.
“Crema-d: Crowd-sourced emotional multimodal actors dataset,”
Houwei Cao, David G Cooper, Michael K Keutmann, Ruben C Gur, Ani Nenkova, and Ragini Verma, · 2014
Earlier work this paper cites.
“Attending at a low intensity increases impulsivity in an auditory sustained attention to response task,”
Hettie Roebuck, Kun Guo, and Patrick Bourke, · 2015
Earlier work this paper cites.
“librosa: Audio and music signal analysis in python,”
Brian McFee, Colin Raffel, Dawen Liang, Daniel P Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto, · 2015
Earlier work this paper cites.
“Speech rate in parkinson’s disease: A controlled study,”
F Martínez-Sánchez, JJG Meilán, J Carro, C Gómez Íñiguez, L Millian-Morell, IM Pujante Valverde, T López-Alburquerque, and DE López, · 2016
Cited alongside, same era.
“Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,”
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency, · 2018
Cited alongside, same era.
“Meld: A multimodal multi-party dataset for emotion recognition in conversations,”
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea, · 2018
Cited alongside, same era.
“The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,”
Steven R Livingstone and Frank A Russo, · 2018
Cited alongside, same era.
“Detecting gender differences in perception of emotion in crowdsourced data,”
“Clotho: an audio captioning dataset,”
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen, · 2020
Later among the works it cites.
“What is the ground truth? reliability of multi-annotator data for audio tagging,”
Irene Martín-Morató and Annamaria Mesaros, · 2021
Later among the works it cites.
“A proposal for multimodal emotion recognition using aural transformers and action units on ravdess dataset,”
Cristina Luna-Jiménez, Ricardo Kleinlein, David Griol, Zoraida Callejas, Juan M Montero, and Fernando Fernández-Martínez, · 2021
Later among the works it cites.
“Clap: Learning audio concepts from natural language supervision,”
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang, · 2022
Later among the works it cites.
“Audio retrieval with wavtext5k and clap training,”
Soham Deshmukh, Benjamin Elizalde, and Huaming Wang, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shahan Ali Memon, Hira Dhamyal, Oren Wright, Daniel Justice, Vijaykumar Palat, William Boler, Bhiksha Raj, and Rita Singh, · 2019
Cited alongside, same era.
“Huggingface’s transformers: State-of-the-art natural language processing,”
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al., · 2019
Cited alongside, same era.
“AudioCaps: Generating Captions for Audios in The Wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
Cited alongside, same era.
“Improving content-based audio retrieval by vocal imitation feedback,”
Bongjun Kim and Bryan Pardo, · 2019
Cited alongside, same era.
“The phonetic bases of vocal expressed emotion: Natural versus acted,”
Hira Dhamyal, Shahan Ali Memon, Bhiksha Raj, and Rita Singh, · 2020
Cited alongside, same era.
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley, · 2020
Cited alongside, same era.
“Praat,” https://www.fon.hum.uva.nl/praat
Cited in the paper.
“Positional encoding for capturing modality specific cadence for emotion detection,”
Hira Dhamyal, Bhiksha Raj, and Rita Singh, · 2022
Later among the works it cites.
“Fsd50k: An open dataset of human-labeled sound events,”
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra, · 2022
Later among the works it cites.
“Research on emotional semantic retrieval of attention mechanism oriented to audio-visual synesthesia,”
Weixing Wang, Qianqian Li, Jingwen Xie, Ningfeng Hu, Ziao Wang, and Ning Zhang, · 2023
Closest in time.
“Emomv: Affective music-video correspondence learning datasets for classification and retrieval,”
Ha Thi Phuong Thao, Gemma Roig, and Dorien Herremans, · 2023
Closest in time.