Fetching the paper…
Reading the bibliography…
Speech Emotion Recognition (SER) typically relies on utterance-level solutions.
“Universals and cultural differences in facial expressions of emotion.,”
Paul Ekman, · 1971
Earlier work this paper cites.
“Unmasking the face: A guide to recognizing emotions from facial clues.,”
Paul Ekman and Wallace V Friesen, · 1975
Earlier work this paper cites.
“A circumplex model of affect.,”
James A Russell, · 1980
Earlier work this paper cites.
“Pleasure-arousal-dominance: A general framework for describing and measuring individual differences in temperament,”
Albert Mehrabian, · 1996
Earlier work this paper cites.
“A database of german emotional speech.,”
Felix Burkhardt, Astrid Paeschke, Miriam Rolfes, Walter F Sendlmeier, Benjamin Weiss, et al., · 2005
Earlier work this paper cites.
“Frame vs. turn-level: emotion recognition from speech considering static and dynamic processing,”
Bogdan Vlasenko, Björn Schuller, Andreas Wendemuth, and Gerhard Rigoll, · 2007
Earlier work this paper cites.
“Conscious emotional experience emerges as a function of multilevel, appraisal-driven response synchronization,”
Didier Grandjean, David Sander, and Klaus R Scherer, · 2008
Earlier work this paper cites.
“Abandoning emotion classes-towards continuous emotion recognition with modelling of long-range dependencies,”
Martin Wöllmer, Florian Eyben, Stephan Reiter, Björn Schuller, Cate Cox, Ellen Douglas-Cowie, and Roddy Cowie, · 2008
Earlier work this paper cites.
“Iemocap: Interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan, · 2008
Earlier work this paper cites.
“Combining long short-term memory and dynamic bayesian networks for incremental emotion-sensitive artificial listening,”
Martin Wöllmer, Björn Schuller, Florian Eyben, and Gerhard Rigoll, · 2010
Earlier work this paper cites.
“Emotion representation, analysis and synthesis in continuous space: A survey,”
Hatice Gunes, Björn Schuller, Maja Pantic, and Roddy Cowie, · 2011
Earlier work this paper cites.
“The semaine database: Annotated multimodal records of emotionally colored conversations between a person and a limited agent,”
Gary McKeown, Michel Valstar, Roddy Cowie, Maja Pantic, and Marc Schroder, · 2011
Earlier work this paper cites.
“Introducing the recola multimodal corpus of remote collaborative and affective interactions,”
Fabien Ringeval, Andreas Sonderegger, Juergen Sauer, and Denis Lalanne, · 2013
Earlier work this paper cites.
“Speech emotion recognition using deep neural network and extreme learning machine,”
Kun Han, Dong Yu, and Ivan Tashev, · 2014
Cited alongside, same era.
“Towards real-time speech emotion recognition using deep neural networks,”
Haytham M Fayek, Margaret Lech, and Lawrence Cavedon, · 2015
Cited alongside, same era.
“Emotion recognition from speech with recurrent neural networks,”
Vladimir Chernykh and Pavel Prikhodko, · 2017
Cited alongside, same era.
“Automatic speech emotion recognition using recurrent neural networks with local attention,”
Seyedmahdad Mirsamadi, Emad Barsoum, and Cha Zhang, · 2017
Cited alongside, same era.
“Evaluating deep learning architectures for speech emotion recognition,”
Haytham M Fayek, Margaret Lech, and Lawrence Cavedon, · 2017
Cited alongside, same era.
“Sewa db: A rich database for audio-visual emotion and sentiment research in the wild,”
Jean Kossaifi, Robert Walecki, Yannis Panagakis, Jie Shen, Maximilian Schmitt, Fabien Ringeval, Jing Han, Vedhas Pandit, Antoine Toisoul, Björn Schuller, et al., · 2019
Later among the works it cites.
“Speech emotion recognition: Emotional models, databases, features, preprocessing methods, supporting modalities, and classifiers,”
Mehmet Berkehan Akçay and Kaya Oğuz, · 2020
Later among the works it cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Later among the works it cites.
“A comprehensive review of speech emotion recognition systems,”
Taiba Majid Wani, Teddy Surya Gunawan, Syed Asif Ahmad Qadri, Mira Kartiwi, and Eliathamby Ambikairajah, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“An attention pooling based representation learning method for speech emotion recognition,”
Pengcheng Li, Yan Song, Ian Vince McLoughlin, Wu Guo, and Li-Rong Dai, · 2018
Cited alongside, same era.
“The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,”
Steven R Livingstone and Frank A Russo, · 2018
Cited alongside, same era.
Adaeze Adigwe, Noé Tits, Kevin El Haddad, Sarah Ostadabbas, and Thierry Dutoit, · 2018
Cited alongside, same era.
“An open source emotional speech corpus for human robot interaction applications,”
Jesin James, Li Tian, and Catherine Watson, · 2018
Cited alongside, same era.
“Speech emotion recognition using deep learning techniques: A review,”
Ruhul Amin Khalil, Edward Jones, Mohammad Inayatullah Babar, Tariqullah Jan, Mohammad Haseeb Zafar, and Thamer Alhussain, · 2019
Cited alongside, same era.
“Speech emotion recognition using capsule networks,”
Xixin Wu, Songxiang Liu, Yuewen Cao, Xu Li, Jianwei Yu, Dongyang Dai, Xi Ma, Shoukang Hu, Zhiyong Wu, Xunying Liu, et al., · 2019
Cited alongside, same era.
Yingzhi Wang, Abdelmoumene Boumadane, and Abdelwahab Heba, · 2021
Later among the works it cites.
“Speechbrain: A general-purpose speech toolkit,”
Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, et al., · 2021
Later among the works it cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Later among the works it cites.
“Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,”
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li, · 2021
Later among the works it cites.
“Speaker normalization for self-supervised speech emotion recognition,”
Itai Gat, Hagai Aronowitz, Weizhong Zhu, Edmilson Morais, and Ron Hoory, · 2022
Later among the works it cites.
“On fine-grained temporal emotion recognition in video: How to trade off recognition accuracy with annotation complexity?,”
Tianyi Zhang, · 2022
Later among the works it cites.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2022
Later among the works it cites.
“Boosting self-supervised embeddings for speech enhancement,”
Kuo-Hsuan Hung, Szu-wei Fu, Huan-Hsin Tseng, Hsin-Tien Chiang, Yu Tsao, and Chii-Wann Lin, · 2022
Later among the works it cites.