Fetching the paper…
Reading the bibliography…
Speech emotion recognition (SER) is a pivotal technology for human-computer interaction systems.
An unsupervised autoregressive model for speech representation learning
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass. 2019 · 1904
Earlier work this paper cites.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019 · 1904
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli. 2019 · 1910
Earlier work this paper cites.
Juri Opitz and Sebastian Burst. 2019 · 1911
Earlier work this paper cites.
Vector-quantized autoregressive predictive coding
Yu-An Chung, Hao Tang, and James Glass. 2020 · 2005
Earlier work this paper cites.
IEMOCAP: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan. 2008 · 2008
Earlier work this paper cites.
Non-autoregressive predictive coding for learning speech representations from local dependencies
Alexander H Liu, Yu-An Chung, and James Glass. 2020a · 2011
Earlier work this paper cites.
The Shifting Meaning of Happiness
Cassie Mogilner, Sepandar D. Kamvar, and Jennifer Aaker. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine Learning in Python
Fabian Pedregosa et al. 2011 · 2011
Earlier work this paper cites.
Decoar 2.0: Deep contextualized acoustic representations with vector quantization
Shaoshi Ling and Yuzong Liu. 2020 · 2012
Earlier work this paper cites.
CREMA-D: Crowd-Sourced Emotional Multimodal Actors Dataset
Houwei Cao, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova, and Ragini Verma. 2014 · 2014
Earlier work this paper cites.
Increasing the Reliability of Crowdsourcing Evaluations Using Online Quality Assessment
Alec Burmania, Srinivas Parthasarathy, and Carlos Busso. 2016 · 2016
Earlier work this paper cites.
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Earlier work this paper cites.
MSP-IMPROV: An Acted Corpus of Dyadic Interactions to Study Emotion Perception
Carlos Busso, Srinivas Parthasarathy, Alec Burmania, Mohammed AbdelWahab, Najmeh Sadoughi, and Emily Mower Provost. 2017 · 2017
Earlier work this paper cites.
NNIME: The NTHU-NTUA Chinese interactive multimodal emotion corpus
Huang-Cheng Chou, Wei-Cheng Lin, Lien-Chiang Chang, Chyi-Chang Li, Hsi-Pin Ma, and Chi-Chun Lee. 2017 · 2017
Cited alongside, same era.
Self-report captures 27 distinct categories of emotion bridged by continuous gradients
Alan S. Cowen and Dacher Keltner. 2017 · 2017
Cited alongside, same era.
Formulating Emotion Perception as a Probabilistic Model with Application to Categorical Emotion Classification
Reza Lotfian and Carlos Busso. 2017 · 2017
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Cited alongside, same era.
Class-Balanced Loss Based on Effective Number of Samples
Y. Cui, M. Jia, T.-Y. Lin, Y. Song, and S. Belongie. 2019 · 2019
Cited alongside, same era.
Superb: Speech processing universal performance benchmark
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al. 2021 · 2021
Later among the works it cites.
Evaluating Self-Supervised Speech Representations for Speech Emotion Recognition
Bagus Tris Atmaja and Akira Sasou. 2022 · 2022
Later among the works it cites.
Data2vec: A general framework for self-supervised learning in speech, vision and language
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. 2022 · 2022
Later among the works it cites.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al. 2022 · 2022
Later among the works it cites.
Exploiting Annotators’ Typed Description of Emotion Perception to Maximize Utilization of Ratings for Speech Emotion Recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
Building Naturalistic Emotionally Balanced Speech Corpus by Retrieving Emotional Speech From Existing Podcast Recordings
Reza Lotfian and Carlos Busso. 2019 · 2019
Cited alongside, same era.
No Sample Left Behind: Towards a Comprehensive Evaluation of Speech Emotion Recognition Systems
Pablo Riera, Luciana Ferrer, Agustín Gravano, and Lara Gauder. 2019 · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Cited alongside, same era.
Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders
Andy T Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee. 2020b · 2020
Cited alongside, same era.
Xls-r: Self-supervised cross-lingual speech representation learning at scale
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, et al. 2021 · 2021
Cited alongside, same era.
Semantic Space Theory: A Computational Approach to Emotion
Alan S. Cowen and Dacher Keltner. 2021 · 2021
Cited alongside, same era.
Huang-Cheng Chou, Wei-Cheng Lin, Chi-Chun Lee, and Carlos Busso. 2022 · 2022
Later among the works it cites.
Superb@ slt 2022: Challenge on generalization and efficiency of self-supervised speech representation learning
Tzu-hsun Feng, Annie Dong, Ching-Feng Yeh, Shu-wen Yang, Tzu-Quan Lin, Jiatong Shi, Kai-Wei Chang, Zili Huang, Haibin Wu, Xuankai Chang, et al. 2023 · 2022
Later among the works it cites.
Exploration of a Self-Supervised Speech Model: A Study on Emotional Corpora
Yuanchao Li, Yumnah Mohamied, Peter Bell, and Catherine Lai. 2023 · 2022
Later among the works it cites.
Hsiang-Sheng Tsai, Heng-Jui Chang, Wen-Chin Huang, Zili Huang, Kushal Lakhotia, Shu-wen Yang, Shuyan Dong, Andy T Liu, Cheng-I Jeff Lai, Jiatong Shi, et al. 2022 · 2022
Later among the works it cites.
Josh Achiam et al. 2023 · 2023
Later among the works it cites.
Designing and Evaluating Speech Emotion Recognition Systems: A Reality Check Case Study with IEMOCAP
Nikolaos Antoniou, Athanasios Katsamanis, Theodoros Giannakopoulos, and Shrikanth Narayanan. 2023 · 2023
Later among the works it cites.
Kiana Kheiri and Hamid Karimi. 2023 · 2023
Later among the works it cites.
An Intelligent Infrastructure Toward Large Scale Naturalistic Affective Speech Corpora Collection
Shreya G Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Lucas Goncalves, Ya-Tse Wu, Ali N Salman, Carlos Busso, and Chi-Chun Lee. 2023 · 2023
Later among the works it cites.
Dawn of the Transformer Era in Speech Emotion Recognition: Closing the Valence Gap
Johannes Wagner, Andreas Triantafyllopoulos, Hagen Wierstorf, Maximilian Schmitt, Felix Burkhardt, Florian Eyben, and Björn W. Schuller. 2023 · 2023
Later among the works it cites.