Fetching the paper…
Reading the bibliography…
We propose emotion2vec, a universal speech emotion representation model.
A database of German emotional speech
Felix Burkhardt, Astrid Paeschke, Miriam Rolfes, Walter F Sendlmeier, Benjamin Weiss, et al. 2005 · 2005
Earlier work this paper cites.
IEMOCAP: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008 · 2008
Earlier work this paper cites.
EMOVO corpus: an Italian emotional speech database
Giovanni Costantini, Iacopo Iaderola, Andrea Paoloni, and Massimiliano Todisco. 2014 · 2014
Earlier work this paper cites.
Surrey audio-visual expressed emotion (SAVEE) database
Philip Jackson and SJUoSG Haq. 2014 · 2014
Earlier work this paper cites.
MOSI: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos
Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency. 2016 · 2016
Earlier work this paper cites.
Speech-based emotion recognition and next reaction prediction
Fatemeh Noroozi, Neda Akrami, and Gholamreza Anbarjafari. 2017 · 2017
Earlier work this paper cites.
A Canadian French emotional speech dataset
Philippe Gournay, Olivier Lahaie, and Roch Lefebvre. 2018 · 2018
Earlier work this paper cites.
Cross lingual speech emotion recognition: Urdu vs. western languages
Siddique Latif, Adnan Qayyum, Muhammad Usman, and Junaid Qadir. 2018 · 2018
Earlier work this paper cites.
The Ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English
Steven R Livingstone and Frank A Russo. 2018 · 2018
Earlier work this paper cites.
UMAP: Uniform manifold approximation and projection for dimension reduction
Leland McInnes, John Healy, and James Melville. 2018 · 2018
Earlier work this paper cites.
Speech emotion recognition for performance interaction
Nikolaos Vryzas, Rigas Kotsakis, Aikaterini Liatsou, Charalampos A Dimoulas, and George Kalliris. 2018 · 2018
Earlier work this paper cites.
Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018 · 2018
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
ShEMO: a large-scale validated database for Persian speech emotion detection
Omid Mohamad Nezami, Paria Jamshid Lou, and Mansoureh Karami. 2019 · 2019
Earlier work this paper cites.
MELD: A multimodal multi-party dataset for emotion recognition in conversations
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019 · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2020
Cited alongside, same era.
Bootstrap your own latent: a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. 2020 · 2020
Cited alongside, same era.
data2vec: A general framework for self-supervised learning in speech, vision and language
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. 2022 · 2022
Later among the works it cites.
WavLM: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al. 2022 · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. 2022 · 2022
Later among the works it cites.
Exploration of a self-supervised speech model: A study on emotional corpora
Yuanchao Li, Yumnah Mohamied, Peter Bell, and Catherine Lai. 2022 · 2022
Later among the works it cites.
Speech emotion recognition using self-supervised features
Edmilson Morais, Ron Hoory, Weizhong Zhu, Itai Gat, Matheus Damasceno, and Hagai Aronowitz. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020 · 2020
Cited alongside, same era.
The MSP-conversation corpus
Luz Martinez-Lucas, Mohammed Abdelwahab, and Carlos Busso. 2020 · 2020
Cited alongside, same era.
Dimensional emotion prediction based on interactive context in conversation
Xiaohan Shi, Sixia Li, and Jianwu Dang. 2020 · 2020
Cited alongside, same era.
MEAD: A large-scale audio-visual dataset for emotional talking-face generation
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. 2020 · 2020
Cited alongside, same era.
BEiT: BERT pre-training of image Transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. 2021 · 2021
Cited alongside, same era.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Cited alongside, same era.
Comparison and analysis of deep audio embeddings for music emotion recognition
Eunjeong Koh and Shlomo Dubnov. 2021 · 2021
Cited alongside, same era.
Supervision-guided codebooks for masked prediction in speech pre-training
Chengyi Wang, Yiming Wang, Yu Wu, Sanyuan Chen, Jinyu Li, Shujie Liu, and Furu Wei. 2022 · 2022
Later among the works it cites.
M3ED: Multi-modal multi-scene multi-label emotional dialogue database
Jinming Zhao, Tenggan Zhang, Jingwen Hu, Yuchen Liu, Qin Jin, Xinchao Wang, and Haizhou Li. 2022 · 2022
Later among the works it cites.
Efficient self-supervised learning with contextualized target representations for vision, speech and language
Alexei Baevski, Arun Babu, Wei-Ning Hsu, and Michael Auli. 2023 · 2023
Closest in time.
Exploring wav2vec 2.0 fine tuning for improved speech emotion recognition
Li-Wei Chen and Alexander Rudnicky. 2023 · 2023
Closest in time.
Self-supervised learning with cluster-aware-DINO for high-performance robust speaker verification
Bing Han, Zhengyang Chen, and Yanmin Qian. 2023 · 2023
Closest in time.
Towards paralinguistic-only speech representations for end-to-end speech emotion recognition
George Ioannides, Michael Owen, Andrew Fletcher, Viktor Rozgic, and Chao Wang. 2023 · 2023
Closest in time.
MER 2023: Multi-label learning, modality robustness, and semi-supervised learning
Zheng Lian, Haiyang Sun, Licai Sun, Kang Chen, Mngyu Xu, Kexin Wang, Ke Xu, Yu He, Ying Li, Jinming Zhao, et al. 2023 · 2023
Closest in time.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023 · 2023
Closest in time.
A vector quantized masked autoencoder for speech emotion recognition
Samir Sadok, Simon Leglaive, and Renaud Séguier. 2023 · 2023
Closest in time.
Emotion awareness in multi-utterance turn for improving emotion prediction in multi-speaker conversation
Xiaohan Shi, Xingfeng Li, and Tomoki Toda. 2023 · 2023
Closest in time.
Temporal modeling matters: A novel temporal emotional modeling approach for speech emotion recognition
Jiaxin Ye, Xin-Cheng Wen, Yujie Wei, Yong Xu, Kunhong Liu, and Hongming Shan. 2023 · 2023
Closest in time.