Fetching the paper…
Reading the bibliography…
Recent advances in transformer-based architectures which are pre-trained in self-supervised manner have shown great promise in several machine learning tasks.
“Evidence for a three-factor theory of emotions”
James Russell and Albert Mehrabian · 1977
Earlier work this paper cites.
“A concordance correlation coefficient to evaluate reproducibility”
Lawrence-Kuei Lin · 1989
Earlier work this paper cites.
“An argument for basic emotions”
Paul Ekman · 1992
Earlier work this paper cites.
“SHEEP, GOATS, LAMBS and WOLVES: a statistical analysis of speaker performance in the NIST 1998 speaker recognition evaluation”
George Doddington, Walter Liggett, Alvin Martin, Mark Przybocki and Douglas. Reynolds · 1998
Earlier work this paper cites.
“IEMOCAP: Interactive emotional dyadic motion capture database”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette Chang, Sungbok Lee and Shrikanth Narayanan · 2008
Earlier work this paper cites.
“Visualizing Data using t-SNE”
Laurens van Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
“A survey of affect recognition methods: Audio, visual, and spontaneous expressions”
Zhihong Zeng, Maja Pantic, Glenn. Roisman and Thomas. Huang · 2009
Earlier work this paper cites.
“Affect detection: An interdisciplinary review of models, methods, and their applications”
Rafael Calvo and Sidney D’Mello · 2010
Earlier work this paper cites.
“Why does unsupervised pre-training help deep learning?”
Dumitru Erhan, Aaron Courville, Yoshua Bengio and Pascal Vincent · 2010
Earlier work this paper cites.
“Are they different? Affect, feeling, emotion, sentiment, and opinion detection in text”
Myriam Munezero, Calkin Montero, Erkki Sutinen and John Pajunen · 2014
Earlier work this paper cites.
“ESC: Dataset for Environmental Sound Classification”
Karol. Piczak · 2015
Earlier work this paper cites.
“Adieu Features? End-to-End Speech Emotion Recognition using a Deep Convolutional Recurrent Network”
George Trigeorgis, Fabien Ringeval, Raymond Brückner, Erik Marchi, Mihalis Nicolaou, Björn Schuller and Stefanos Zafeiriou · 2016
Earlier work this paper cites.
“Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages”
Amir Zadeh, Rowan Zellers, Eli Pincus and Louis-Philippe Morency · 2016
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“The Perception of Emotions in Noisified Nonsense Speech”
Emilia Parada-Cabaleiro, Alice Baird, Anton Batliner, Nicholas Cummins, Simone Hantke and Björn Schuller · 2017
Earlier work this paper cites.
“Speech Emotion Recognition: Two Decades in a Nutshell, Benchmarks, and Ongoing Trends”
Björn Schuller · 2018
Earlier work this paper cites.
“Polarity and Intensity: the Two Aspects of Sentiment Analysis”
Leimin Tian, Catherine Lai and Johanna Moore · 2018
Earlier work this paper cites.
“AVEC 2018 workshop and challenge: Bipolar disorder and cross-cultural affect recognition”
Fabien Ringeval, Björn Schuller, Michel Valstar, Roddy Cowie, Heysem Kaya, Maximilian Schmitt, Shahin Amiriparian, Nicholas Cummins, Denis Lalanne and Adrien Michaud · 2018
Earlier work this paper cites.
“The measure and mismeasure of fairness: A critical review of fair machine learning”
Sam Corbett-Davies and Sharad Goel · 2018
Earlier work this paper cites.
“Personalized machine learning for robot perception of affect and engagement in autism therapy”
Ognjen Rudovic, Jaeryoung Lee, Miles Dai, Björn Schuller and Rosalind Picard · 2018
Earlier work this paper cites.
Felix Burkhardt, Benjamin Weiss, Florian Eyben, Jun Deng and Björn Schuller · 2018
Earlier work this paper cites.
“Multi-Modal Learning for Speech Emotion Recognition: An Analysis and Comparison of ASR Outputs with Ground Truth Transcription”
Saurabh Sahu, Vikramjit Mitra, Nadee Seneviratne and Carol Espy-Wilson · 2019
Earlier work this paper cites.
“Robust Speech Emotion Recognition under Different Encoding Conditions”
Christopher Oates, Andreas Triantafyllopoulos, Ingmar Steiner and Björn. Schuller · 2019
Earlier work this paper cites.
“Towards Robust Speech Emotion Recognition using Deep Residual Networks for Speech Enhancement”
Andreas Triantafyllopoulos, Gil Keren, Johannes Wagner, Ingmar Steiner and Björn. Schuller · 2019
Cited alongside, same era.
“Building Naturalistic Emotionally Balanced Speech Corpus by Retrieving Emotional Speech from Existing Podcast Recordings”
Reza Lotfian and Carlos Busso · 2019
Cited alongside, same era.
“Gender De-Biasing in Speech Emotion Recognition”
Cristina Gorrostieta, Reza Lotfian, Kye Taylor, Richard Brutti and John Kane · 2019
Cited alongside, same era.
“A General Framework for Fair Regression”
Jack. Fitzsimons, AbdulRahman Ali, Michael. Osborne and Stephen. Roberts · 2019
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Cited alongside, same era.
“Using Large Pre-Trained Models with Cross-Modal Attention for Multi-Modal Emotion Recognition”
D.. Krishna · 2021
Later among the works it cites.
“The Role of Phonetic Units in Speech Emotion Recognition”
Jiahong Yuan, Xingyu Cai, Renjie Zheng, Liang Huang and Kenneth Church · 2021
Later among the works it cites.
“SUPERB: Speech processing Universal PERformance Benchmark”
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Lai, Kushal Lakhotia, Yist. Lin, Andy. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko-tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed and Hung-yi Lee · 2021
Later among the works it cites.
“Emotion Recognition from Speech Using wav2vec 2.0 Embeddings”
Leonardo Pepino, Pablo Riera and Luciana Ferrer · 2021
Later among the works it cites.
“Exploring Wav2vec 2.0 fine-tuning for improved speech emotion recognition”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anton Batliner, Simone Hantke and Bjoern Schuller · 2020
Cited alongside, same era.
“A Simple Framework for Contrastive Learning of Visual Representations”
Ting Chen, Simon Kornblith, Mohammad Norouzi and Geoffrey Hinton · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed and Michael Auli · 2020
Cited alongside, same era.
“PANNs: Large-scale pretrained audio neural networks for audio pattern recognition”
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang and Mark Plumbley · 2020
Cited alongside, same era.
“Machine Learning Testing: Survey, Landscapes and Horizons”
Jie. Zhang, Mark Harman, Lei Ma and Yang Liu · 2020
Cited alongside, same era.
“Fusion approaches for emotion recognition from speech using acoustic and text-based features”
Leonardo Pepino, Pablo Riera, Luciana Ferrer and Agustı́n Gravano · 2020
Cited alongside, same era.
“Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders”
Andy Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu and Hung-yi Lee · 2020
Cited alongside, same era.
Li-Wei Chen and Alexander Rudnicky · 2021
Later among the works it cites.
“Multimodal Emotion Recognition with High-level Speech and Text Features”
Mariana Makiuchi, Kuniaki Uto and Koichi Shinoda · 2021
Later among the works it cites.
Sundararajan Srinivasan, Zhaocheng Huang and Katrin Kirchhoff · 2021
Later among the works it cites.
“Contrastive Unsupervised Learning for Speech Emotion Recognition”
Mao Li, Bo Yang, Joshua Levy, Andreas Stolcke, Viktor Rozgic, Spyros Matsoukas, Constantinos Papayiannis, Daniel Bone and Chao Wang · 2021
Later among the works it cites.
Mimansa Jaiswal and Emily Provost · 2021
Later among the works it cites.
“CopyPaste: An augmentation method for speech emotion recognition”
Raghavendra Pappagari, Jesús Villalba, Piotr Żelasko, Laureano Moro-Velazquez and Najim Dehak · 2021
Later among the works it cites.
“AequeVox: Automated Fairness Testing of Speech Recognition Systems”
Sai Rajan, Sakshi Udeshi and Sudipta Chattopadhyay · 2021
Later among the works it cites.
“Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training”
Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski, Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Jacob Kahn, Ann Lee, Ronan Collobert, Gabriel Synnaeve and Michael Auli · 2021
Later among the works it cites.
“VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation”
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino and Emmanuel Dupoux · 2021
Later among the works it cites.
“XLS-R: Self-supervised cross-lingual speech representation learning at scale”
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, Alexei Baevski, Alexis Conneau and Michael Auli · 2021
Later among the works it cites.
“The Role of Task and Acoustic Similarity in Audio Transfer Learning: Insights from the Speech Emotion Recognition Case”
Andreas Triantafyllopoulos and Björn Schuller · 2021
Later among the works it cites.
“Deep speaker conditioning for speech emotion recognition”
Andreas Triantafyllopoulos, Shuo Liu and Björn Schuller · 2021
Later among the works it cites.
“Survey on bimodal speech emotion recognition from acoustic and linguistic information fusion”
Bagus Atmaja, Akira Sasou and Masato Akagi · 2022
Closest in time.
“A survey on vision transformer”
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, Zhaohui Yang, Yiman Zhang and Dacheng Tao · 2022
Closest in time.
“Model for Dimensional Speech Emotion Recognition based on Wav2vec 2.0”
Johannes Wagner, Andreas Triantafyllopoulos, Hagen Wierstorf, Maximilian Schmitt, Florian Eyben and Björn. Schuller · 2022
Closest in time.
“Probing Speech Emotion Recognition Transformers for Linguistic Knowledge”
Andreas Triantafyllopoulos, Johannes Wagner, Hagen Wierstorf, Maximilian Schmitt, Uwe Reichel, Florian Eyben, Felix Burkhardt and Björn Schuller · 2022
Closest in time.
Kusha Sridhar and Carlos Busso · 2022
Closest in time.
“Gender annotations for Multimodal Opinion-level Sentiment Intensity dataset (MOSI)”
Hagen Wierstorf · 2023
Closest in time.