Fetching the paper…
Reading the bibliography…
Machine learning models for speech emotion recognition (SER) can be trained for different tasks and are usually evaluated based on a few available datasets per task.
“TIMIT Acoustic-Phonetic Continuous Speech Corpus”
J.. Garofolo, L.. Lamel, W.. Fisher, J.. Fiscus, D.. Pallett and N.. Dahlgren · 1983
Earlier work this paper cites.
“Switchboard-1 Release 2 LDC97S62”
John Godfrey and Edward Holliman · 1993
Earlier work this paper cites.
“Design, recording and verification of a danish emotional speech database”
Inger. Engberg, Anya Hansen, Ove Andersen and Paul Dalsgaard · 1997
Earlier work this paper cites.
“A new metric for probability distributions”
Dominik Endres and Johannes Schindelin · 2003
Earlier work this paper cites.
“Fisher English training speech part 1 transcripts”
Christopher Cieri, David Graff, Owen Kimball, Dave Miller and Kevin Walker · 2004
Earlier work this paper cites.
“A database of German emotional speech.”
Felix Burkhardt, Astrid Paeschke, Miriam Rolfes, Walter Sendlmeier and Benjamin Weiss · 2005
Earlier work this paper cites.
“Evaluation of speech dereverberation algorithms using the MARDY database”
Jimi Wen, Nikolay Gaubitch, Emanuel Habets, Tony Myatt and Patrick Naylor · 2006
Earlier work this paper cites.
“An approach to software testing of machine learning applications”
Christian Murphy, Gail Kaiser and Marta Arias · 2007
Earlier work this paper cites.
“The World of Emotions is not Two-Dimensional”
Johnny.J. Fontaine, Klaus. Scherer, Etienne. Roesch and Phoebe. Ellsworth · 2007
Earlier work this paper cites.
“IEMOCAP: Interactive emotional dyadic motion capture database”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette Chang, Sungbok Lee and Shrikanth Narayanan · 2008
Earlier work this paper cites.
“A Binaural Room Impulse Response Database for the Evaluation of Dereverberation Algorithms”
Marco Jeub, Magnus Schäfer and Peter Vary · 2009
Earlier work this paper cites.
“Mapping discrete emotions into the dimensional space: An empirical approach”
Holger Hoffmann, Andreas Scheck, Timo Schuster, Steffen Walter, Kerstin Limbrecht, Harald Traue and Henrik Kessler · 2012
Earlier work this paper cites.
“Intriguing properties of neural networks”
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow and Rob Fergus · 2013
Earlier work this paper cites.
“Crema-d: Crowd-sourced emotional multimodal actors dataset”
Houwei Cao, David Cooper, Michael Keutmann, Ruben Gur, Ani Nenkova and Ragini Verma · 2014
Earlier work this paper cites.
“EMOVO corpus: an Italian emotional speech database”
Giovanni Costantini, Iacopo Iaderola, Andrea Paoloni and Massimiliano Todisco · 2014
Earlier work this paper cites.
“Speech recognition and keyword spotting for low-resource languages: Babel project research at cued”
Mark Gales, Kate Knill, Anton Ragni and Shakti Rath · 2014
Earlier work this paper cites.
“Speech accent archive” Retrieved from http://accent.gmu.edu , 2015
Steven Weinberger · 2015
Earlier work this paper cites.
“MUSAN: A Music, Speech, and Noise Corpus” arXiv:1510.08484v1, 2015
David Snyder, Guoguo Chen and Daniel Povey · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books”
Vassil Panayotov, Guoguo Chen, Daniel Povey and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
“Distilling the knowledge in a neural network” NIPS 2014 Deep Learning Workshop
Geoffrey Hinton, Oriol Vinyals and Jeff Dean · 2015
Earlier work this paper cites.
“Introduction to software testing”
Paul Ammann and Jeff Offutt · 2016
Earlier work this paper cites.
“Mapping emotion terms into affective space”
Christelle Gillioz, Johnny Fontaine, Cristina Soriano and Klaus Scherer · 2016
Earlier work this paper cites.
“Measuring neural net robustness with constraints”
Osbert Bastani, Yani Ioannou, Leonidas Lampropoulos, Dimitrios Vytiniotis, Aditya Nori and Antonio Criminisi · 2016
Earlier work this paper cites.
“Affect representation and recognition in 3D continuous valence–arousal–dominance space”
Gyanendra Verma and Uma Tiwary · 2017
Earlier work this paper cites.
“CAST a database: Rapid targeted large-scale big data acquisition via small-world modelling of social media platforms”
Shahin Amiriparian, Sergey Pugachevskiy, Nicholas Cummins, Simone Hantke, Jouni Pohjalainen, Gil Keren and Björn. Schuller · 2017
Earlier work this paper cites.
“Audio Set: An ontology and human-labeled dataset for audio events”
Jort. Gemmeke, Daniel.. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Moore, Manoj Plakal and Marvin Ritter · 2017
Cited alongside, same era.
“Deeptest: Automated testing of deep-neural-network-driven autonomous cars”
Yuchi Tian, Kexin Pei, Suman Jana and Baishakhi Ray · 2018
Cited alongside, same era.
Rachel Bellamy, Kuntal Dey, Michael Hind, Samuel Hoffman, Stephanie Houde, Kalapriya Kannan, Pranay Lohia, Jacquelyn Martino, Sameep Mehta and Aleksandra Mojsilovic · 2018
Cited alongside, same era.
“EmotionLines: An Emotion Corpus of Multi-Party Conversations”
Chao-Chun Hsu, Sheng-Yeh Chen, Chuan-Chun Kuo, Ting-Hao Huang and Lun-Wei Ku · 2018
Cited alongside, same era.
“The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English”
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition”
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang and Mark Plumbley · 2020
Later among the works it cites.
“On the opportunities and risks of foundation models”
Rishi Bommasani, Drew Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael Bernstein, Jeannette Bohg, Antoine Bosselut and Emma Brunskill · 2021
Later among the works it cites.
Mimansa Jaiswal and Emily Provost · 2021
Later among the works it cites.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Tsai, Kushal Lakhotia, Ruslan Salakhutdinov and Abdelrahman Mohamed · 2021
Later among the works it cites.
“A Survey on Bias and Fairness in Machine Learning”
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman and Aram Galstyan · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Steven Livingstone and Frank Russo · 2018
Cited alongside, same era.
“The measure and mismeasure of fairness: A critical review of fair machine learning”
Sam Corbett-Davies and Sharad Goel · 2018
Cited alongside, same era.
“Model Cards for Model Reporting”
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Raji and Timnit Gebru · 2019
Cited alongside, same era.
“Machine Learning Testing: Survey, Landscapes and Horizons”
Jie. Zhang, Mark Harman, Lei Ma and Yang Liu · 2019
Cited alongside, same era.
“Affective and behavioural computing: Lessons learnt from the first computational paralinguistics challenge”
Björn Schuller, Felix Weninger, Yue Zhang, Fabien Ringeval, Anton Batliner, Stefan Steidl, Florian Eyben, Erik Marchi, Alessandro Vinciarelli and Klaus Scherer · 2019
Cited alongside, same era.
“Towards testing of deep learning systems with training set reduction”
Helge Spieker and Arnaud Gotlieb · 2019
Cited alongside, same era.
“MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations”
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria and Rada Mihalcea · 2019
Cited alongside, same era.
“Building Naturalistic Emotionally Balanced Speech Corpus by Retrieving Emotional Speech from Existing Podcast Recordings”
Reza Lotfian and Carlos Busso · 2019
Cited alongside, same era.
Later among the works it cites.
“coqui-ai/TTS, version 0.6.1”, 2021
Gölge Eren and TheTTS Team · 2021
Later among the works it cites.
“GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio”
Guoguo Chen et al · 2021
Later among the works it cites.
“VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation”
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino and Emmanuel Dupoux · 2021
Later among the works it cites.
“VoxLingua107: a dataset for spoken language recognition”
Jörgen Valk and Tanel Alumäe · 2021
Later among the works it cites.
“Scientific machine learning benchmarks”
Jeyan Thiyagalingam, Mallikarjun Shankar, Geoffrey Fox and Tony Hey · 2022
Later among the works it cites.
“Underspecification Presents Challenges for Credibility in Modern Machine Learning”
Alexander D’Amour et al · 2022
Later among the works it cites.
“SERAB: A Multi-Lingual Benchmark for Speech Emotion Recognition”
Neil Scheidwasser-Clow, Mikolaj Kegler, Pierre Beckmann and Milos Cernak · 2022
Later among the works it cites.
“Probing speech emotion recognition transformers for linguistic knowledge”
Andreas Triantafyllopoulos, Johannes Wagner, Hagen Wierstorf, Maximilian Schmitt, Uwe Reichel, Florian Eyben, Felix Burkhardt and Björn. Schuller · 2022
Later among the works it cites.
“Bias and Fairness on Multimodal Emotion Detection Algorithms”
Matheus Schmitz, Rehan Ahmed and Jimi Cao · 2022
Later among the works it cites.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka and Xiong Xiao · 2022
Later among the works it cites.
“Data2vec: A general framework for self-supervised learning in speech, vision and language”
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu and Michael Auli · 2022
Later among the works it cites.
“Fast Yet Effective Speech Emotion Recognition with Self-distillation”
Zhao Ren, Thanh Nguyen, Yi Chang and Björn Schuller · 2022
Later among the works it cites.
“Test splits for CREMA-D, emoDB, IEMOCAP, MELD, RAVDESS”, 2023
Hagen Wierstorf and Anna Derington · 2023
Closest in time.
“Argos Translate”, 2023
P.J. Finlay and Contributors Argos · 2023
Closest in time.
“Praat: doing phonetics by computer [Computer program]” Retrieved from http://www.praat.org/ , 2023
Paul Boersma · 2023
Closest in time.
“Dawn of the transformer era in speech emotion recognition: closing the valence gap”
Johannes Wagner, Andreas Triantafyllopoulos, Hagen Wierstorf, Maximilian Schmitt, Florian Eyben and Björn Schuller · 2023
Closest in time.
“xLSTM: Extended Long Short-Term Memory”
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter and Sepp Hochreiter · 2024
Closest in time.
“Audio xLSTMs: Learning Self-supervised audio representations with xLSTMs”
Sarthak Yadav, Sergios Theodoridis and Zheng-Hua Tan · 2024
Closest in time.
“Odyssey2024 - Speech Emotion Recognition Challenge: Dataset, Baseline Framework, and Results”
L. Goncalves, A.. Salman, A. Reddy Naini, L. Moro-Velazquez, T. Thebaud, L. Paola Garcia, N. Dehak, B. Sisman and C. Busso · 2024
Closest in time.