Fetching the paper…
Reading the bibliography…
Speech emotion recognition is an important component of any human centered system.
Practical hidden voice attacks against speech and speaker recognition systems
Hadi Abdullah, Washington Garcia, Christian Peeters, Patrick Traynor, Kevin RB Butler, and Joseph Wilson. 2019 · 1904
Earlier work this paper cites.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019 · 1904
Earlier work this paper cites.
Measuring emotion: the self-assessment manikin and the semantic differential
Margaret M Bradley and Peter J Lang. 1994 · 1994
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008 · 2008
Earlier work this paper cites.
Shrikanth narayanan fundamental frequency analysis for speech emotion processing
Carlos Busso, Murtaza Bulut, and Sungbok Lee. 2009 · 2009
Earlier work this paper cites.
Automatic detection of “g-dropping” in american english using forced alignment
Jiahong Yuan and Mark Liberman · 2011
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al. 2014 · 2014
Earlier work this paper cites.
An overview of noise-robust automatic speech recognition
Jinyu Li, Li Deng, Yifan Gong, and Reinhold Haeb-Umbach. 2014 · 2014
Earlier work this paper cites.
Robust speaker identification in noisy and reverberant conditions
Xiaojia Zhao, Yuxuan Wang, and DeLiang Wang. 2014 · 2014
Earlier work this paper cites.
Audio augmentation for speech recognition
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Human emotions track changes in the acoustic environment
Weiyi Ma and William Forde Thompson. 2015 · 2015
Earlier work this paper cites.
Speech emotion recognition in noisy environment
Farah Chenchah and Zied Lachiri. 2016 · 2016
Earlier work this paper cites.
Determining the energetic and informational components of speech-on-speech masking
Gerald Kidd Jr, Christine R Mason, Jayaganesh Swaminathan, Elin Roverud, Kameron K Clayton, and Virginia Best. 2016 · 2016
Earlier work this paper cites.
Speech masking speech in everyday communication: The role of inhibitory control and working memory capacity
Victoria Stenback. 2016 · 2016
Cited alongside, same era.
Improving the robustness of deep neural networks via stability training
Stephan Zheng, Yang Song, Thomas Leung, and Ian Goodfellow. 2016 · 2016
Cited alongside, same era.
Pooling acoustic and lexical features for the prediction of valence
Zakaria Aldeneh, Soheil Khorram, Dimitrios Dimitriadis, and Emily Mower Provost. 2017 · 2017
Cited alongside, same era.
Using regional saliency for speech emotion recognition
Zakaria Aldeneh and Emily Mower Provost. 2017 · 2017
Cited alongside, same era.
Hotflip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2017 · 2017
Cited alongside, same era.
A hybrid dsp/deep learning approach to real-time full-band speech enhancement
Jean-Marc Valin. 2018 · 2018
Later among the works it cites.
Front-end feature compensation and denoising for noise robust speech emotion recognition
Rupayan Chakraborty, Ashish Panda, Meghna Pandharipande, Sonal Joshi, and Sunil Kumar Kopparapu. 2019 · 2019
Later among the works it cites.
Muse-ing on the impact of utterance ordering on crowdsourced emotion annotations
Mimansa Jaiswal, Zakaria Aldeneh, Cristian-Paul Bara, Yuanhang Luo, Mihai Burzo, Rada Mihalcea, and Emily Mower Provost. 2019 · 2019
Later among the works it cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Later among the works it cites.
Real-time speech emotion recognition using a pre-trained image classification network: Effects of bandwidth reduction and companding
Margaret Lech, Melissa Stolar, Christopher Best, and Robert Bolia. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, et al. 2017 · 2017
Cited alongside, same era.
Evaluating discourse annotation: Some recent insights and new approaches
Jet Hoek and Merel Scholman. 2017 · 2017
Cited alongside, same era.
Capturing long-term temporal dependencies with convolutional networks for continuous emotion recognition
Soheil Khorram, Zakaria Aldeneh, Dimitrios Dimitriadis, Melvin McInnis, and Emily Mower Provost. 2017 · 2017
Cited alongside, same era.
The perception of emotions in noisified nonsense speech
Emilia Parada-Cabaleiro, Alice Baird, Anton Batliner, Nicholas Cummins, et al. 2017 · 2017
Cited alongside, same era.
Audio adversarial examples: Targeted attacks on speech-to-text
Nicholas Carlini and David Wagner · 2018
Cited alongside, same era.
A study of all-convolutional encoders for connectionist temporal classification
Kalpesh Krishna, Liang Lu, Kevin Gimpel, and Karen Livescu. 2018 · 2018
Cited alongside, same era.
The effect of noise on emotion perception in an unknown language
Odette Scharenborg, Sofoklis Kakouros, and Jiska Koemans. 2018 · 2018
Cited alongside, same era.
Siamese capsule network for end-to-end speaker recognition in the wild
Amirhossein Hajavi and Ali Etemad. 2021 · 2021
Closest in time.
Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training
Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski, Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Jacob Kahn, Ann Lee, Ronan Collobert, Gabriel Synnaeve, et al. 2021 · 2021
Closest in time.
Arabic speech emotion recognition employing wav2vec2. 0 and hubert based on baved dataset
Omar Mohamed and Salah A Aly. 2021 · 2021
Closest in time.
Copypaste: An augmentation method for speech emotion recognition
Raghavendra Pappagari, Jesús Villalba, Piotr Żelasko, Laureano Moro-Velazquez, and Najim Dehak. 2021 · 2021
Closest in time.
Emotion recognition from speech using wav2vec 2.0 embeddings
Leonardo Pepino, Pablo Riera, and Luciana Ferrer. 2021 · 2021
Closest in time.
Head fusion: Improving the accuracy and robustness of speech emotion recognition on the iemocap and ravdess dataset
Mingke Xu, Fan Zhang, and Wei Zhang. 2021 · 2021
Closest in time.
Data augmentation techniques for speech emotion recognition and deep learning
José Antonio Nicolás, Javier de Lope, and Manuel Graña. 2022 · 2022
Closest in time.