Fetching the paper…
Reading the bibliography…
Data augmentation is a widely used strategy for training robust machine learning models.
“Speech emotion perception by human and machine,”
Szabolcs Levente Tóth, David Sztahó, and Klára Vicsi, · 2008
Earlier work this paper cites.
“Iemocap: Interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan, · 2008
Earlier work this paper cites.
“Emotion recognition from speech: a review,”
Shashidhar G Koolagudi and K Sreenivasa Rao, · 2012
Earlier work this paper cites.
“Recent developments in opensmile, the munich open-source multimedia feature extractor,”
Florian Eyben, Felix Weninger, Florian Gross, and Björn Schuller, · 2013
Earlier work this paper cites.
“Speech emotion recognition using cnn,”
Zhengwei Huang, Ming Dong, Qirong Mao, and Yongzhao Zhan, · 2014
Earlier work this paper cites.
“Feature selection for automatic analysis of emotional response based on nonlinear speech modeling suitable for diagnosis of alzheimer’s disease,”
Karmele Lopez-de Ipiña, JB Alonso-Hernández, Jordi Solé-Casals, Carlos Manuel Travieso-González, Aitzol Ezeiza, Marcos Faundez-Zanuy, Pilar M Calvo, and Blanca Beitia, · 2015
Earlier work this paper cites.
“Musan: A music, speech, and noise corpus,”
David Snyder, Guoguo Chen, and Daniel Povey, · 2015
Earlier work this paper cites.
“Robust speech recognition in unknown reverberant and noisy conditions,”
Roger Hsiao, Jeff Ma, William Hartmann, Martin Karafiát, František Grézl, Lukáš Burget, Igor Szöke, Jan Honza Černockỳ, Shinji Watanabe, Zhuo Chen, et al., · 2015
Earlier work this paper cites.
“Speech emotion recognition using convolutional and recurrent neural networks,”
Wootaek Lim, Daeyoung Jang, and Taejin Lee, · 2016
Earlier work this paper cites.
“Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network,”
George Trigeorgis, Fabien Ringeval, Raymond Brueckner, Erik Marchi, Mihalis A Nicolaou, Björn Schuller, and Stefanos Zafeiriou, · 2016
Cited alongside, same era.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Cited alongside, same era.
“A survey on hate speech detection using natural language processing,”
Anna Schmidt and Michael Wiegand, · 2017
Cited alongside, same era.
“End-to-end multimodal emotion recognition using deep neural networks,”
Panagiotis Tzirakis, George Trigeorgis, Mihalis A Nicolaou, Björn W Schuller, and Stefanos Zafeiriou, · 2017
Cited alongside, same era.
“Deep learning scaling is predictable, empirically,”
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou, · 2017
Cited alongside, same era.
“On the robustness of speech emotion recognition for human-robot interaction with deep neural networks,”
Egor Lakomkin, Mohammad Ali Zamani, Cornelius Weber, Sven Magg, and Stefan Wermter, · 2018
Later among the works it cites.
“Cnn+ lstm architecture for speech emotion recognition with data augmentation,”
Caroline Etienne, Guillaume Fidanza, Andrei Petrovskii, Laurence Devillers, and Benoit Schmauch, · 2018
Later among the works it cites.
“Speech emotion recognition using deep 1d & 2d cnn lstm networks,”
Jianfeng Zhao, Xia Mao, and Lijiang Chen, · 2019
Later among the works it cites.
“Cyclegan-based emotion style transfer as data augmentation for speech emotion recognition.,”
Fang Bao, Michael Neumann, and Ngoc Thang Vu, · 2019
Later among the works it cites.
“State-of-the-art speaker recognition for telephone and video speech: The jhu-mit submission for nist sre18,”
Jesús Villalba, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, Fred Richardson, Suwon Shon, François Grondin, et al., · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings,”
Reza Lotfian and Carlos Busso, · 2017
Cited alongside, same era.
“Deep neural networks for emotion recognition combining audio and transcripts.,”
Jaejin Cho, Raghavendra Pappagari, Purva Kulkarni, Jesús Villalba, Yishay Carmiel, and Najim Dehak, · 2018
Cited alongside, same era.
“Emotion identification from raw speech signals using dnns.,”
Mousmita Sarma, Pegah Ghahremani, Daniel Povey, Nagendra Kumar Goel, Kandarpa Kumar Sarma, and Najim Dehak, · 2018
Cited alongside, same era.
“Reusing neural speech representations for auditory emotion recognition,”
Egor Lakomkin, Cornelius Weber, Sven Magg, and Stefan Wermter, · 2018
Cited alongside, same era.
Later among the works it cites.
“Clinical state tracking in serious mental illness through computational analysis of speech,”
Armen C Arevian, Daniel Bone, Nikolaos Malandrakis, Victor R Martinez, Kenneth B Wells, David J Miklowitz, and Shrikanth Narayanan, · 2020
Closest in time.
“x-vectors meet emotions: A study on dependencies between emotion and speaker recognition,”
Raghavendra Pappagari, Tianzi Wang, Jesus Villalba, Nanxin Chen, and Najim Dehak, · 2020
Closest in time.
“Stargan for emotional speech conversion: Validated by data augmentation of end-to-end emotion recognition,”
Georgios Rizos, Alice Baird, Max Elliott, and Björn Schuller, · 2020
Closest in time.