Fetching the paper…
Reading the bibliography…
Speech emotion recognition (SER) is a key technology to enable more natural human-machine communication.
“A concordance correlation coefficient to evaluate reproducibility,”
Lawrence I-Kuei Lin, · 1989
Earlier work this paper cites.
“Describing the emotional states that are expressed in speech,”
Roddy Cowie and Randolph R Cornelius, · 2003
Earlier work this paper cites.
“Extracting moods from pictures and sounds: Towards truly personalized TV,”
Alan Hanjalic, · 2006
Earlier work this paper cites.
“IEMOCAP: interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower Provost, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan, · 2008
Earlier work this paper cites.
“Interpreting ambiguous emotional expressions,”
Emily Mower, Angeliki Metallinou, Chi-Chun Lee, Abe Kazemzadeh, Carlos Busso, Sungbok Lee, and Shrikanth Narayanan, · 2009
Earlier work this paper cites.
“Introducing shared-hidden-layer autoencoders for transfer learning and their application in acoustic emotion recognition,”
Jun Deng, Rui Xia, Zixing Zhang, Yang Liu, and Björn Schuller, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network,”
George Trigeorgis, Fabien Ringeval, Raymond Brueckner, Erik Marchi, Mihalis A Nicolaou, Björn Schuller, and Stefanos Zafeiriou, · 2016
Cited alongside, same era.
AUTOMATIC SPEECH RECOGNITION
Dong Yu and Li Deng, · 2016
Cited alongside, same era.
“Discriminatively trained recurrent neural networks for continuous dimensional emotion recognition from audio.,”
Felix Weninger, Fabien Ringeval, Erik Marchi, and Björn W Schuller, · 2016
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Speech emotion recognition: Two decades in a nutshell, benchmarks, and ongoing trends,”
Björn W Schuller, · 2018
Cited alongside, same era.
“Representation learning with contrastive predictive coding,”
“The ordinal nature of emotions: An emerging approach,”
Georgios N Yannakakis, Roddy Cowie, and Carlos Busso, · 2018
Later among the works it cites.
“Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings,”
Reza Lotfian and Carlos Busso, · 2019
Later among the works it cites.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Later among the works it cites.
Zheng Lian, Jianhua Tao, Bin Liu, and Jian Huang, · 2019
Later among the works it cites.
“On variational bounds of mutual information,”
Ben Poole, Sherjil Ozair, Aaron van den Oord, A. Alemi, and G. Tucker, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aäron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Cited alongside, same era.
“Unsupervised learning approach to feature analysis for automatic speech emotion recognition,”
Sefik Emre Eskimez, Zhiyao Duan, and Wendi Heinzelman, · 2018
Cited alongside, same era.
Tom B. Brown et.al., · 2020
Later among the works it cites.
“Momentum contrast for unsupervised visual representation learning,”
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick, · 2020
Later among the works it cites.