Fetching the paper…
Reading the bibliography…
Automatic emotion recognition is one of the central concerns of the Human-Computer Interaction field as it can bridge the gap between humans and machines.
“Continuously variable duration hidden markov models for automatic speech recognition,”
Stephen E Levinson, · 1986
Earlier work this paper cites.
“Speech emotion recognition using hidden markov models,”
Tin Lay Nwe, Say Wei Foo, and Liyanage C De Silva, · 2003
Earlier work this paper cites.
“Speech emotion recognition based on hmm and svm,”
Yi-Lin Lin and Gang Wei, · 2005
Earlier work this paper cites.
“Emotion recognition from assamese speeches using mfcc features and gmm classifier,”
Aditya Bihar Kandali, Aurobinda Routray, and Tapan Kumar Basu, · 2008
Earlier work this paper cites.
“Iemocap: Interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan, · 2008
Earlier work this paper cites.
“Opensmile: the munich versatile and fast open-source audio feature extractor,”
Florian Eyben, Martin Wöllmer, and Björn Schuller, · 2010
Earlier work this paper cites.
“Survey on speech emotion recognition: Features, classification schemes, and databases,”
Moataz El Ayadi, Mohamed S Kamel, and Fakhri Karray, · 2011
Earlier work this paper cites.
“Speech emotion recognition using support vector machine,”
Yixiong Pan, Peipei Shen, and Liping Shen, · 2012
Earlier work this paper cites.
“Speech emotion recognition using deep neural network and extreme learning machine,”
Kun Han, Dong Yu, and Ivan Tashev, · 2014
Earlier work this paper cites.
“Explaining and harnessing adversarial examples,”
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy, · 2014
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville, · 2016
Earlier work this paper cites.
“Unsupervised learning of disentangled representations from video,”
Emily L Denton et al., · 2017
Cited alongside, same era.
“Voxceleb: a large-scale speaker identification dataset,”
A. Nagrani, J. S. Chung, and A. Zisserman, · 2017
Cited alongside, same era.
“Sentiment analysis by capsules,”
Yequan Wang, Aixin Sun, Jialong Han, Ying Liu, and Xiaoyan Zhu, · 2018
Cited alongside, same era.
“Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,”
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron J Weiss, Rob Clark, and Rif A Saurous, · 2018
Cited alongside, same era.
“Generalized end-to-end loss for speaker verification,”
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno, · 2018
Cited alongside, same era.
“Voxceleb2: Deep speaker recognition,”
J. S. Chung, A. Nagrani, and A. Zisserman, · 2018
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Later among the works it cites.
“Jointly Fine-Tuning “BERT-Like” Self Supervised Models to Improve Multimodal Speech Emotion Recognition,”
Shamane Siriwardhana, Andrew Reis, Rivindu Weerasekera, and Suranga Nanayakkara, · 2020
Later among the works it cites.
“Electra: Pre-training text encoders as discriminators rather than generators,”
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning, · 2020
Later among the works it cites.
“Unsupervised speech decomposition via triple information bottleneck,”
Kaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson, and David Cox, · 2020
Later among the works it cites.
“Speaker-invariant affective representation learning via adversarial training,”
Haoqi Li, Ming Tu, Jing Huang, Shrikanth Narayanan, and Panayiotis Georgiou, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“Xlnet: Generalized autoregressive pretraining for language understanding,”
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le, · 2019
Cited alongside, same era.
“Autovc: Zero-shot voice style transfer with only autoencoder loss,”
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson, · 2019
Cited alongside, same era.
“Disentangling style factors from speaker representations.,”
Jennifer Williams and Simon King, · 2019
Cited alongside, same era.
“Multimodal emotion recognition using cross-modal attention and 1d convolutional neural networks,”
DN Krishna and Ankita Patil, · 2020
Cited alongside, same era.
“Speech representation learning for emotion recognition using end-to-end asr with factorized adaptation,”
Sung-Lin Yeh, Yun-Shao Lin, and Chi-Chun Lee, · 2020
Cited alongside, same era.
“Fusion approaches for emotion recognition from speech using acoustic and text-based features,”
Leonardo Pepino, Pablo Riera, Luciana Ferrer, and Agustín Gravano, · 2020
Later among the works it cites.
“Advancing multiple instance learning with attention modeling for categorical speech emotion recognition,”
Shuiyang Mao, PC Ching, C-C Jay Kuo, and Tan Lee, · 2020
Later among the works it cites.
“Emotion recognition from speech using wav2vec 2.0 embeddings,”
Leonardo Pepino, Pablo Riera, and Luciana Ferrer, · 2021
Closest in time.
“On the use of self-supervised pre-trained acoustic and linguistic features for continuous speech emotion recognition,”
Manon Macary, Marie Tahon, Yannick Estève, and Anthony Rousseau, · 2021
Closest in time.
“A novel attention-based gated recurrent unit and its efficacy in speech emotion recognition,”
Srividya Tirunellai Rajamani, Kumar T Rajamani, Adria Mallol-Ragolta, Shuo Liu, and Björn Schuller, · 2021
Closest in time.
“Transformer based unsupervised pre-training for acoustic representation learning,”
Ruixiong Zhang, Haiwei Wu, Wubo Li, Dongwei Jiang, Wei Zou, and Xiangang Li, · 2021
Closest in time.
“Emotion recognition by fusing time synchronous and time asynchronous representations,”
Wen Wu, Chao Zhang, and Philip C. Woodland, · 2021
Closest in time.
“Exploring disentangled feature representation beyond face identification,”
Yu Liu, Fangyin Wei, Jing Shao, Lu Sheng, Junjie Yan, and Xiaogang Wang, · 2089
Closest in time.