Fetching the paper…
Reading the bibliography…
In this work, we explore the dependencies between speaker recognition and emotion recognition.
“Iemocap: Interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan, · 2008
Earlier work this paper cites.
“Recent developments in opensmile, the munich open-source multimedia feature extractor,”
Florian Eyben, Felix Weninger, Florian Gross, and Björn Schuller, · 2013
Earlier work this paper cites.
“Unsupervised methods for speaker diarization: An integrated and iterative approach,”
Stephen H Shum, Najim Dehak, Réda Dehak, and James R Glass, · 2013
Earlier work this paper cites.
“Speech emotion recognition using cnn,”
Zhengwei Huang, Ming Dong, Qirong Mao, and Yongzhao Zhan, · 2014
Earlier work this paper cites.
“Speaker diarization with plda i-vector scoring and unsupervised calibration,”
Gregory Sell and Daniel Garcia-Romero, · 2014
Earlier work this paper cites.
“The geneva minimalistic acoustic parameter set (gemaps) for voice research and affective computing,”
Florian Eyben, Klaus R Scherer, Björn W Schuller, Johan Sundberg, Elisabeth André, Carlos Busso, Laurence Y Devillers, Julien Epps, Petri Laukka, Shrikanth S Narayanan, et al., · 2015
Earlier work this paper cites.
“Musan: A music, speech, and noise corpus,”
David Snyder, Guoguo Chen, and Daniel Povey, · 2015
Earlier work this paper cites.
“Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network,”
George Trigeorgis, Fabien Ringeval, Raymond Brueckner, Erik Marchi, Mihalis A Nicolaou, Björn Schuller, and Stefanos Zafeiriou, · 2016
Earlier work this paper cites.
“Speech emotion recognition using convolutional and recurrent neural networks,”
Wootaek Lim, Daeyoung Jang, and Taejin Lee, · 2016
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“End-to-end multimodal emotion recognition using deep neural networks,”
Panagiotis Tzirakis, George Trigeorgis, Mihalis A Nicolaou, Björn W Schuller, and Stefanos Zafeiriou, · 2017
Cited alongside, same era.
“Sphereface: Deep hypersphere embedding for face recognition,”
Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li, Bhiksha Raj, and Le Song, · 2017
Cited alongside, same era.
“Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings,”
Reza Lotfian and Carlos Busso, · 2017
Cited alongside, same era.
“Emotion identification from raw speech signals using dnns.,”
Mousmita Sarma, Pegah Ghahremani, Daniel Povey, Nagendra Kumar Goel, Kandarpa Kumar Sarma, and Najim Dehak, · 2018
Cited alongside, same era.
“Deep neural networks for emotion recognition combining audio and transcripts.,”
Jaejin Cho, Raghavendra Pappagari, Purva Kulkarni, Jesús Villalba, Yishay Carmiel, and Najim Dehak, · 2018
Cited alongside, same era.
“Characterizing performance of speaker diarization systems on far-field speech using standard methods,”
Matthew Maciejewski, David Snyder, Vimal Manohar, Najim Dehak, and Sanjeev Khudanpur, · 2018
Later among the works it cites.
“Diarization is hard: Some experiences and lessons learned for the jhu team in the inaugural dihard challenge.,”
Gregory Sell, David Snyder, Alan McCree, Daniel Garcia-Romero, Jesús Villalba, Matthew Maciejewski, Vimal Manohar, Najim Dehak, Daniel Povey, Shinji Watanabe, et al., · 2018
Later among the works it cites.
“X-vectors: Robust dnn embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Later among the works it cites.
“Speech emotion recognition using deep 1d & 2d cnn lstm networks,”
Jianfeng Zhao, Xia Mao, and Lijiang Chen, · 2019
Later among the works it cites.
“Improving emotion classification through variational inference of latent variables,”
Srinivas Parthasarathy, Viktor Rozgic, Ming Sun, and Chao Wang, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Siddique Latif, Rajib Rana, and Junaid Qadir, · 2018
Cited alongside, same era.
“Towards conditional adversarial training for predicting emotions from speech,”
Jing Han, Zixing Zhang, Zhao Ren, Fabien Ringeval, and Björn Schuller, · 2018
Cited alongside, same era.
“On enhancing speech emotion recognition using generative adversarial networks,”
Saurabh Sahu, Rahul Gupta, and Carol Espy-Wilson, · 2018
Cited alongside, same era.
“Transfer learning for improving speech emotion classification accuracy,”
Siddique Latif, Rajib Rana, Shahzad Younis, Junaid Qadir, and Julien Epps, · 2018
Cited alongside, same era.
“Reusing neural speech representations for auditory emotion recognition,”
Egor Lakomkin, Cornelius Weber, Sven Magg, and Stefan Wermter, · 2018
Cited alongside, same era.
“Probing the information encoded in x-vectors,”
Desh Raj, David Snyder, Daniel Povey, and Sanjeev Khudanpur, · 2019
Later among the works it cites.
“Disentangling style factors from speaker representations,”
Jennifer Williams and Simon King, · 2019
Later among the works it cites.
“State-of-the-art speaker recognition for telephone and video speech: The jhu-mit submission for nist sre18,”
Jesús Villalba, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, Fred Richardson, Suwon Shon, François Grondin, et al., · 2019
Later among the works it cites.
“Curriculum learning for speech emotion recognition from crowdsourced labels,”
Reza Lotfian and Carlos Busso, · 2019
Later among the works it cites.
“A study of x-vector based speaker recognition on short utterances,”
Ahilan Kanagasundaram, Sridha Sridharan, Sriram Ganapathy, Prachi Singh, and Clinton B Fookes, · 2019
Later among the works it cites.