Fetching the paper…
Reading the bibliography…
Multimodal emotion recognition from speech is an important area in affective computing.
A. Kolesnikov, X. Zhai, and L. Beyer, “Revisiting self-supervised visual representation learning,” in
1929
Earlier work this paper cites.
R. W. Picard,
2000
Earlier work this paper cites.
N. Sebe, I. Cohen, T. Gevers, and T. S. Huang, “Multimodal approaches for emotion recognition: a survey,” in
2005
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,”
2008
Earlier work this paper cites.
G. Degottex, J. Kane, T. Drugman, T. Raitio, and S. Scherer, “Covarep—a collaborative voice analysis repository for speech technologies,” in
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. van den Oord, O. Vinyals
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
S. Poria, E. Cambria, D. Hazarika, N. Majumder, A. Zadeh, and L.-P. Morency, “Context-dependent sentiment analysis in user-generated videos,” in
2017
Earlier work this paper cites.
J. Kim and R. A. Saurous, “Emotion recognition from human speech using temporal information and deep learning.” in
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
A. Zadeh, P. P. Liang, S. Poria, P. Vij, E. Cambria, and L.-P. Morency, “Multi-attention recurrent network for human communication comprehension,” in
2018
Cited alongside, same era.
M. Swain, A. Routray, and P. Kabisatpathy, “Databases, features and classifiers for speech emotion recognition: a review,”
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”
2018
Cited alongside, same era.
A. B. Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency, “Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,” in
2018
Cited alongside, same era.
J.-L. Li and C.-C. Lee, “Attentive to individual: A multimodal emotion recognition network with personalized attention profile,”
2019
Cited alongside, same era.
A. Singh, V. Kadyan, M. Kumar, and N. Bassan, “Asroil: a comprehensive survey for automatic speech recognition of indian languages,”
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Z. Han, H. Zhao, and R. Wang, “Transfer learning for speech emotion recognition,” in
2019
Cited alongside, same era.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,”
2019
Later among the works it cites.
2019
Later among the works it cites.
K. Feng and T. Chaspari, “A review of generalizable transfer learning in automatic emotion recognition,”
2020
Closest in time.