Fetching the paper…
Reading the bibliography…
Multimodal emotion recognition study is hindered by the lack of labelled corpora in terms of scale and diversity, due to the high annotation cost and label ambiguity.
“Emotion recognition in human–computer interaction,”
Nickolaos Fragopanagos and John G Taylor, · 2005
Earlier work this paper cites.
“Iemocap: interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, and et al., · 2008
Earlier work this paper cites.
“Msp-improv: An acted corpus of dyadic interactions to study emotion perception,”
Carlos Busso, Srinivas Parthasarathy, Alec Burmania, and et al., · 2016
Earlier work this paper cites.
“Training deep networks for facial expression recognition with crowd-sourced label distribution,”
Emad Barsoum, Cha Zhang, Cristian Canton Ferrer, and et al., · 2016
Earlier work this paper cites.
“Densely connected convolutional networks,”
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and et al., · 2017
Earlier work this paper cites.
“Memory fusion network for multi-view sequential learning,”
Amir Zadeh, Paul Pu Liang, Navonil Mazumder, and et al., · 2018
Earlier work this paper cites.
“Multi-attention recurrent network for human communication comprehension,”
Amir Zadeh, Paul Pu Liang, Soujanya Poria, and et al., · 2018
Earlier work this paper cites.
“Deep contextualized word representations,”
Matthew E Peters, Mark Neumann, Mohit Iyyer, and et al., · 2018
Earlier work this paper cites.
“Multimodal transformer for unaligned multimodal language sequences,”
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, and et al., · 2019
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and et al., · 2019
Earlier work this paper cites.
“Vl-bert: Pre-training of generic visual-linguistic representations,”
Weijie Su, Xizhou Zhu, Yue Cao, and et al., · 2019
Cited alongside, same era.
“Lxmert: Learning cross-modality encoder representations from transformers,”
Hao Tan and Mohit Bansal, · 2019
Cited alongside, same era.
“Videobert: A joint model for video and language representation learning,”
Chen Sun, Austin Myers, Carl Vondrick, and et al., · 2019
Cited alongside, same era.
“Language models as knowledge bases?,”
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, and et al., · 2019
Cited alongside, same era.
“Pre-training with whole word masking for chinese bert,”
Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Ziqing Yang, Shijin Wang, and Guoping Hu, · 2019
“Multi-modal attention for speech emotion recognition,”
Zexu Pan, Zhaojie Luo, Jichen Yang, and et al., · 2020
Later among the works it cites.
“Semi-supervised multi-modal emotion recognition with cross-modal distribution matching,”
Jingjun Liang, Ruichen Li, and Qin Jin, · 2020
Later among the works it cites.
“Multimodal pretraining unmasked: A meta-analysis and a unified framework of vision-and-language berts,”
Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, and et al., · 2021
Closest in time.
“Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,”
Hassan Akbari, Linagzhe Yuan, Rui Qian, and et al., · 2021
Closest in time.
Pengfei Liu, Weizhe Yuan, Jinlan Fu, and et al., · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Uniter: Universal image-text representation learning,” 2020
Yen-Chun Chen, Linjie Li, Licheng Yu, and et al., · 2020
Cited alongside, same era.
“Actbert: Learning global-local video-text representations,”
Linchao Zhu and Yi Yang, · 2020
Cited alongside, same era.
“Few-shot text generation with pattern-exploiting training,”
Timo Schick and Hinrich Schütze, · 2020
Cited alongside, same era.
“Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,”
Andy T Liu, Shu-wen Yang, Po-Han Chi, and et al., · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and et al., · 2020
Cited alongside, same era.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeffrey Wu, Rewon Child, and et al.,
Cited in the paper.
“Adapting language models for zero-shot learning by meta-tuning on dataset and prompt collections,” 2021
Ruiqi Zhong, Kristy Lee, Zheng Zhang, and et al., · 2021
Closest in time.
“Exploiting cloze-questions for few-shot text classification and natural language inference,”
Timo Schick and Hinrich Schütze, · 2021
Closest in time.
“A further study of unsupervised pretraining for transformer based speech recognition,”
Dongwei Jiang, Wubo Li, Ruixiong Zhang, and et al., · 2021
Closest in time.
“Missing modality imagination network for emotion recognition with uncertain missing modalities,”
Jinming Zhao, Ruichen Li, and Qin Jin, · 2021
Closest in time.
“Hero: Hierarchical encoder for video+ language omni-representation pre-training,”
Linjie Li, Yen-Chun Chen, Yu Cheng, and et al., · 2065
Closest in time.