Fetching the paper…
Reading the bibliography…
The wide application of smart devices enables the availability of multimodal data, which can be utilized in many tasks.
D. Olson, “From utterance to text: The bias of language in speech and writing,” Harvard Educational Review , vol. 47, no. 3, pp. 257–281, 1977
1977
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997. [Online]. Available: https://doi.org/10.1162/neco.1997.9.8.1735
1997
Earlier work this paper cites.
K. Q. Weinberger and L. K. Saul, “Distance metric learning for large margin nearest neighbor classification.” Journal of machine learning research , vol. 10, no. 2, 2009
2009
Earlier work this paper cites.
C. H. Wu and W. B. Liang, “Emotion recognition of affective speech based on multiple classifiers using acoustic-prosodic information and semantic labels,” IEEE Transactions on Affective Computing , vol. 2, no. 1, pp. 10–21, 2010
2010
Earlier work this paper cites.
J. Wagner, E. Andre, F. Lingenfelser, and J. Kim, “Exploring fusion methods for multimodal emotion recognition with missing data,” IEEE Transactions on Affective Computing , vol. 2, no. 4, pp. 206–218, 2011
2011
Earlier work this paper cites.
V. Rozgic, S. Ananthakrishnan, S. Saleem, R. Kumar, and R. Prasad, “Ensemble of svm trees for multimodal emotion recognition,” in Signal and Information Processing Association Summit and Conference , 2012, pp. 1–4
2012
Earlier work this paper cites.
M. Wollmer, F. Weninger, T. Knaup, B. Schuller, C. Sun, K. Sagae, and L. P. Morency, “Youtube movie reviews: Sentiment analysis in an audio-visual context,” IEEE Intelligent Systems , vol. 28, no. 3, pp. 46–53, 2013
2013
Earlier work this paper cites.
K. Cho, B. Van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder-decoder for statistical machine translation,” in EMNLP , 2014, pp. 1724–1734
2014
Earlier work this paper cites.
G. Degottex, J. Kane, T. Drugman, T. Raitio, and S. Scherer, “Covarep: A collaborative voice analysis repository for speech technologies,” in ICASSP , 2014, pp. 960–964
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proceedings of International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
A. Zadeh, R. Zellers, E. Pincus, and L. P. Morency, “Mosi: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos,” IEEE Intelligent Systems , vol. 31, no. 6, pp. 82–88, 2016
2016
Earlier work this paper cites.
S. Poria, I. Chaturvedi, E. Cambria, and A. Hussain, “Convolutional mkl based multimodal emotion recognition and sentiment analysis,” in Proceedings of IEEE International Conference on Data Mining (ICDM) , 2016, pp. 439–448
2016
Earlier work this paper cites.
B. Nojavanasghari, D. Gopinath, J. Koushik, and L. P. Morency, “Deep multimodal fusion for persuasiveness prediction,” in Proceedings of ACM International Conference on Multimodal Interaction , 2016, pp. 284–288
2016
Earlier work this paper cites.
A. Zadeh, R. Zellers, E. Pincus, and L. P. Morency, “Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages,” IEEE Intelligent Systems , vol. 31, no. 6, pp. 82–88, 11 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
S. Poria, E. Cambria, R. Bajpai, and A. Hussain, “A review of affective computing: From unimodal analysis to multimodal fusion,” Information Fusion , vol. 37, pp. 98–125, 2017
2017
Cited alongside, same era.
S. Poria, E. Cambria, D. Hazarika, N. Majumder, A. Zadeh, and L. P. Morency, “Context-dependent sentiment analysis in user-generated videos,” in ACL , 2017, pp. 873–883
2017
Cited alongside, same era.
A. Zadeh, M. Chen, S. Poria, E. Cambria, and L. P. Morency, “Tensor fusion network for multimodal sentiment analysis,” in EMNLP , 2017, pp. 1114–1125
2017
Cited alongside, same era.
A. Zadeh, P. P. Liang, N. Mazumder, S. Poria, E. Cambria, and L. P. Morency, “Memory fusion network for multi-view sequential learning,” in AAAI , 2018, pp. 5634–5641
2018
Cited alongside, same era.
A. Zadeh, P. P. Liang, S. Poria, P. Vij, E. Cambria, and L. P. Morency, “Multi-attention recurrent network for human communication comprehension,” in AAAI , 2018, pp. 5642–5649
S. Mai, H. Hu, J. Xu, and S. Xing, “Multi-fusion residual memory network for multimodal human sentiment comprehension,” IEEE Transactions on Affective Computing , pp. 1–1, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
W. Rahman, M. Hasan, S. Lee, A. Zadeh, C. Mao, L.-P. Morency, and E. Hoque, “Integrating multimodal information in large pretrained transformers,” Proceedings of the conference. Association for Computational Linguistics. Meeting , vol. 2020, pp. 2359–2369, 2020
2020
Later among the works it cites.
D. Hazarika, R. Zimmermann, and S. Poria, “Misa: Modality-invariant and -specific representations for multimodal sentiment analysis,” Proceedings of the 28th ACM International Conference on Multimedia , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Z. Liu, Y. Shen, P. P. Liang, A. Zadeh, and L. P. Morency, “Efficient low-rank multimodal fusion with modality-specific factors,” in ACL , 2018, pp. 2247–2256
2018
Cited alongside, same era.
A. Zadeh, P. P. Liang, J. Vanbriesen, S. Poria, E. Tong, E. Cambria, M. Chen, and L. P. Morency, “Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,” in ACL , 2018, pp. 2236–2246
2018
Cited alongside, same era.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 41, no. 2, pp. 423–443, Feb 2019
2019
Cited alongside, same era.
S. Mai, H. Hu, and S. Xing, “Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing,” in ACL , Jul. 2019
2019
Cited alongside, same era.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in ACL , Jul. 2019
2019
Cited alongside, same era.
M. Hou, J. Tang, J. Zhang, W. Kong, and Q. Zhao, “Deep multimodal multilinear fusion with high-order polynomial pooling,” in Advances in Neural Information Processing Systems , 2019, pp. 12 113–12 122
2019
Cited alongside, same era.
K. Yang, H. Xu, and K. Gao, “Cm-bert: Cross-modal bert for text-audio sentiment analysis,” in Proceedings of the 28th ACM International Conference on Multimedia , 2020, pp. 521–528
2020
Later among the works it cites.
S. Mai, H. Hu, and S. Xing, “Modality to modality translation: An adversarial representation learning and graph fusion network for multimodal fusion,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 01, 2020, pp. 164–172
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
Z. Sun, P. Sarma, W. Sethares, and Y. Liang, “Learning relationships between text, audio, and video via deep canonical correlation for multimodal language analysis,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 05, 2020, pp. 8992–8999
2020
Later among the works it cites.
L. Stappen, A. Baird, L. Schumann, and S. Bjorn, “The multimodal sentiment analysis in car reviews (muse-car) dataset: Collection, insights and improvements,” IEEE Transactions on Affective Computing , pp. 1–1, 2021
2021
Closest in time.
Y. Zhao, X. Cao, J. Lin, D. Yu, and X. Cao, “Multimodal affective states recognition based on multiscale cnns and biologically inspired decision fusion model,” IEEE Transactions on Affective Computing , pp. 1–1, 2021
2021
Closest in time.
Q. Li, D. Gkoumas, C. Lioma, and M. Melucci, “Quantum-inspired multimodal fusion for video sentiment analysis,” Information Fusion , vol. 65, pp. 58 – 71, 2021. [Online]. Available: http://www.sciencedirect.com/science/article/pii/S1566253520303365
2021
Closest in time.
D. Gkoumas, Q. Li, C. Lioma, Y. Yu, and D. wei Song, “What makes the difference? an empirical comparison of fusion strategies for multimodal language analysis,” Information Fusion , vol. 66, pp. 184–197, 2021
2021
Closest in time.