Fetching the paper…
Reading the bibliography…
In this paper, we study the task of multimodal sequence analysis which aims to draw inferences from visual, language and acoustic sequences.
1903
Earlier work this paper cites.
D. Olson, “From utterance to text: The bias of language in speech and writing,” Harvard Educational Review , vol. 47, no. 3, pp. 257–281, 1977
1977
Earlier work this paper cites.
M. W. Goudreau, C. L. Giles, S. T. Chakradhar, and . Chen, D., “First-order versus second-order single-layer recurrent neural networks,” IEEE Transactions on Neural Networks , vol. 5, no. 3, pp. 511–513, 1994
1994
Earlier work this paper cites.
Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” IEEE Transactions on Neural Networks , vol. 5, no. 2, pp. 157–166, 1994
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
J. Yuan and M. Liberman, “Speaker identification on the SCOTUS corpus,” Acoustical Society of America Journal , vol. 123, p. 3878, 2008
2008
Earlier work this paper cites.
A. Micheli, “Neural network for graphs: A contextual constructive approach,” IEEE Transactions on Neural Networks , vol. 20, no. 3, pp. p.498–511, 2009
2009
Earlier work this paper cites.
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini, “The graph neural network model,” IEEE Transactions on Neural Networks , vol. 20, no. 1, pp. 61–80, 2009
2009
Earlier work this paper cites.
C. H. Wu and W. B. Liang, “Emotion recognition of affective speech based on multiple classifiers using acoustic-prosodic information and semantic labels,” IEEE Transactions on Affective Computing , vol. 2, no. 1, pp. 10–21, 2010
2010
Earlier work this paper cites.
V. Rozgic, S. Ananthakrishnan, S. Saleem, R. Kumar, and R. Prasad, “Ensemble of svm trees for multimodal emotion recognition,” in Signal and Information Processing Association Summit and Conference , 2012, pp. 1–4
2012
Earlier work this paper cites.
S. Sun, “A survey of multi-view machine learning,” Neural Computing and Applications , vol. 23, no. 7-8, pp. 2031–2038, 2013
2013
Earlier work this paper cites.
M. Wollmer, F. Weninger, T. Knaup, B. Schuller, C. Sun, K. Sagae, and L. P. Morency, “Youtube movie reviews: Sentiment analysis in an audio-visual context,” IEEE Intelligent Systems , vol. 28, no. 3, pp. 46–53, 2013
2013
Earlier work this paper cites.
K. Cho, B. Van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder-decoder for statistical machine translation,” in EMNLP , 2014, pp. 1724–1734
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in Empirical Methods in Natural Language Processing (EMNLP) , 2014, pp. 1532–1543. [Online]. Available: http://www.aclweb.org/anthology/D14-1162
2014
Earlier work this paper cites.
G. Degottex, J. Kane, T. Drugman, T. Raitio, and S. Scherer, “Covarep: A collaborative voice analysis repository for speech technologies,” in ICASSP , 2014, pp. 960–964
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proceedings of International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
S. Poria, I. Chaturvedi, E. Cambria, and A. Hussain, “Convolutional mkl based multimodal emotion recognition and sentiment analysis,” in Proceedings of IEEE International Conference on Data Mining (ICDM) , 2016, pp. 439–448
2016
Earlier work this paper cites.
B. Nojavanasghari, D. Gopinath, J. Koushik, and L. P. Morency, “Deep multimodal fusion for persuasiveness prediction,” in Proceedings of ACM International Conference on Multimodal Interaction , 2016, pp. 284–288
2016
Earlier work this paper cites.
A. Zadeh, R. Zellers, E. Pincus, and L. P. Morency, “Mosi: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos,” IEEE Intelligent Systems , vol. 31, no. 6, pp. 82–88, 2016
2016
Earlier work this paper cites.
T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” in ICLR , 2016
2016
Earlier work this paper cites.
A. Zadeh, R. Zellers, E. Pincus, and L. P. Morency, “Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages,” IEEE Intelligent Systems , vol. 31, no. 6, pp. 82–88, 11 2016
2016
Cited alongside, same era.
A. Zadeh, M. Chen, S. Poria, E. Cambria, and L. P. Morency, “Tensor fusion network for multimodal sentiment analysis,” in EMNLP , 2017, pp. 1114–1125
2017
Cited alongside, same era.
W. Hamilton, Z. Ying, and J. Leskovec, “Inductive representation learning on large graphs,” in Advances in neural information processing systems , 2017, pp. 1024–1034
2017
Cited alongside, same era.
S. Poria, E. Cambria, D. Hazarika, N. Majumder, A. Zadeh, and L. P. Morency, “Context-dependent sentiment analysis in user-generated videos,” in ACL , 2017, pp. 873–883
2017
Cited alongside, same era.
S. Bai, J. Kolter, and V. Koltun, “Trellis networks for sequence modeling,” in ICLR , 2019
2019
Later among the works it cites.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 41, no. 2, pp. 423–443, Feb 2019
2019
Later among the works it cites.
G. Aguilar, V. Rozgic, W. Wang, and C. Wang, “Multimodal and multi-view models for emotion recognition,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 991–1002
2019
Later among the works it cites.
P. P. Liang, Z. Liu, Y.-H. H. Tsai, Q. Zhao, R. Salakhutdinov, and L.-P. Morency, “Learning representations from imperfect time series data via tensor rank regularization,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Florence, Italy: Association for Computational Linguistics, Jul. 2019, pp. 1569–1576. [Online]. Available: https://www.aclweb.org/anthology/P19-1152
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Chen, S. Wang, P. P. Liang, T. Baltrus a ˇ \check{a} itis, A. Zadeh, and L. P. Morency, “Multimodal sentiment analysis with word-level fusion and reinforcement learning,” in 19th ACM International Conference on Multimodal Interaction (ICMI’17) , 2017, pp. 163–171
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Cited alongside, same era.
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin, “Convolutional sequence to sequence learning,” in Proceedings of the 34th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, D. Precup and Y. W. Teh, Eds., vol. 70. International Convention Centre, Sydney, Australia: PMLR, 06–11 Aug 2017, pp. 1243–1252. [Online]. Available: http://proceedings.mlr.press/v70/gehring17a.html
2017
Cited alongside, same era.
P. P. Liang, Z. Liu, A. Zadeh, and L. P. Morency, “Multimodal language analysis with recurrent multistage fusion,” in EMNLP , 2018, pp. 150–161
2018
Cited alongside, same era.
A. Zadeh, P. P. Liang, N. Mazumder, S. Poria, E. Cambria, and L. P. Morency, “Memory fusion network for multi-view sequential learning,” in AAAI , 2018, pp. 5634–5641
2018
Cited alongside, same era.
A. Zadeh, P. P. Liang, J. Vanbriesen, S. Poria, E. Tong, E. Cambria, M. Chen, and L. P. Morency, “Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,” in ACL , 2018, pp. 2236–2246
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
S. Mai, H. Hu, and S. Xing, “Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing,” in Proceedings of the 57th Conference of the Association for Computational Linguistics . Florence, Italy: Association for Computational Linguistics, Jul. 2019, pp. 481–492. [Online]. Available: https://www.aclweb.org/anthology/P19-1046
2019
Later among the works it cites.
M. Hou, J. Tang, J. Zhang, W. Kong, and Q. Zhao, “Deep multimodal multilinear fusion with high-order polynomial pooling,” in Advances in Neural Information Processing Systems , 2019, pp. 12 113–12 122
2019
Later among the works it cites.
H. Pham, P. P. Liang, T. Manzini, L. P. Morency, and P. Barnabǎs, “Found in translation: Learning robust joint representations by cyclic translations between modalities,” in AAAI , 2019, pp. 6892–6899
2019
Later among the works it cites.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Florence, Italy: Association for Computational Linguistics, Jul. 2019, pp. 6558–6569. [Online]. Available: https://www.aclweb.org/anthology/P19-1656
2019
Later among the works it cites.
Y. H. H. Tsai, P. P. Liang, A. Zadeh, L. P. Morency, and R. Salakhutdinov, “Learning factorized multimodal representations,” in ICLR , 2019
2019
Later among the works it cites.
P. P. Liang, Y. C. Lim, Y. H. Tsai, R. R. Salakhutdinov, and L.-P. Morency, “Strong and simple baselines for multimodal utterance embeddings,” in NAACL , 2019, pp. 2599–2609
2019
Later among the works it cites.
D. S. Chauhan, M. S. Akhtar, A. Ekbal, and P. Bhattacharyya, “Context-aware interactive attention for multi-modal sentiment and emotion analysis,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 5651–5661
2019
Later among the works it cites.
M. S. Akhtar, D. Chauhan, D. Ghosal, S. Poria, A. Ekbal, and P. Bhattacharyya, “Multi-task learning for multi-modal emotion recognition and sentiment analysis,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019, pp. 370–379
2019
Later among the works it cites.
E. Georgiou, C. Papaioannou, and A. Potamianos, “Deep hierarchical fusion with application in sentiment analysis,” Proc. Interspeech 2019 , pp. 1646–1650, 2019
2019
Later among the works it cites.
A. Pandey and D. Wang, “Tcnn: Temporal convolutional neural network for real-time speech enhancement in the time domain,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6875–6879
2019
Later among the works it cites.
2020
Closest in time.
S. Mai, H. Hu, J. Xu, and S. Xing, “Multi-fusion residual memory network for multimodal human sentiment comprehension,” IEEE Transactions on Affective Computing , pp. 1–1, 2020
2020
Closest in time.
S. Mai, S. Xing, and H. Hu, “Locally confined modality fusion network with a global perspective for multimodal human affective computing,” IEEE Transactions on Multimedia , vol. 22, no. 1, pp. 122–137, 2020
2020
Closest in time.
S. Mai, H. Hu, and S. Xing, “Modality to modality translation: An adversarial representation learning and graph fusion network for multimodal fusion,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 01, 2020, pp. 164–172
2020
Closest in time.
H. Yuan and S. Ji, “Structpool: Structured graph pooling via conditional random fields,” in ICLR , 2020
2020
Closest in time.
D. Gkoumas, Q. Li, C. Lioma, Y. Yu, and D. wei Song, “What makes the difference? an empirical comparison of fusion strategies for multimodal language analysis,” Information Fusion , vol. 66, pp. 184–197, 2021
2021
Closest in time.