Fetching the paper…
Reading the bibliography…
Multi-modal large language models (MLLMs) have achieved remarkable performance on objective multimodal perception tasks, but their ability to interpret subjective, emotionally nuanced multimodal content remains largely unexplored.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language resources and evaluation , 2008
2008
Earlier work this paper cites.
2014
Earlier work this paper cites.
K.-C. Peng, T. Chen, A. Sadovnik, and A. C. Gallagher, “A mixed bag of emotions: Model, predict, and transfer emotion distributions,” in CVPR , 2015
2015
Earlier work this paper cites.
T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” arXiv preprint , 2016
2016
Earlier work this paper cites.
S. Zhang, X. Zhu, Z. Lei, H. Shi, X. Wang, and S. Z. Li, “S3fd: Single shot scale-invariant face detector,” in ICCV , 2017
2017
Earlier work this paper cites.
R. Kosti, J. M. Alvarez, A. Recasens, and A. Lapedriza, “Emotion recognition in context,” in CVPR , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv preprint , 2017
2017
Earlier work this paper cites.
A. B. Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency, “Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,” in ACL , 2018
2018
Earlier work this paper cites.
A. Coucke, A. Saade, A. Ball, T. Bluche, A. Caulier, D. Leroy, C. Doumouro, T. Gisselbrecht, F. Caltagirone, T. Lavril et al. , “Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces,” arXiv preprint , 2018
2018
Earlier work this paper cites.
S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea, “Meld: A multimodal multi-party dataset for emotion recognition in conversations,” in ACL , 2019
2019
Earlier work this paper cites.
R. Kosti, J. M. Alvarez, A. Recesens, and A. Lapedriza, “Context based emotion recognition using emotic dataset,” IEEE TAPMI , 2019
2019
Earlier work this paper cites.
S. Larson, A. Mahendran, J. J. Peper, C. Clarke, A. Lee, P. Hill, J. K. Kummerfeld, K. Leach, M. A. Laurenzano, L. Tang et al. , “An evaluation dataset for intent classification and out-of-scope prediction,” in EMNLP , 2019
2019
Earlier work this paper cites.
X. Liu, A. Eshghi, P. Swietojanski, and V. Rieser, “Benchmarking natural language understanding services for building conversational agents,” arXiv preprint , 2019
2019
Earlier work this paper cites.
J. Kruk, J. Lubin, K. Sikka, X. Lin, D. Jurafsky, and A. Divakaran, “Integrating text and image: Determining multimodal document intent in Instagram posts,” in EMNLP , 2019
2019
Earlier work this paper cites.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in ACL , 2019
2019
Earlier work this paper cites.
J. Lee, S. Kim, S. Kim, J. Park, and K. Sohn, “Context-aware emotion recognition networks,” in ICCV , 2019
2019
Earlier work this paper cites.
J. Jiang, Y. Chen, X. Meng, L. Wang, and K. Li, “A novel density peaks clustering algorithm based on k nearest neighbors for improving assignment process,” Physica A: Statistical Mechanics and its Applications , 2019
2019
Earlier work this paper cites.
D. Demszky, D. Movshovitz-Attias, J. Ko, A. Cowen, G. Nemade, and S. Ravi, “Goemotions: A dataset of fine-grained emotions,” arXiv preprint , 2020
2020
Cited alongside, same era.
I. Casanueva, T. Temcinas, D. Gerz, M. Henderson, and I. Vulic, “Efficient intent detection with dual sentence encoders,” in ACL WorkShop , 2020
2020
Cited alongside, same era.
D. Hazarika, R. Zimmermann, and S. Poria, “Misa: Modality-invariant and-specific representations for multimodal sentiment analysis,” in ACM MM , 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in NIPS , 2020
2020
Cited alongside, same era.
M. Jia, Z. Wu, A. Reiter, C. Cardie, S. Belongie, and S.-N. Lim, “Intentonomy: a dataset and study towards human intent understanding,” in CVPR , 2021
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in ICML , 2023
2023
Later among the works it cites.
P. Jin, R. Takanobu, C. Zhang, X. Cao, and L. Yuan, “Chat-univi: Unified visual representation empowers large language models with image and video understanding,” arXiv preprint , 2023
2023
Later among the works it cites.
C. Lyu, M. Wu, L. Wang, X. Huang, B. Liu, Z. Du, S. Shi, and Z. Tu, “Macaw-llm: Multi-modal language modeling with image, audio, video, and text integration,” arXiv preprint , 2023
2023
Later among the works it cites.
J. Han, K. Gong, Y. Zhang, J. Wang, K. Zhang, D. Lin, Y. Qiao, P. Gao, and X. Yue, “Onellm: One framework to align all modalities with language,” arXiv preprint , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Cited alongside, same era.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” arXiv preprint , 2021
2021
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” NIPS , 2022
2022
Cited alongside, same era.
A. Jia, Y. He, Y. Zhang, S. Uprety, D. Song, and C. Lioma, “Beyond emotion: A multi-modal dataset for human desire understanding,” in NACCL , 2022
2022
Cited alongside, same era.
H. Zhang, H. Xu, X. Wang, Q. Zhou, S. Zhao, and J. Teng, “Mintrec: A new dataset for multimodal intent recognition,” in ACM MM , 2022
2022
Cited alongside, same era.
D. Yang, H. Kuang, S. Huang, and L. Zhang, “Learning modality-specific and-agnostic representations for asynchronous multimodal language sequences,” in ACM MM , 2022
2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv preprint , 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
L. Hyun, K. Sung-Bin, S. Han, Y. Yu, and T.-H. Oh, “Smile: Multimodal dataset for understanding laughter in video with language models,” arXiv preprint , 2023
2023
Later among the works it cites.
M.-H. Van and X. Wu, “Detecting and correcting hate speech in multimodal memes with large visual language model,” arXiv preprint , 2023
2023
Later among the works it cites.
L. Qin, S. Huang, Q. Chen, C. Cai, Y. Zhang, B. Liang, W. Che, and R. Xu, “Mmsd2. 0: Towards a reliable multi-modal sarcasm detection system,” arXiv preprint , 2023
2023
Later among the works it cites.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez et al. , “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” https://vicuna.lmsys.org , 2023
2023
Later among the works it cites.
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. N. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” NIPS , 2024
2024
Closest in time.
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma et al. , “Scaling instruction-finetuned language models,” JMLR , 2024
2024
Closest in time.
K. Sun, Z. Xie, M. Ye, and H. Zhang, “Contextual augmented global contrast for multimodal intent recognition,” in CVPR , 2024
2024
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
W. Hu, Y. Xu, Y. Li, W. Li, Z. Chen, and Z. Tu, “Bliva: A simple multimodal llm for better handling of text-rich visual questions,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 3, 2024, pp. 2256–2264
2024
Closest in time.