Fetching the paper…
Reading the bibliography…
Recent advancements in multimodal large language models (MLLMs) have aimed to integrate and interpret data across diverse modalities.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
H. Chernoff · 1952
Earlier work this paper cites.
What size test set gives good error rate estimates?
I. Guyon, J. Makhoul, R. Schwartz, and V. Vapnik · 1998
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang · 2017
Earlier work this paper cites.
Aishell-2: Transforming mandarin asr research into industrial scale
J. Du, X. Na, X. Liu, and H. Bu · 2018
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy · 2020
Earlier work this paper cites.
Clotho: An audio captioning dataset
K. Drossos, S. Lipping, and T. Virtanen · 2020
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention. arxiv 2020
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2020
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Emergency vehicles audio detection and localization in autonomous driving
H. Sun, X. Liu, K. Xu, J. Miao, and Q. Luo · 2021
Earlier work this paper cites.
Comparison of feature extraction methods for sound-based classification of honey bee activity
A. Terenzi, N. Ortolani, I. Nolasco, E. Benetos, and S. Cecchi · 2021
Earlier work this paper cites.
Pano-avqa: Grounded audio-visual question answering on 360deg videos
H. Yun, Y. Yu, W. Yang, K. Lee, and G. Kim · 2021
Earlier work this paper cites.
Beats: Audio pre-training with acoustic tokenizers
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, and F. Wei · 2022
Earlier work this paper cites.
Vocalsound: A dataset for improving human vocal sounds recognition
Y. Gong, J. Yu, and J. Glass · 2022
Earlier work this paper cites.
Learning to answer questions in dynamic audio-visual scenarios
G. Li, Y. Wei, Y. Tian, C. Xu, J.-R. Wen, and D. Hu · 2022
Earlier work this paper cites.
Robust speech recognition via large-scale weak supervision. arxiv 2022
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2022
Earlier work this paper cites.
Avqa: A dataset for audio-visual question answering on videos
P. Yang, X. Wang, X. Duan, H. Chen, R. Hou, C. Jin, and W. Zhu · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Qwen-vl: A frontier large vision-language model with versatile abilities
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Cited alongside, same era.
Valor: Vision-audio-language omni-perception pretraining model and dataset
S. Chen, X. He, L. Guo, X. Zhu, W. Wang, J. Tang, and J. Liu · 2023
Cited alongside, same era.
Musilingo: Bridging music and text with pre-trained language models for music captioning and query response
Z. Deng, Y. Ma, Y. Liu, R. Guo, G. Zhang, W. Chen, W. Huang, and E. Benetos · 2023
Cited alongside, same era.
What matters when building vision-language models?
H. Laurençon, L. Tronchon, M. Cord, and V. Sanh · 2024
Closest in time.
Llava-onevision: Easy visual task transfer
B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, Y. Li, Z. Liu, and C. Li · 2024
Closest in time.
A survey on benchmarks of multimodal large language models
J. Li and W. Lu · 2024
Closest in time.
Learning from taxonomy: Multi-label few-shot classification for everyday sound recognition
J. Liang, H. Phan, and E. Benetos · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. Glass · 2023
Cited alongside, same era.
Cr19: A framework for preliminary detection of covid-19 in cough audio signals using machine learning algorithms for automated medical diagnosis applications
E. E.-D. Hemdan, W. El-Shafai, and A. Sayed · 2023
Cited alongside, same era.
The impact of multimodal large language models on health care’s future
B. Meskó · 2023
Cited alongside, same era.
Recent advancements in multimodal human–robot interaction
H. Su, W. Qi, J. Chen, C. Yang, J. Sandoval, and M. A. Laribi · 2023
Cited alongside, same era.
Salmonn: Towards generic hearing abilities for large language models
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Cited alongside, same era.
On decoder-only architecture for speech-to-text and large language model integration
J. Wu, Y. Gaur, Z. Chen, L. Zhou, Y. Zhu, T. Wang, J. Li, S. Liu, B. Ren, L. Liu, et al · 2023
Cited alongside, same era.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov · 2023
Cited alongside, same era.
Foundation models for music: A survey
Y. Ma, A. Øland, A. Ragni, B. M. Del Sette, C. Saitis, C. Donahue, C. Lin, C. Plachouras, E. Benetos, E. Quinton, et al · 2024
Closest in time.
Reka core, flash, and edge: A series of powerful multimodal language models
A. Ormazabal, C. Zheng, C. d. M. d’Autume, D. Yogatama, D. Fu, D. Ong, E. Chen, E. Lamprecht, H. Pham, I. Ong, et al · 2024
Closest in time.
video-salmonn: Speech-enhanced audio-visual large language models
G. Sun, W. Yu, C. Tang, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, Y. Wang, and C. Zhang · 2024
Closest in time.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024
S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, A. Wang, R. Fergus, Y. LeCun, and S. Xie · 2024
Closest in time.
Air-bench: Benchmarking large audio-language models via generative comprehension
Q. Yang, J. Xu, W. Liu, Y. Chu, Z. Jiang, X. Zhou, Y. Leng, Y. Lv, Z. Zhao, C. Zhou, and J. Zhou · 2024
Closest in time.
Yi: Open foundation models by 01. ai, 2024
A. Young, B. Chen, C. Li, C. Huang, G. Zhang, G. Zhang, H. Li, J. Zhu, J. Chen, J. Chang, et al · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al · 2024
Closest in time.
Anygpt: Unified multimodal llm with discrete sequence modeling
J. Zhan, J. Dai, J. Ye, Y. Zhou, D. Zhang, Z. Liu, X. Zhang, R. Yuan, G. Zhang, L. Li, et al · 2024
Closest in time.
Cmmmu: A chinese massive multi-discipline multimodal understanding benchmark
G. Zhang, X. Du, B. Chen, Y. Liang, T. Luo, T. Zheng, K. Zhu, Y. Cheng, C. Xu, S. Guo, H. Zhang, X. Qu, J. Wang, R. Yuan, Y. Li, Z. Wang, Y. Liu, Y.-H. Tsai, F. Zhang, C. Lin, W. Huang, W. Chen, and J. Fu · 2024
Closest in time.
Vita-1.5: Towards gpt-4o level real-time vision and speech interaction
C. Fu, H. Lin, X. Wang, Y.-F. Zhang, Y. Shen, X. Liu, Y. Li, Z. Long, H. Gao, K. Li, et al · 2025
Closest in time.
Baichuan-omni-1.5 technical report
Y. Li, J. Liu, T. Zhang, S. Chen, T. Li, Z. Li, L. Liu, L. Ming, G. Dong, D. Pan, et al · 2025
Closest in time.
R. Luo, T.-E. Lin, H. Zhang, Y. Wu, X. Liu, M. Yang, Y. Li, L. Chen, J. Li, L. Zhang, et al · 2025
Closest in time.
Cmi-bench: A comprehensive benchmark for evaluating music instruction following
Y. Ma, S. Li, J. Yu, E. Benetos, and A. Maezawa · 2025
Closest in time.
Qwen2.5-omni technical report, 2025
J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Dang, B. Zhang, X. Wang, Y. Chu, and J. Lin · 2025
Closest in time.