Fetching the paper…
Reading the bibliography…
In an era defined by the explosive growth of data and rapid technological advancements, Multimodal Large Language Models (MLLMs) stand at the forefront of artificial intelligence (AI) systems.
S. K. Gaikwad, B. W. Gawali, and P. Yannawar, “A review on speech recognition technique,” International Journal of Computer Applications , vol. 10, no. 3, pp. 16–24, 2010
2010
Earlier work this paper cites.
2013
Earlier work this paper cites.
G. Laccetti, R. Montella, C. Palmieri, and V. Pelliccia, “The high performance internet of things: using gvirtus to share high-end gpus with arm based cluster computing nodes,” in International Conference on Parallel Processing and Applied Mathematics . Springer, 2013, pp. 734–744
2013
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. Xu, T. Mei, T. Yao, and Y. Rui, “Msr-vtt: A large video description dataset for bridging video and language,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 5288–5296
2016
Earlier work this paper cites.
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang, “Video question answering via gradually refined attention over appearance and motion,” in Proceedings of the 25th ACM international conference on Multimedia , 2017, pp. 1645–1653
2017
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 2556–2565
2018
Earlier work this paper cites.
P. Bahar, T. Bieschke, and H. Ney, “A comparative study on end-to-end speech to text translation,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 792–799
2019
Earlier work this paper cites.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of NAACL-HLT , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition,” Interspeech 2019 , 2019
2019
Earlier work this paper cites.
H. Gao, J. Pei, and H. Huang, “Progan: Network embedding via proximity generative adversarial network,” in Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2019, pp. 1308–1316
2019
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019, pp. 119–132
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
K. Gadzicki, R. Khamsehashari, and C. Zetzsche, “Early vs late fusion in multimodal convolutional neural networks,” in 2020 IEEE 23rd international conference on information fusion (FUSION) . IEEE, 2020, pp. 1–6
2020
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Communications of the ACM , vol. 63, no. 11, pp. 139–144, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Malik, M. K. Malik, K. Mehmood, and I. Makhdoom, “Automatic speech recognition: a survey,” Multimedia Tools and Applications , vol. 80, pp. 9411–9457, 2021
2021
Earlier work this paper cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM transactions on audio, speech, and language processing , vol. 29, pp. 3451–3460, 2021
2021
Earlier work this paper cites.
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 1728–1738
2021
Earlier work this paper cites.
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut, “Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 3558–3568
2021
Earlier work this paper cites.
D. Neimark, O. Bar, M. Zohar, and D. Asselmann, “Video transformer network,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 3163–3172
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Xiao, X. Shang, A. Yao, and T.-S. Chua, “Next-qa: Next phase of question-answering to explaining temporal actions,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 9777–9786
2021
Earlier work this paper cites.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video diffusion models,” Advances in Neural Information Processing Systems , vol. 35, pp. 8633–8646, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Wu, W. Gan, Z. Chen, S. Wan, and S. Y. Philip, “Multimodal large language models: A survey,” in 2023 IEEE International Conference on Big Data (BigData) . IEEE, 2023, pp. 2247–2256
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Wang, Z. Liu, L. Zhao, Z. Wu, C. Ma, S. Yu, H. Dai, Q. Yang, Y. Liu, S. Zhang et al. , “Review of large vision models and visual prompt engineering,” Meta-Radiology , p. 100047, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Liu, T. Han, S. Ma, J. Zhang, Y. Yang, J. Tian, H. He, A. Li, M. He, Z. Liu et al. , “Summary of chatgpt-related research and perspective towards the future of large language models,” Meta-Radiology , p. 100017, 2023
2023
Cited alongside, same era.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing et al. , “Judging llm-as-a-judge with mt-bench and chatbot arena,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
M. Mulimani and A. Mesaros, “Class-incremental learning for multi-label audio classification,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 916–920
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning . PMLR, 2023, pp. 28 492–28 518
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, “Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
R. Huang, M. Li, D. Yang, J. Shi, X. Chang, Z. Ye, Y. Wu, Z. Hong, J. Huang, J. Liu et al. , “Audiogpt: Understanding and generating speech, music, sound, and talking head,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 21, 2024, pp. 23 802–23 804
2024
Closest in time.
C. Li, Z. Gan, Z. Yang, J. Yang, L. Li, L. Wang, J. Gao et al. , “Multimodal foundation models: From specialists to general-purpose assistants,” Foundations and Trends® in Computer Graphics and Vision , vol. 16, no. 1-2, pp. 1–214, 2024
2024
Closest in time.
2024
Closest in time.
Y. Li, M. E. H. Daho, P.-H. Conze, R. Zeghlache, H. Le Boité, R. Tadayoni, B. Cochener, M. Lamard, and G. Quellec, “A review of deep learning-based information fusion techniques for multimodal medical image classification,” Computers in Biology and Medicine , p. 108635, 2024
2024
Closest in time.
D. Caffagni, F. Cocchi, N. Moratelli, S. Sarto, M. Cornia, L. Baraldi, and R. Cucchiara, “Wiki-llava: Hierarchical retrieval-augmented generation for multimodal llms,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 1818–1826
2024
Closest in time.
2024
Closest in time.
Z. Li, Q. Xu, D. Zhang, H. Song, Y. Cai, Q. Qi, R. Zhou, J. Pan, Z. Li, V. T. Vu et al. , “Groundinggpt: Language enhanced multi-modal grounding model.” CoRR , 2024
2024
Closest in time.
X. Wang, B. Zhuang, and Q. Wu, “Modaverse: Efficiently transforming modalities with llms,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 606–26 616
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
“Ideogram 1.0, Feb 2024.” [Online]. Available: https://about.ideogram.ai/1.0
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Lin, H. Yin, W. Ping, P. Molchanov, M. Shoeybi, and S. Han, “Vila: On pre-training for visual language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 689–26 699
2024
Closest in time.
Y. Zhang, B. Li, h. Liu, Y. j. Lee, L. Gui, D. Fu, J. Feng, Z. Liu, and C. Li, “Llava-next: A strong zero-shot video understanding model,” April 2024. [Online]. Available: https://llava-vl.github.io/blog/2024-04-30-llava-next-video/
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo et al. , “Mvbench: A comprehensive multi-modal video understanding benchmark,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 22 195–22 206
2024
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.
Z. Tang, Z. Yang, C. Zhu, M. Zeng, and M. Bansal, “Any-to-any generation via composable diffusion,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
H. Fei, S. Wu, H. Zhang, T.-S. Chua, and S. Yan, “Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing,” 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. Pan, H. Li, X. Zhou, J. Jiao, C. Zhu, and Q. Zhang, “Research on pig sound recognition based on deep neural network and hidden markov models,” Sensors , vol. 24, no. 4, p. 1269, 2024
2024
Closest in time.
A. Zhang, H. Fei, Y. Yao, W. Ji, L. Li, Z. Liu, and T.-S. Chua, “Vpgtrans: Transfer visual prompt generator across llms,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 296–26 306
2024
Closest in time.