Fetching the paper…
Reading the bibliography…
The rapid advancements in the development of multimodal large language models (MLLMs) have consistently led to new breakthroughs on various benchmarks.
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P. and Zitnick, C. L. [2014], Microsoft coco: Common objects in context
2014
Earlier work this paper cites.
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L. and Parikh, D. [2015], Vqa: Visual question answering
2015
Earlier work this paper cites.
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J. and Lazebnik, S. [2015], Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
2015
Earlier work this paper cites.
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D. and Parikh, D. [2017], Making the v in vqa matter: Elevating the role of image understanding in visual question answering
2017
Earlier work this paper cites.
Kafle, K. and Kanan, C. [2017], An analysis of visual question answering algorithms
2017
Earlier work this paper cites.
Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S. and Anderson, P. [2019], Nocaps: Novel object captioning at scale
2019
Earlier work this paper cites.
Hossain, M. Z., Sohel, F., Shiratuddin, M. F. and Laga, H. [2019], ‘A comprehensive survey of deep learning for image captioning’, ACM Computing Surveys (CsUR)
2019
Earlier work this paper cites.
Hudson, D. A. and Manning, C. D. [2019], Gqa: A new dataset for real-world visual reasoning and compositional question answering
2019
Earlier work this paper cites.
Singh, A., Natarjan, V., Shah, M., Jiang, Y., Chen, X., Parikh, D. and Rohrbach, M. [2019], Towards vqa models that can read
2019
Earlier work this paper cites.
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A. and Choi, Y. [2019], ‘Hellaswag: Can a machine really finish your sentence?’
2019
Earlier work this paper cites.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T. et al. [2020], ‘An image is worth 16x16 words: Transformers for image recognition at scale’, ICLR
2020
Earlier work this paper cites.
Rahman, W., Hasan, M. K., Lee, S., Zadeh, A., Mao, C., Morency, L.-P. and Hoque, E. [2020], Integrating multimodal information in large pretrained transformers
2020
Earlier work this paper cites.
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M. et al. [2021], ‘Training verifiers to solve math word problems’
2021
Earlier work this paper cites.
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D. and Steinhardt, J. [2021], ‘Measuring massive multitask language understanding’, ICLR
2021
Earlier work this paper cites.
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D. and Steinhardt, J. [2021], ‘Measuring mathematical problem solving with the math dataset’, NeurIPS
2021
Earlier work this paper cites.
Mathew, M., Karatzas, D. and Jawahar, C. V. [2021], ‘Docvqa: A dataset for vqa on document images’
2021
Earlier work this paper cites.
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X. et al. [2021], ‘Towards vqa models that can read’
2021
Earlier work this paper cites.
Desai, P., Chakraborty, T. and Akhtar, M. S. [2022], ‘Nice perfume. how long did you marinate in it? multimodal sarcasm explanation’, AAAI
2022
Earlier work this paper cites.
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P. and Kalyan, A. [2022], Learn to explain: Multimodal reasoning via thought chains for science question answering
2022
Earlier work this paper cites.
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W. et al. [2022], ‘Challenging big-bench tasks and whether chain-of-thought can solve them’
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E. et al. [2023], ‘Sparks of artificial general intelligence: Early experiments with gpt-4’, arXiv preprint arXiv: 2303.12712
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
Ghandi, T., Pourreza, H. and Mahyar, H. [2023], ‘Deep learning approaches on image captioning: A review’, ACM Computing Surveys
2023
Cited alongside, same era.
Hessel, J., Marasovic, A., Hwang, J. D., Lee, L., Da, J., Zellers, R., Mankoff, R. and Choi, Y. [2023], Do androids laugh at electric sheep? humor “understanding” benchmarks from the new yorker caption contest
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
Chen, L., Li, J., Dong, X., Zhang, P., Zang, Y. et al. [2024], ‘Are we on the right way for evaluating large vision-language models?’
2024
Closest in time.
2024
Closest in time.
Cui, C., Ma, Y., Cao, X., Ye, W., Zhou, Y., Liang, K., Chen, J., Lu, J., Yang, Z., Liao, K.-D. et al. [2024], A survey on multimodal large language models for autonomous driving
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Lu, S., Liu, M., Yin, L., Yin, Z., Liu, X. and Zheng, W. [2023], ‘The multi-modal fusion in visual question answering: a review of attention mechanisms’, PeerJ Computer Science
2023
Cited alongside, same era.
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P. N. and Hoi, S. [2024], ‘Instructblip: Towards general-purpose vision-language models with instruction tuning’, NIPS
2024
Closest in time.
2024
Closest in time.
He, Z., Wu, X., Zhou, P., Xuan, R., Liu, G., Yang, X., Zhu, Q. and Huang, H. [2024], ‘Cmmu: A benchmark for chinese multi-modal multi-type question understanding and reasoning’
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Li, H., Zhang, Y., Koto, F., Yang, Y., Zhao, H. et al. [2024], ‘Cmmlu: Measuring massive multitask language understanding in chinese’
2024
Closest in time.
Liu, H., Zheng, Z., Qiao, Y., Duan, H., Fei, Z., Zhou, F., Zhang, W. et al. [2024], ‘Mathbench: Evaluating the theory and application proficiency of llms with a hierarchical mathematics benchmark’
2024
Closest in time.
Liu, Y., Li, Z., Yang, B., Li, C., Yin, X. et al. [2024], ‘On the hidden mystery of ocr in large multimodal models’
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Qian, T., Chen, J., Zhuo, L., Jiao, Y. and Jiang, Y.-G. [2024], Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario
2024
Closest in time.
Strachan, J. W., Albergo, D., Borghini, G., Pansardi, O., Scaliti, E., Gupta, S., Saxena, K., Rufo, A. et al. [2024], ‘Testing theory of mind in large language models and humans’, Nature Human Behaviour
2024
Closest in time.
2024
Closest in time.
Yang, Y., Li, Z., Dong, Q., Xia, H. and Sui, Z. [2024], ‘Can large multimodal models uncover deep semantics behind images?’
2024
Closest in time.
2024
Closest in time.
Zhang, G., Du, X., Chen, B., Liang, Y., Luo, T., Zheng, T., Zhu, K., Cheng, Y. et al. [2024], ‘Cmmmu: A chinese massive multi-discipline multimodal understanding benchmark’
2024
Closest in time.
Zhong, S., Huang, Z., Gao, S., Wen, W., Lin, L., Zitnik, M. and Zhou, P. [2024], ‘Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation’
2024
Closest in time.
2025
Closest in time.