Fetching the paper…
Reading the bibliography…
This paper focuses on monolithic Multimodal Large Language Models (MLLMs), which integrate visual encoding and language decoding into a single model.
S. E. Yuksel, J. N. Wilson, and P. D. Gader, “Twenty years of mixture of experts,” IEEE transactions on neural networks and learning systems , vol. 23, no. 8, pp. 1177–1193, 2012
2012
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” TACL , vol. 2, pp. 67–78, 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi, “A diagram is worth a dozen images,” in ECCV , 2016, pp. 235–251
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in ECCV , vol. 9906, 2016, pp. 69–85
2016
Earlier work this paper cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in CVPR , 2016, pp. 11–20
2016
Earlier work this paper cites.
A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+ d: A large scale dataset for 3d human activity analysis,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 1010–1019
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
B. Shi, C. Yao, M. Liao, M. Yang, P. Xu, L. Cui, S. Belongie, S. Lu, and X. Bai, “Icdar2017 competition on reading chinese text in the wild (rctw-17),” in ICDAR , vol. 1, 2017, pp. 1429–1434
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in visual question answering,” in CVPR , 2017, pp. 6325–6334
2017
Earlier work this paper cites.
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra, “Visual dialog,” in CVPR , 2017, pp. 326–335
2017
Earlier work this paper cites.
A. Kembhavi, M. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Hajishirzi, “Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,” in CVPR , 2017, pp. 4999–5007
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei, “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” IJCV , vol. 123, no. 1, pp. 32–73, 2017
2017
Earlier work this paper cites.
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele, “Movie description,” International Journal of Computer Vision , 2017. [Online]. Available: http://link.springer.com/article/10.1007/s11263-016-0987-1?wt_mc=Internal.Event.1.SEM.ArticleAuthorOnlineFirst
2017
Earlier work this paper cites.
C. Clark and M. Gardner, “Simple and effective multi-paragraph reading comprehension,” in ACL , 2018, pp. 845–855
2018
Earlier work this paper cites.
K. Kafle, B. Price, S. Cohen, and C. Kanan, “Dvqa: Understanding data visualizations via question answering,” in CVPR , 2018, pp. 5648–5656
2018
Earlier work this paper cites.
B. Zhang and R. Sennrich, “Root mean square layer normalization,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun, “Objects365: A large-scale, high-quality dataset for object detection,” in ICCV , 2019, pp. 8430–8439
2019
Earlier work this paper cites.
Y. Sun, Z. Ni, C.-K. Chng, Y. Liu, C. Luo, C. C. Ng, J. Han, E. Ding, J. Liu, D. Karatzas et al. , “Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt,” in ICDAR , 2019, pp. 1557–1562
2019
Earlier work this paper cites.
A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, E. Valveny, C. Jawahar, and D. Karatzas, “Scene text visual question answering,” in ICCV , 2019, pp. 4291–4301
2019
Earlier work this paper cites.
R. Zhang, Y. Zhou, Q. Jiang, Q. Song, N. Li, K. Zhou, L. Wang, D. Wang, M. Liao, M. Yang et al. , “Icdar 2019 robust reading challenge on reading chinese text on signboard,” in ICDAR , 2019, pp. 1577–1581
2019
Earlier work this paper cites.
C. K. Chng, Y. Liu, Y. Sun, C. C. Ng, C. Luo, Z. Ni, C. Fang, S. Zhang, J. Han, E. Ding et al. , “Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art,” in ICDAR , 2019, pp. 1571–1576
2019
Earlier work this paper cites.
T.-L. Yuan, Z. Zhu, K. Xu, C.-J. Li, T.-J. Mu, and S.-M. Hu, “A large chinese text dataset in the wild,” Journal of Computer Science and Technology , vol. 34, pp. 509–521, 2019
2019
Earlier work this paper cites.
D. A. Hudson and C. D. Manning, “GQA: A new dataset for real-world visual reasoning and compositional question answering,” in CVPR , 2019, pp. 6700–6709
2019
Earlier work this paper cites.
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “Ok-vqa: A visual question answering benchmark requiring external knowledge,” in CVPR , 2019, pp. 3195–3204
2019
Earlier work this paper cites.
S. Shah, A. Mishra, N. Yadati, and P. P. Talukdar, “Kvqa: Knowledge-aware visual question answering,” in AAAI , vol. 33, no. 01, 2019, pp. 8876–8884
2019
Earlier work this paper cites.
A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty, “Ocr-vqa: Visual question answering by reading text in images,” in ICDAR , 2019, pp. 947–952
2019
Earlier work this paper cites.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards VQA models that can read,” in CVPR , 2019
2019
Earlier work this paper cites.
Z. Huang, K. Chen, J. He, X. Bai, D. Karatzas, S. Lu, and C. Jawahar, “Icdar2019 competition on scanned receipt ocr and information extraction,” in 2019 International Conference on Document Analysis and Recognition (ICDAR) . IEEE, 2019, pp. 1516–1520
2019
Earlier work this paper cites.
J.-P. T. Guillaume Jaume, Hazim Kemal Ekenel, “Funsd: A dataset for form understanding in noisy scanned documents,” in Accepted to ICDAR-OST , 2019
2019
Earlier work this paper cites.
O. Sidorov, R. Hu, M. Rohrbach, and A. Singh, “Textcaps: A dataset for image captioning with reading comprehension,” in ECCV , vol. 12347, 2020, pp. 742–758
2020
Earlier work this paper cites.
N. Methani, P. Ganguly, M. M. Khapra, and P. Kumar, “Plotqa: Reasoning over scientific plots,” in WACV , 2020, pp. 1527–1536
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in ICML , vol. 139, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Earlier work this paper cites.
A. Singh, G. Pang, M. Toh, J. Huang, W. Galuba, and T. Hassner, “Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text,” in CVPR , 2021, pp. 8802–8812
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
B. Wu, S. Yu, Z. Chen, J. B. Tenenbaum, and C. Gan, “Star: A benchmark for situated reasoning in real-world videos,” in Thirty-fifth Conference on Neural Information Processing Systems (NeurIPS) , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Li, D. Li, C. Xiong, and S. C. H. Hoi, “BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,” in ICLR , vol. 162, 2022, pp. 12 888–12 900
2022
Earlier work this paper cites.
H. Bao, W. Wang, L. Dong, Q. Liu, O. K. Mohammed, K. Aggarwal, S. Som, S. Piao, and F. Wei, “Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,” Advances in Neural Information Processing Systems , vol. 35, pp. 32 897–32 912, 2022
2022
Earlier work this paper cites.
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som, and F. Wei, “Image as a foreign language: Beit pretraining for all vision and vision-language tasks,” arXiv: 2208.10442 , 2022
2022
Earlier work this paper cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al. , “Laion-5b: An open large-scale dataset for training next generation image-text models,” NeurIPS , vol. 35, pp. 25 278–25 294, 2022
2022
Earlier work this paper cites.
M. Byeon, B. Park, H. Kim, S. Lee, W. Baek, and S. Kim, “Coyo-700m: Image-text pair dataset,” https://github.com/kakaobrain/coyo-dataset , 2022
2022
Cited alongside, same era.
J. Gu, X. Meng, G. Lu, L. Hou, N. Minzhe, X. Liang, L. Yao, R. Huang, W. Zhang, X. Jiang et al. , “Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark,” NeurIPS , vol. 35, pp. 26 418–26 431, 2022
2022
Cited alongside, same era.
C. Schuhmann, A. Köpf, R. Vencu, T. Coombes, and R. Beaumont, “Laion coco: 600m synthetic captions from laion2b-en.” https://laion.ai/blog/laion-coco/ , 2022
2022
Cited alongside, same era.
G. Kim, T. Hong, M. Yim, J. Nam, J. Park, J. Yim, W. Hwang, S. Yun, D. Han, and S. Park, “Ocr-free document understanding transformer,” in ECCV , 2022
2022
Cited alongside, same era.
W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, and L. Wang, “Mm-vet: Evaluating large multimodal models for integrated capabilities,” arXiv: 2308.02490 , 2023
2023
Later among the works it cites.
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun et al. , “Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,” arXiv: 2311.16502 , 2023
2023
Later among the works it cites.
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao, “Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts,” arXiv: 2310.02255 , 2023
2023
Later among the works it cites.
B. Li, R. Wang, G. Wang, Y. Ge, Y. Ge, and Y. Shan, “Seed-bench: Benchmarking multimodal llms with generative comprehension,” arXiv: 2307.16125 , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar, “Infographicvqa,” in WACV , 2022, pp. 1697–1706
2022
Cited alongside, same era.
P. Lu, S. Mishra, T. Xia, L. Qiu, K. Chang, S. Zhu, O. Tafjord, P. Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” in NeurIPS , 2022
2022
Cited alongside, same era.
J. Cao and J. Xiao, “An augmented benchmark dataset for geometric question answering through dual parallel text encoding,” in COLING , 2022, pp. 1511–1520
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi, “A-okvqa: A benchmark for visual question answering using world knowledge,” in ECCV , 2022, pp. 146–162
2022
Cited alongside, same era.
P. Lerner, O. Ferret, C. Guinaudeau, H. Le Borgne, R. Besançon, J. G. Moreno, and J. Lovón Melgarejo, “Viquae, a dataset for knowledge-based visual question answering about named entities,” in SIGIR , 2022, pp. 3108–3120
2022
Cited alongside, same era.
2023
Later among the works it cites.
T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob et al. , “Hallusionbench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models,” arXiv: 2310.14566 , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Contributors, “Opencompass: A universal evaluation platform for foundation models,” https://github.com/open-compass/opencompass , 2023
2023
Later among the works it cites.
LMDeployContributors, “Lmdeploy: A toolkit for compressing, deploying, and serving llm,” https://github.com/InternLM/lmdeploy , 2023
2023
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Wang, X. Zhang, Z. Luo, Q. Sun, Y. Cui, J. Wang, F. Zhang, Y. Wang, Z. Li, Q. Yu, Y. Zhao, Y. Ao, X. Min, T. Li, B. Wu, B. Zhao, B. Zhang, L. Wang, G. Liu, Z. He, X. Yang, J. Liu, Y. Lin, T. Huang, and Z. Wang, “Emu3: Next-token prediction is all you need,” arXiv: 2409.18869 , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee, “Llava-next: Improved reasoning, ocr, and world knowledge,” 2024. [Online]. Available: https://llava-vl.github.io/blog/2024-01-30-llava-next/
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Gao, Z. Chen, E. Cui, Y. Ren, W. Wang, J. Zhu, H. Tian, S. Ye, J. He, X. Zhu et al. , “Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance,” Visual Intelligence , vol. 2, no. 1, pp. 1–17, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
W. Wang, M. Shi, Q. Li, W. Wang, Z. Huang, L. Xing, Z. Chen, H. Li, X. Zhu, Z. Cao et al. , “The all-seeing project: Towards panoptic visual recognition and understanding of the open world,” in ICLR , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing et al. , “Judging llm-as-a-judge with mt-bench and chatbot arena,” NeurIPS , vol. 36, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Li, Y. Zhang, C. Wang, Z. Zhong, Y. Chen, R. Chu, S. Liu, and J. Jia, “Mini-gemini: Mining the potential of multi-modality vision language models,” arXiv: 2403.18814 , 2024
2024
Later among the works it cites.
B. McKinzie, Z. Gan, J. Fauconnier, S. Dodge, B. Zhang, P. Dufter, D. Shah, X. Du, F. Peng, F. Weers, A. Belyi, H. Zhang, K. Singh, D. Kang, A. Jain, H. Hè, M. Schwarzer, T. Gunter, X. Kong, A. Zhang, J. Wang, C. Wang, N. Du, T. Lei, S. Wiseman, G. Yin, M. Lee, Z. Wang, R. Pang, P. Grasch, A. Toshev, and Y. Yang, “MM1: methods, analysis & insights from multimodal LLM pre-training,” arXiv: 2403.09611 , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
G. Luo, X. Yang, W. Dou, Z. Wang, J. Liu, J. Dai, Y. Qiao, and X. Zhu, “Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training,” in CVPR , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.