Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) excel in solving text-based mathematical problems, but they struggle with mathematical diagrams since they are primarily trained on natural scene images.
Analysing mathematical reasoning abilities of neural models
Saxton, D.; Grefenstette, E.; Hill, F.; and Kohli, P. 2019 · 1904
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Learning Transferable Visual Models from Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Chen, W.; Ma, X.; Wang, X.; and Cohen, W. W. 2022 · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Gray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Earlier work this paper cites.
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023 · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, G. 2023 · 2023
Earlier work this paper cites.
Tora: A tool-integrated reasoning agent for mathematical problem solving
Gou, Z.; Shao, Z.; Gong, Y.; Yang, Y.; Huang, M.; Duan, N.; Chen, W.; et al. 2023 · 2023
Earlier work this paper cites.
Visual Instruction Tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Earlier work this paper cites.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L.; Jiang, W.; Shi, H.; Yu, J.; Liu, Z.; Zhang, Y.; Kwok, J. T.; Li, Z.; Weller, A.; and Liu, W. 2023 · 2023
Cited alongside, same era.
Mammoth: Building math generalist models through hybrid instruction tuning
Yue, X.; Qu, X.; Zhang, G.; Fu, Y.; Huang, W.; Sun, H.; Su, Y.; and Chen, W. 2023 · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
Zhai, X.; Mustafa, B.; Kolesnikov, A.; and Beyer, L. 2023 · 2023
Cited alongside, same era.
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
Cai, S.; Bao, K.; Guo, H.; Zhang, J.; Song, J.; and Zheng, B. 2024 · 2024
Cited alongside, same era.
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Qiao, R.; Tan, Q.; Dong, G.; Wu, M.; Sun, C.; Song, X.; GongQue, Z.; Lei, S.; Wei, Z.; Zhang, M.; et al. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M.; Savinov, N.; Teplyashin, D.; Lepikhin, D.; Lillicrap, T.; Alayrac, J.-b.; Soricut, R.; Lazaridou, A.; Firat, O.; Schrittwieser, J.; et al. 2024 · 2024
Closest in time.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Zhang, M.; Li, Y.; Wu, Y.; and Guo, D. 2024 · 2024
Closest in time.
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
Shi, W.; Hu, Z.; Bin, Y.; Liu, J.; Yang, Y.; Ng, S.-K.; Bing, L.; and Lee, R. K.-W. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, Z.; Wang, W.; Tian, H.; Ye, S.; Gao, Z.; Cui, E.; Tong, W.; Hu, K.; Luo, J.; Ma, Z.; et al. 2024 · 2024
Cited alongside, same era.
Dong, X.; Zhang, P.; Zang, Y.; Cao, Y.; Wang, B.; Ouyang, L.; Wei, X.; Zhang, S.; Duan, H.; Cao, M.; et al. 2024 · 2024
Cited alongside, same era.
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Gao, P.; Zhang, R.; Liu, C.; Qiu, L.; Huang, S.; Lin, W.; Zhao, S.; Geng, S.; Lin, Z.; Jin, P.; et al. 2024 · 2024
Cited alongside, same era.
NuminaMath
LI, J.; Beeching, E.; Tunstall, L.; Lipkin, B.; Soletskyi, R.; Huang, S. C.; Rasul, K.; Yu, L.; Jiang, A.; Shen, Z.; Qin, Z.; Dong, B.; Zhou, L.; Fleureau, Y.; Lample, G.; and Polu, S. 2024 · 2024
Cited alongside, same era.
Augmenting math word problems via iterative question composing
Liu, H.; and Yao, A. C.-C. 2024 · 2024
Cited alongside, same era.
Orca-math: Unlocking the potential of slms in grade school math
Mitra, A.; Khanpour, H.; Rosset, C.; and Awadallah, A. 2024 · 2024
Cited alongside, same era.
MiniGPT-V2: Large Language Model as a Unified Interface for Vision-Language Multi-task Learning
Chen, J.; Li, D. Z. X. S. X.; Zhang, Z. L. P.; Xiong, R. K. V. C. Y.; and Elhoseiny, M. 2023a
Cited in the paper.
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Chen, L.; Li, J.; wen Dong, X.; Zhang, P.; He, C.; Wang, J.; Zhao, F.; and Lin, D. 2023b
Cited in the paper.
VisualWebInstruct
TIGER-Lab. 2024 · 2024
Closest in time.
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
Tong, Y.; Zhang, X.; Wang, R.; Wu, R.; and He, J. 2024 · 2024
Closest in time.
Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models
Wu, T.; Tao, C.; Wang, J.; Zhao, Z.; and Wong, N. 2024 · 2024
Closest in time.
Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; et al. 2024 · 2024
Closest in time.
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Ye, Q.; Xu, H.; Ye, J.; Yan, M.; Hu, A.; Liu, H.; Qian, Q.; Zhang, J.; and Huang, F. 2024 · 2024
Closest in time.
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Zhang, R.; Jiang, D.; Zhang, Y.; Lin, H.; Guo, Z.; Qiu, P.; Zhou, A.; Lu, P.; Chang, K.-W.; Gao, P.; et al. 2024 · 2024
Closest in time.