Fetching the paper…
Reading the bibliography…
Large Multimodal Models (LMMs) demonstrate significant cross-modal reasoning capabilities.
Generative adversarial nets
Goodfellow, I. J., J. Pouget-Abadie, M. Mirza, et al · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., A. Agrawal, J. Lu, et al · 2015
Earlier work this paper cites.
Solving geometry problems: Combining text and diagram interpretation
Seo, M., H. Hajishirzi, A. Farhadi, et al · 2015
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A., M. Salvato, E. Kolve, et al · 2016
Earlier work this paper cites.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Kembhavi, A., M. Seo, D. Schwenk, et al · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., T. Khot, D. Summers-Stay, et al · 2017
Earlier work this paper cites.
Figureqa: An annotated figure dataset for visual reasoning
Kahou, S. E., V. Michalski, A. Atkinson, et al · 2018
Earlier work this paper cites.
Dvqa: Understanding data visualizations via question answering
Kafle, K., B. Price, S. Cohen, et al · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Q. Li, A. J. Stangl, et al · 2018
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
Lau, J. J., S. Gayen, A. Ben Abacha, et al · 2018
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., V. Natarajan, M. Shah, et al · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N., I. Gurevych · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., M.-W. Chang, K. Lee, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., B. Mann, N. Ryder, et al · 2020
Earlier work this paper cites.
Plotqa: Reasoning over scientific plots
Methani, N., P. Ganguly, M. M. Khapra, et al · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., J. W. Kim, C. Hallacy, et al · 2021
Earlier work this paper cites.
Finqa: A dataset of numerical reasoning over financial data
Chen, Z., W. Chen, C. Smiley, et al · 2021
Earlier work this paper cites.
Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning
Lu, P., L. Qiu, J. Chen, et al · 2021
Earlier work this paper cites.
Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Lu, P., R. Gong, S. Jiang, et al · 2021
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., X. Wang, D. Schuurmans, et al · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., J. Donahue, P. Luc, et al · 2022
Earlier work this paper cites.
When flue meets flang: Benchmarks and large pre-trained language model for financial domain
Shah, R. S., K. Chawla, D. Eidnani, et al · 2022
Earlier work this paper cites.
Infographicvqa
Mathew, M., V. Bagal, R. Tito, et al · 2022
Earlier work this paper cites.
Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning
Lu, P., L. Qiu, K.-W. Chang, et al · 2022
Earlier work this paper cites.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Masry, A., D. X. Long, J. Q. Tan, et al · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., S. Mishra, T. Xia, et al · 2022
Earlier work this paper cites.
An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Cao, J., J. Xiao · 2022
Earlier work this paper cites.
Mapqa: A dataset for question answering on choropleth maps
Chang, S., D. Palzer, J. Li, et al · 2022
Earlier work this paper cites.
Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression
Chen, J., T. Li, J. Qin, et al · 2022
Earlier work this paper cites.
Clevr-math: A dataset for compositional language, visual and mathematical reasoning
Lindström, A. D., S. S. Abraham · 2022
Earlier work this paper cites.
Finvis-gpt: A multimodal large language model for financial chart analysis
Wang, Z., Y. Li, J. Wu, et al · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., D. Li, S. Savarese, et al · 2023
Cited alongside, same era.
Visual instruction tuning
Liu, H., C. Li, Q. Wu, et al · 2023
Cited alongside, same era.
Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models
Zheng, G., B. Yang, J. Tang, et al · 2023
Cited alongside, same era.
Chameleon: Plug-and-play compositional reasoning with large language models
Lu, P., B. Peng, H. Cheng, et al · 2023
Cited alongside, same era.
Mm-react: Prompting chatgpt for multimodal reasoning and action
Grattafiori, A., A. Dubey, A. Jauhri, et al · 2024
Later among the works it cites.
Investorbench: A benchmark for financial decision-making tasks with llm-based agent
Li, H., Y. Cao, Y. Yu, et al · 2024
Later among the works it cites.
Finben: A holistic financial benchmark for large language models
Xie, Q., W. Han, Z. Chen, et al · 2024
Later among the works it cites.
Xu, L., L. Zhu, Y. Wu, et al · 2024
Later among the works it cites.
Mme-finance: A multimodal finance benchmark for expert-level understanding and reasoning
Gan, Z., Y. Lu, D. Zhang, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang, Z., L. Li, J. Wang, et al · 2023
Cited alongside, same era.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Qin, Y., S. Liang, Y. Ye, et al · 2023
Cited alongside, same era.
Bai, J., S. Bai, Y. Chu, et al · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G., R. Anil, S. Borgeaud, et al · 2023
Cited alongside, same era.
Bloomberggpt: A large language model for finance
Wu, S., O. Irsoy, S. Lu, et al · 2023
Cited alongside, same era.
Fingpt: Instruction tuning benchmark for open-source large language models in financial datasets
Wang, N., H. Yang, C. D. Wang · 2023
Cited alongside, same era.
Xuanyuan 2.0: A large chinese financial chat model with hundreds of billions parameters
Zhang, X., Q. Yang · 2023
Cited alongside, same era.
Later among the works it cites.
Yizhao-findataset, 2024
CMB AILab · 2024
Later among the works it cites.
Liu, A., B. Feng, B. Xue, et al · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., P. Wang, Q. Zhu, et al · 2024
Later among the works it cites.
Measuring multimodal mathematical reasoning with math-vision dataset
Wang, K., J. Pan, W. Shi, et al · 2024
Later among the works it cites.
Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?
Zhang, R., D. Jiang, Y. Zhang, et al · 2024
Later among the works it cites.
Are we on the right way for evaluating large vision-language models?
Chen, L., J. Li, X. Dong, et al · 2024
Later among the works it cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., H. Bansal, T. Xia, et al · 2024
Later among the works it cites.
Gpqa: A graduate-level google-proof q&a benchmark
Rein, D., B. L. Hou, A. C. Stickland, et al · 2024
Later among the works it cites.
Chen, Z., W. Wang, Y. Cao, et al · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone
Yao, Y., T. Yu, A. Zhang, et al · 2024
Later among the works it cites.
Hybridflow: A flexible and efficient rlhf framework
Sheng, G., C. Zhang, Z. Ye, et al · 2024
Later among the works it cites.
Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl
Peng, Y., G. Zhang, M. Zhang, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., D. Yang, H. Zhang, et al · 2025
Closest in time.
Srpo: A cross-domain implementation of large-scale reinforcement learning on llm
Zhang, X., J. Wang, Z. Cheng, et al · 2025
Closest in time.
Skywork r1v2: Multimodal hybrid reinforcement learning for reasoning
Wei, Y., Y. Peng, X. Wang, et al · 2025
Closest in time.
Visualprm: An effective process reward model for multimodal reasoning
Wang, W., Z. Gao, L. Chen, et al · 2025
Closest in time.
Echoink-r1: Exploring audio-visual reasoning in multimodal llms via reinforcement learning
Xing, Z., X. Hu, C.-W. Fu, et al · 2025
Closest in time.
R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization
Yang, Y., X. He, H. Pan, et al · 2025
Closest in time.
Fin-r1: A large language model for financial reasoning through reinforcement learning
Liu, Z., X. Guo, F. Lou, et al · 2025
Closest in time.
FAMMA: A benchmark for financial multilingual multimodal question answering, 2025
Xue, S., T. Chen, F. Zhou, et al · 2025
Closest in time.
Bai, S., K. Chen, X. Liu, et al · 2025
Closest in time.
Towards thinking-optimal scaling of test-time compute for llm reasoning
Yang, W., S. Ma, Y. Lin, et al · 2025
Closest in time.
Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl
Luo, M., S. Tan, J. Wong, et al · 2025
Closest in time.
Famma: A benchmark for financial domain multilingual multimodal question answering
Xue, S., X. Li, F. Zhou, et al · 2025
Closest in time.
Easyr1: An efficient, scalable, multi-modality rl training framework
Yaowei, Z., L. Junting, W. Shenzhi, et al · 2025
Closest in time.