Fetching the paper…
Reading the bibliography…
Recent advancements in large language models (LLMs) have demonstrated substantial progress in reasoning capabilities, such as DeepSeek-R1, which leverages rule-based reinforcement learning to enhance logical reasoning significantly.
Proximal policy optimization algorithms
Schulman, J., F. Wolski, P. Dhariwal, et al · 2017
Earlier work this paper cites.
Competence-based curriculum learning for neural machine translation
Platanios, E. A., O. Stretcu, G. Neubig, et al · 2019
Earlier work this paper cites.
A survey on curriculum learning
Wang, X., Y. Chen, W. Zhu · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., J. Donahue, P. Luc, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., J. Wu, X. Jiang, et al · 2022
Earlier work this paper cites.
Let’s verify step by step
Lightman, H., V. Kosaraju, Y. Burda, et al · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., D. Yu, J. Zhao, et al · 2023
Earlier work this paper cites.
Visual instruction tuning
Liu, H., C. Li, Q. Wu, et al · 2023
Earlier work this paper cites.
Yang, A., B. Yang, B. Hui, et al · 2024
Earlier work this paper cites.
Jaech, A., A. Kalai, A. Lerer, et al · 2024
Earlier work this paper cites.
Step-dpo: Step-wise preference optimization for long-chain reasoning of llms
Lai, X., Z. Tian, Y. Chen, et al · 2024
Earlier work this paper cites.
Atomthink: A slow thinking framework for multimodal mathematical reasoning
Xiang, K., Z. Liu, Z. Jiang, et al · 2024
Earlier work this paper cites.
Improve vision language model chain-of-thought reasoning
Zhang, R., B. Zhang, Y. Li, et al · 2024
Cited alongside, same era.
Gme: Improving universal multimodal retrieval by multimodal llms, 2024
Zhang, X., Y. Zhang, W. Xie, et al · 2024
Cited alongside, same era.
Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?
Zhang, R., D. Jiang, Y. Zhang, et al · 2024
Cited alongside, same era.
Measuring multimodal mathematical reasoning with math-vision dataset
Wang, K., J. Pan, W. Shi, et al · 2024
Cited alongside, same era.
Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
He, C., R. Luo, Y. Bai, et al · 2024
Cited alongside, same era.
Vision-r1: Incentivizing reasoning capability in multimodal large language models
Huang, W., B. Jia, Z. Zhai, et al · 2025
Closest in time.
R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization
Yang, Y., X. He, H. Pan, et al · 2025
Closest in time.
X-reasoner: Towards generalizable reasoning across modalities and domains
Liu, Q., S. Zhang, G. Qin, et al · 2025
Closest in time.
Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl
Peng, Y., G. Zhang, M. Zhang, et al · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hurst, A., A. Lerer, A. P. Goucher, et al · 2024
Cited alongside, same era.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Wang, P., S. Bai, S. Tan, et al · 2024
Cited alongside, same era.
Yao, H., J. Huang, W. Wu, et al · 2024
Cited alongside, same era.
Chen, Z., W. Wang, Y. Cao, et al · 2024
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., D. Yang, H. Zhang, et al · 2025
Cited alongside, same era.
Comt: A novel benchmark for chain of multi-modal thought on large vision-language models
Cheng, Z., Q. Chen, J. Zhang, et al · 2025
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., H. Bansal, T. Xia, et al
Cited in the paper.
Lu, Y., J. Yuan, Z. Li, et al · 2025
Closest in time.
Bai, S., K. Chen, X. Liu, et al · 2025
Closest in time.
Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl, 2025
Luo, M., S. Tan, J. Wong, et al · 2025
Closest in time.
Vl-rethinker: Incentivizing self-reflection of vision-language models with reinforcement learning
Wang, H., C. Qu, Z. Huang, et al · 2025
Closest in time.
Light-r1: Curriculum sft, dpo and rl for long cot from scratch and beyond
Wen, L., Y. Cai, F. Xiao, et al · 2025
Closest in time.
Mm-eureka: Exploring the frontiers of multimodal reasoning with rule-based reinforcement learning
Meng, F., L. Du, Z. Liu, et al · 2025
Closest in time.
Fast-slow thinking for large vision-language model reasoning
Xiao, W., L. Gan, W. Dai, et al · 2025
Closest in time.