Fetching the paper…
Reading the bibliography…
This paper investigates the role of explicit thinking process in rule-based reinforcement fine-tuning (RFT) for MLLMs.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
L. Fei-Fei, R. Fergus, and P. Perona · 2004
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba · 2010
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Cats and dogs
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
J. Krause, M. Stark, J. Deng, and L. Fei-Fei · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
S. Maji, E. Rahtu, J. Kannala, M. Blaschko, and A. Vedaldi · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
L. Bossard, M. Guillaumin, and L. Van Gool · 2014
Earlier work this paper cites.
Describing textures in the wild
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
P. Helber, B. Bischke, A. Dengel, and D. Borth · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
M. Tan and Q. Le · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. Hoi · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Conditional prompt learning for vision-language models
K. Zhou, J. Yang, C. C. Loy, and Z. Liu · 2022
Cited alongside, same era.
Learning to prompt for vision-language models
K. Zhou, J. Yang, C. C. Loy, and Z. Liu · 2022
Cited alongside, same era.
B. Joshi, Z. Liu, S. Ramnath, A. Chan, Z. Tong, S. Nie, Q. Wang, Y. Choi, and X. Ren · 2023
Cited alongside, same era.
Maple: Multi-modal prompt learning
M. U. Khattak, H. Rasheed, M. Maaz, S. Khan, and F. S. Khan · 2023
Cited alongside, same era.
Class balanced adaptive pseudo labeling for federated semi-supervised learning
M. Li, Q. Li, and Y. Wang · 2023
Cited alongside, same era.
Let’s verify step by step
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning
Z. Sprague, F. Yin, J. D. Rodriguez, D. Jiang, M. Wadhwa, P. Singhal, X. Zhao, X. Ye, K. Mahowald, and G. Durrett · 2024
Later among the works it cites.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
P. Tong, E. Brown, P. Wu, S. Woo, A. J. V. IYER, S. C. Akula, S. Yang, J. Yang, M. Middepogu, Z. Wang, et al · 2024
Later among the works it cites.
Measuring multimodal mathematical reasoning with math-vision dataset
K. Wang, J. Pan, W. Shi, Z. Lu, H. Ren, A. Zhou, M. Zhan, and H. Li · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, Y. Fan, K. Dang, M. Du, X. Ren, R. Men, D. Liu, C. Zhou, J. Zhou, and J. Lin · 2024
Later among the works it cites.
Why are visually-grounded language models bad at image classification?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2023
Cited alongside, same era.
Attention is all you need, 2023
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2023
Cited alongside, same era.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations
P. Wang, L. Li, Z. Shao, R. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui · 2023
Cited alongside, same era.
Do not think that much for 2+ 3=? on the overthinking of o1-like llms
X. Chen, J. Xu, T. Liang, Z. He, J. Pang, D. Yu, L. Song, Q. Liu, M. Zhou, Z. Zhang, et al · 2024
Cited alongside, same era.
Z. Chen, W. Wang, Y. Cao, Y. Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu, et al · 2024
Cited alongside, same era.
Y. Zhang, A. Unell, X. Wang, D. Ghosh, Y. Su, L. Schmidt, and S. Yeung-Levy · 2024
Later among the works it cites.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin · 2025
Closest in time.
R1-v: Reinforcing super generalization ability in vision-language models with less than $3
L. Chen, L. Li, H. Zhao, Y. Song, and Vinci · 2025
Closest in time.
Sft memorizes, rl generalizes: A comparative study of foundation model post-training
T. Chu, Y. Zhai, J. Yang, S. Tong, S. Xie, D. Schuurmans, Q. V. Le, S. Levine, and Y. Ma · 2025
Closest in time.
The danger of overthinking: Examining the reasoning-action dilemma in agentic tasks
A. Cuadron, D. Li, W. Ma, X. Wang, Y. Wang, S. Zhuang, S. Liu, L. G. Schroeder, T. Xia, H. Mao, et al · 2025
Closest in time.
Gemini 2.0 flash thinking
DeepMind · 2025
Closest in time.
rstar-math: Small llms can master math reasoning with self-evolved deep thinking
X. Guan, L. L. Zhang, Y. Liu, N. Shang, Y. Sun, Y. Zhu, F. Yang, and M. Yang · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Reinforce++: A simple and efficient approach for aligning large language models
J. Hu · 2025
Closest in time.
Vision-r1: Incentivizing reasoning capability in multimodal large language models
W. Huang, B. Jia, Z. Zhai, S. Cao, Z. Ye, F. Zhao, Z. Xu, Y. Hu, and S. Lin · 2025
Closest in time.
D. Jiang, R. Zhang, Z. Guo, Y. Li, Y. Qi, X. Chen, L. Wang, J. Jin, C. Guo, S. Yan, et al · 2025
Closest in time.
C3ot: Generating shorter chain-of-thought without compromising effectiveness
Y. Kang, X. Sun, L. Chen, and W. Zou · 2025
Closest in time.
Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models
Y. Lai, J. Zhong, M. Li, S. Zhao, and X. Yang · 2025
Closest in time.
Visual-rft: Visual reinforcement fine-tuning
Z. Liu, Z. Sun, Y. Zang, X. Dong, Y. Cao, H. Duan, D. Lin, and J. Wang · 2025
Closest in time.
Mm-eureka: Exploring visual aha moment with rule-based large-scale reinforcement learning
F. Meng, L. Du, Z. Liu, Z. Zhou, Q. Lu, D. Fu, B. Shi, W. Wang, J. He, K. Zhang, et al · 2025
Closest in time.
Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl
Y. Peng, G. Zhang, M. Zhang, Z. You, J. Liu, Q. Zhu, K. Yang, X. Xu, X. Geng, and X. Yang · 2025
Closest in time.
Vlm-r1: A stable and generalizable r1-style large vision-language model
H. Shen, P. Liu, J. Li, C. Fang, Y. Ma, J. Liao, Q. Shen, Z. Zhang, K. Zhao, Q. Zhang, et al · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models
Y. Sui, Y.-N. Chuang, G. Wang, J. Zhang, T. Zhang, J. Yuan, H. Liu, A. Wen, H. Chen, X. Hu, et al · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al · 2025
Closest in time.
When more is less: Understanding chain-of-thought length in llms
Y. Wu, Y. Wang, Z. Ye, T. Du, S. Jegelka, and Y. Wang · 2025
Closest in time.
Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning
T. Xie, Z. Gao, Q. Ren, H. Luo, Y. Hong, B. Dai, J. Zhou, K. Qiu, Z. Wu, and C. Luo · 2025
Closest in time.
Softcot: Soft chain-of-thought for efficient reasoning with llms
Y. Xu, X. Guo, Z. Zeng, and C. Miao · 2025
Closest in time.
Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?
Y. Yue, Z. Chen, R. Lu, A. Zhao, Z. Wang, S. Song, and G. Huang · 2025
Closest in time.
R1-zero’s" aha moment" in visual reasoning on a 2b non-sft model
H. Zhou, X. Li, R. Wang, M. Cheng, T. Zhou, and C.-J. Hsieh · 2025
Closest in time.