Fetching the paper…
Reading the bibliography…
Recent advances in multimodal Reward Models (RMs) have shown significant promise in delivering reward signals to align vision models with human preferences.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Reinforced self-training (rest) for language modeling
C. Gulcehre, T. L. Paine, S. Srinivasan, K. Konyushkova, L. Weerts, A. Sharma, A. Siddhant, A. Ahern, M. Wang, C. Gu, et al · 2023
Earlier work this paper cites.
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Y. Kirstain, A. Polyak, U. Singer, S. Matiana, J. Penna, and O. Levy · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2023
Earlier work this paper cites.
Aligning large multimodal models with factually augmented rlhf
Z. Sun, S. Shen, S. Cao, H. Liu, C. Li, Y. Shen, C. Gan, L.-Y. Gui, Y.-X. Wang, Y. Yang, et al · 2023
Earlier work this paper cites.
Imagereward: Learning and evaluating human preferences for text-to-image generation
J. Xu, X. Liu, Y. Wu, Y. Tong, Q. Li, M. Ding, J. Tang, and Y. Dong · 2023
Earlier work this paper cites.
Finding the subjective truth: Collecting 2 million votes for comprehensive gen-ai model evaluation
D. Christodoulou and M. Kuhlmann-Jørgensen · 2024
Earlier work this paper cites.
S. Han, H. Fan, J. Fu, L. Li, T. Li, J. Cui, Y. Wang, Y. Tai, J. Sun, C. Guo, et al · 2024
Earlier work this paper cites.
Videoscore: Building automatic metrics to simulate fine-grained human feedback for video generation
X. He, D. Jiang, G. Zhang, M. Ku, A. Soni, S. Siu, H. Chen, A. Chandra, Z. Jiang, A. Arulraj, K. Wang, Q. D. Do, Y. Ni, B. Lyu, Y. Narsupalli, R. Fan, Z. Lyu, Y. Lin, and W. Chen · 2024
Earlier work this paper cites.
Qwen2. 5-coder technical report
B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Lu, et al · 2024
Cited alongside, same era.
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al · 2024
Cited alongside, same era.
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al · 2024
Cited alongside, same era.
Genai arena: An open evaluation platform for generative models
D. Jiang, M. Ku, T. Li, Y. Ni, S. Sun, R. Fan, and W. Chen · 2024
Cited alongside, same era.
Llava-critic: Learning to evaluate multimodal models
T. Xiong, X. Wang, D. Guo, Q. Ye, H. Fan, Q. Gu, H. Huang, and C. Li · 2024
Later among the works it cites.
J. Xu, Y. Huang, J. Cheng, Y. Yang, J. Xu, Y. Wang, W. Duan, S. Yang, Q. Jin, S. Li, J. Teng, Z. Yang, W. Zheng, X. Liu, M. Ding, X. Zhang, X. Gu, S. Huang, M. Huang, J. Tang, and Y. Dong · 2024
Later among the works it cites.
Qwen2. 5-math technical report: Toward mathematical expert model via self-improvement
A. Yang, B. Zhang, B. Hui, B. Gao, B. Yu, C. Li, D. Liu, J. Tu, J. Zhou, J. Lin, et al · 2024
Later among the works it cites.
Internlm-math: Open math large language models toward verifiable reasoning
H. Ying, S. Zhang, L. Li, Z. Zhou, Y. Shao, Z. Fei, Y. Ma, J. Hong, K. Liu, Z. Wang, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Jiao, G. Guo, X. Zhang, N. F. Chen, S. Joty, and F. Wei · 2024
Cited alongside, same era.
Videodpo: Omni-preference alignment for video diffusion generation
R. Liu, H. Wu, Z. Ziqiang, C. Wei, Y. He, R. Pi, and Q. Chen · 2024
Cited alongside, same era.
Reft: Reasoning with reinforced fine-tuning
T. Q. Luong, X. Zhang, Z. Jie, P. Sun, X. Jin, and H. Li · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al · 2024
Cited alongside, same era.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Cited alongside, same era.
Lift: Leveraging human feedback for text-to-video model alignment
Y. Wang, Z. Tan, J. Wang, X. Yang, C. Jin, and H. Li · 2024
Cited alongside, same era.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al
Cited in the paper.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al
Cited in the paper.
Y. Zhou, C. Cui, R. Rafailov, C. Finn, and H. Yao · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Q-insight: Understanding image quality via visual reinforcement learning
W. Li, X. Zhang, S. Zhao, Y. Zhang, J. Li, L. Zhang, and J. Zhang · 2025
Closest in time.
Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?
Y. Yue, Z. Chen, R. Lu, A. Zhao, Z. Wang, S. Song, and G. Huang · 2025
Closest in time.
Internlm-xcomposer2.5-reward: A simple yet effective multi-modal reward model
Y. Zang, X. Dong, P. Zhang, Y. Cao, Z. Liu, S. Ding, S. Wu, Y. Ma, H. Duan, W. Zhang, K. Chen, D. Lin, and J. Wang · 2025
Closest in time.