Fetching the paper…
Reading the bibliography…
Recent Large Language Models (LLMs) have significantly advanced natural language processing and automated decision-making.
Maximal flow through a network
Ford, L. R., D. R. Fulkerson · 1956
Earlier work this paper cites.
An integrative theory of prefrontal cortex function
Miller, E. K., J. D. Cohen · 2001
Earlier work this paper cites.
Perplexity—a measure of the difficulty of speech recognition tasks
Jelinek, F., R. L. Mercer, L. R. Bahl, et al · 2005
Earlier work this paper cites.
Sequential sampling models in cognitive neuroscience
Forstmann, B. U., R. Ratcliff, E.-J. Wagenmakers · 2016
Earlier work this paper cites.
Dual-process theories
Evans, J. S. B · 2018
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., C. Burns, S. Kadavath, et al · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
Hendrycks, D., C. Burns, S. Basart, et al · 2021
Earlier work this paper cites.
Students’ use of formalisations for improved logical reasoning
Bronkhorst, H., G. Roorda, C. Suhre, et al · 2022
Earlier work this paper cites.
Detecting language model attacks with perplexity, 2023
Alon, G., M. Kamfonas · 2023
Earlier work this paper cites.
Eureka: Human-level reward design via coding large language models
Ma, Y. J., W. Liang, G. Wang, et al · 2023
Earlier work this paper cites.
Gpqa: A graduate-level google-proof q&a benchmark, 2023
Rein, D., B. L. Hou, A. C. Stickland, et al · 2023
Earlier work this paper cites.
Agieval: A human-centric benchmark for evaluating foundation models, 2023
Zhong, W., R. Cui, Y. Guo, et al · 2023
Earlier work this paper cites.
OpenAI o1, 2024
OpenAI · 2024
Earlier work this paper cites.
Can language models learn to skip steps?, 2024
Liu, T., Q. Guo, X. Hu, et al · 2024
Earlier work this paper cites.
Jaech, A., A. Kalai, A. Lerer, et al · 2024
Earlier work this paper cites.
Vineppo: Unlocking rl potential for llm reasoning through refined credit assignment
Kazemnejad, A., M. Aghajohari, E. Portelance, et al · 2024
Earlier work this paper cites.
On designing effective rl reward at training time for llm reasoning
Gao, J., S. Xu, W. Ye, et al · 2024
Earlier work this paper cites.
Omni-math: A universal olympiad level mathematic benchmark for large language models, 2024
Gao, B., F. Song, Z. Yang, et al · 2024
Earlier work this paper cites.
Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024
He, C., R. Luo, Y. Bai, et al · 2024
Earlier work this paper cites.
Qwen2.5-math technical report: Toward mathematical expert model via self-improvement, 2024
Yang, A., B. Zhang, B. Hui, et al · 2024
Cited alongside, same era.
Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning systems, 2024
Min, Y., Z. Chen, J. Jiang, et al · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Shao, Z., P. Wang, Q. Zhu, et al · 2024
Cited alongside, same era.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, 2025
DeepSeek-AI, D. Guo, D. Yang, et al · 2025
Cited alongside, same era.
QwQ-32B: Embracing the Power of Reinforcement Learning | Qwen, 2025
QwQ · 2025
Cited alongside, same era.
Dapo: An open-source llm reinforcement learning system at scale
Yu, Q., Z. Zhang, R. Zhu, et al · 2025
Closest in time.
Enhancing llm reasoning with iterative dpo: A comprehensive empirical investigation
Tu, S., J. Lin, X. Tian, et al · 2025
Closest in time.
Cppo: Accelerating the training of group relative policy optimization-based reasoning models
Lin, Z., M. Lin, Y. Xie, et al · 2025
Closest in time.
Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks
Yue, Y., Y. Yuan, Q. Yu, et al · 2025
Closest in time.
One framework to rule them all: Unifying rl-based and rl-free methods in rlhf
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Team, K., A. Du, B. Gao, et al · 2025
Cited alongside, same era.
Adaptive group policy optimization: Towards stable training and token-efficient reasoning
Li, C., N. Liu, K. Yang · 2025
Cited alongside, same era.
Training language models to reason efficiently
Arora, D., A. Zanette · 2025
Cited alongside, same era.
L1: Controlling how long a reasoning model thinks with reinforcement learning
Aggarwal, P., S. Welleck · 2025
Cited alongside, same era.
O1-pruner: Length-harmonizing fine-tuning for o1-like reasoning pruning
Luo, H., L. Shen, H. He, et al · 2025
Cited alongside, same era.
Dast: Difficulty-adaptive slow-thinking for large reasoning models
Shen, Y., J. Zhang, J. Huang, et al · 2025
Cited alongside, same era.
Thinkprune: Pruning long chain-of-thought of llms via reinforcement learning
Hou, B., Y. Zhang, J. Ji, et al · 2025
Cited alongside, same era.
Cai, X · 2025
Closest in time.
Exploring data scaling trends and effects in reinforcement learning from human feedback
Shen, W., G. Liu, Z. Wu, et al · 2025
Closest in time.
Light-r1: Curriculum sft, dpo and rl for long cot from scratch and beyond
Wen, L., Y. Cai, F. Xiao, et al · 2025
Closest in time.
Tapered off-policy reinforce: Stable and efficient reinforcement learning for llms
Roux, N. L., M. G. Bellemare, J. Lebensold, et al · 2025
Closest in time.
Process reinforcement through implicit rewards
Cui, G., L. Yuan, Z. Wang, et al · 2025
Closest in time.
A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility, 2025
Hochlehnert, A., H. Bhatnagar, V. Udandarao, et al · 2025
Closest in time.
s1: Simple test-time scaling, 2025
Muennighoff, N., Z. Yang, W. Shi, et al · 2025
Closest in time.
Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl
Luo, M., S. Tan, J. Wong, et al · 2025
Closest in time.
R1-searcher: Stimulating the search capability of llm from zero via reinforcement learning
Song, H., J. Jiang, Y. Min, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, D. Guo, D. Yang, et al · 2025
Closest in time.
There may not be aha moment in r1-zero-like training — a pilot study
Liu, Z., C. Chen, W. Li, et al · 2025
Closest in time.
Fastcurl: Curriculum reinforcement learning with progressive context extension for efficient training r1-like reasoning models, 2025
Song, M., M. Zheng, Z. Li, et al · 2025
Closest in time.
Reinforcement learning for reasoning in small llms: What works and what doesn’t, 2025
Dang, Q.-A., C. Ngo · 2025
Closest in time.
L1: Controlling how long a reasoning model thinks with reinforcement learning, 2025
Aggarwal, P., S. Welleck · 2025
Closest in time.