Fetching the paper…
Reading the bibliography…
Reinforcement learning with verifiable rewards (RLVR) has shown promise in enhancing the reasoning capabilities of large language models by learning directly from outcome-based rewards.
Wang, R., Lehman, J., Clune, J., and Stanley, K. O · 1901
Earlier work this paper cites.
Demski, A. and Garrabrant, S · 1902
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems
Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S · 1906
Earlier work this paper cites.
Regularizing deep multi-task networks using orthogonal gradients
Suteu, M. and Guo, Y · 1912
Earlier work this paper cites.
Elements of Software Science (Operating and programming systems series)
Halstead, M. H · 1977
Earlier work this paper cites.
Verification, the key to ai
Sutton, R. S · 2001
Earlier work this paper cites.
Exploring the predictable
Schmidhuber, J · 2003
Earlier work this paper cites.
Schmidhuber, J · 2011
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Understanding computation - from simple machines to impossible programs
Stuart, T · 2015
Earlier work this paper cites.
Cyclomatic complexity
Ebert, C., Cain, J., Antoniol, G., Counsell, S., and Laplante, P · 2016
Earlier work this paper cites.
Intrinsic motivation, curiosity, and learning: Theory and applications in educational technologies
Oudeyer, P.-Y., Gottlieb, J., and Lopes, M · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T. P., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T. P., Simonyan, K., and Hassabis, D · 2017
Earlier work this paper cites.
Approval-directed bootstrapping
Christiano, P · 2018
Earlier work this paper cites.
Automatic goal generation for reinforcement learning agents
Florensa, C., Held, D., Geng, X., and Abbeel, P · 2018
Earlier work this paper cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, S., Lin, Z., Kostrikov, I., Synnaeve, G., Szlam, A., and Fergus, R · 2018
Earlier work this paper cites.
Capability amplification
Christiano, P · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Emergent complexity and zero-shot transfer via unsupervised environment design
Dennis, M., Jaques, N., Vinitsky, E., Bayen, A. M., Russell, S., Critch, A., and Levine, S · 2020
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. C., and Bengio, Y · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M. I., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C. J., Terry, M., Le, Q. V., and Sutton, C · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
URLB: unsupervised reinforcement learning benchmark
Laskin, M., Yarats, D., Liu, H., Lee, K., Zhan, A., Lu, K., Cang, C., Pinto, L., and Abbeel, P · 2021
Earlier work this paper cites.
Asymmetric self-play for automatic goal discovery in robotic manipulation
OpenAI, Plappert, M., Sampedro, R., Xu, T., Akkaya, I., Kosaraju, V., Welinder, P., D’Sa, R., Petron, A., de Oliveira Pinto, H. P., Paino, A., Noh, H., Weng, L., Yuan, Q., Chu, C., and Zaremba, W · 2021
Earlier work this paper cites.
A survey on multi-task learning
Zhang, Y. and Yang, Q · 2021
Earlier work this paper cites.
Exploration in deep reinforcement learning: A survey
Ladosz, P., Weng, L., Kim, M., and Oh, H · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V. V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y., Neyshabur, B., Gur-Ari, G., and Misra, V · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., Mu, J., and Goodman, N · 2022
Earlier work this paper cites.
A mixture of surprises for unsupervised reinforcement learning
Zhao, A., Lin, M. G., Li, Y., Liu, Y., and Huang, G · 2022
Cited alongside, same era.
Language models can teach themselves to program better
Haluptzok, P., Bowers, M., and Kalai, A. T · 2023
Cited alongside, same era.
Introducing superalignment
Leike, J. and Sutskever, I · 2023
Cited alongside, same era.
TACO: topics in algorithmic code generation dataset
Li, R., Fu, J., Zhang, B., Huang, T., Sun, Z., Lyu, C., Liu, G., Jin, Z., and Li, G · 2023
Cited alongside, same era.
Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
Liu, J., Xia, C. S., Wang, Y., and Zhang, L · 2023
Cited alongside, same era.
Train once, get a family: State-adaptive balances for offline-to-online reinforcement learning
Stateflow: Enhancing LLM task-solving through state-driven workflows
Wu, Y., Yue, T., Zhang, S., Wang, C., and Wu, Q · 2024
Later among the works it cites.
Evolving alignment via asymmetric self-play
Ye, Z., Agarwal, R., Liu, T., Joshi, R., Velury, S., Le, Q. V., Tan, Q., and Liu, Y · 2024
Later among the works it cites.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Later among the works it cites.
Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution
Yue, Y., Wang, Y., Kang, B., Han, Y., Wang, S., Song, S., Feng, J., and Huang, G · 2024
Later among the works it cites.
Expel: LLM agents are experiential learners
Zhao, A., Huang, D., Xu, Q., Lin, M., Liu, Y., and Huang, G · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, S., Yang, Q., Gao, J., Lin, M. G., Chen, H., Wu, L., Jia, N., Song, S., and Huang, G · 2023
Cited alongside, same era.
Autogen: Enabling next-gen LLM applications via multi-agent conversation framework
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Zhang, S., Zhu, E., Li, B., Jiang, L., Zhang, X., and Wang, C · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y · 2023
Cited alongside, same era.
Understanding, predicting and better resolving q-value divergence in offline-rl
Yue, Y., Lu, R., Kang, B., Song, S., and Huang, G · 2023
Cited alongside, same era.
RT-2: vision-language-action models transfer web knowledge to robotic control
Zitkovich, B., Yu, T., Xu, S., Xu, P., Xiao, T., Xia, F., Wu, J., Wohlhart, P., Welker, S., Wahid, A., Vuong, Q., Vanhoucke, V., Tran, H. T., Soricut, R., Singh, A., Singh, J., Sermanet, P., Sanketi, P. R., Salazar, G., Ryoo, M. S., Reymann, K., Rao, K., Pertsch, K., Mordatch, I., Michalewski, H., Lu, Y., Levine, S., Lee, L., Lee, T. E., Leal, I., Kuang, Y., Kalashnikov, D., Julian, R., Joshi, N. J., Irpan, A., Ichter, B., Hsu, J., Herzog, A., Hausman, K., Gopalakrishnan, K., Fu, C., Florence, P., Finn, C., Dubey, K. A., Driess, D., Ding, T., Choromanski, K. M., Chen, X., Chebotar, Y., Carbajal, J., Brown, N., Brohan, A., Arenas, M. G., and Han, K · 2023
Cited alongside, same era.
To code, or not to code? exploring impact of code in pre-training
Aryabumi, V., Su, Y., Ma, R., Morisot, A., Zhang, I., Locatelli, A., Fadaee, M., Üstün, A., and Hooker, S · 2024
Cited alongside, same era.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., Sutskever, I., and Wu, J · 2024
Cited alongside, same era.
Later among the works it cites.
Radon: Python tool for code metrics
Canal, M · 2025
Closest in time.
Spc: Evolving self-play critic via adversarial games for llm reasoning, 2025
Chen, J., Zhang, B., Ma, R., Wang, P., Liang, X., Tu, Z., Li, X., and Wong, K.-Y. K · 2025
Closest in time.
Process reinforcement through implicit rewards
Cui, G., Yuan, L., Wang, Z., Wang, H., Li, W., He, B., Fan, Y., Yu, T., Xu, Q., Chen, W., Yuan, J., Chen, H., Zhang, K., Lv, X., Wang, S., Yao, Y., Han, X., Peng, H., Cheng, Y., Liu, Z., Sun, M., Zhou, B., and Ding, N · 2025
Closest in time.
Gaven, L., Carta, T., Romac, C., Colas, C., Lamprier, S., Sigaud, O., and Oudeyer, P · 2025
Closest in time.
Deepseek-r1 incentivizes reasoning in llms through reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, R., Ma, S., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., Xue, B., Wang, B., Wu, B., Feng, B., Lu, C., Zhao, C., Deng, C., Ruan, C., Dai, D., Chen, D., Ji, D., Li, E., Lin, F., Dai, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Xu, H., Ding, H., Gao, H., Qu, H., Li, H., Guo, J., Li, J., Chen, J., Yuan, J., Tu, J., Qiu, J., Li, J., Cai, J. L., Ni, J., Liang, J., Chen, J., Dong, K., Hu, K., You, K., Gao, K., Guan, K., Huang, K., Yu, K., Wang, L., Zhang, L., Zhao, L., Wang, L., Zhang, L., Xu, L., Xia, L., Zhang, M., Zhang, M., Tang, M., Zhou, M., Li, M., Wang, M., Li, M., Tian, N., Huang, P., Zhang, P., Wang, Q., Chen, Q., Du, Q., Ge, R., Zhang, R., Pan, R., Wang, R., Chen, R. J., Jin, R. L., Chen, R., Lu, S., Zhou, S., Chen, S., Ye, S., Wang, S., Yu, S., Zhou, S., Pan, S., Li, S. S., Zhou, S., Wu, S., Yun, T., Pei, T., Sun, T., Wang, T., Zeng, W., Liu, W., Liang, W., Gao, W., Yu, W., Zhang, W., Xiao, W. L., An, W., Liu, X., Wang, X., Chen, X., Nie, X., Cheng, X., Liu, X., Xie, X., Liu, X., Yang, X., Li, X., Su, X., Lin, X., Li, X. Q., Jin, X., Shen, X., Chen, X., Sun, X., Wang, X., Song, X., Zhou, X., Wang, X., Shan, X., Li, Y. K., Wang, Y. Q., Wei, Y. X., Zhang, Y., Xu, Y., Li, Y., Zhao, Y., Sun, Y., Wang, Y., Yu, Y., Zhang, Y., Shi, Y., Xiong, Y., He, Y., Piao, Y., Wang, Y., Tan, Y., Ma, Y., Liu, Y., Guo, Y., Ou, Y., Wang, Y., Gong, Y., Zou, Y., He, Y., Xiong, Y., Luo, Y., You, Y., Liu, Y., Zhou, Y., Zhu, Y. X., Huang, Y., Li, Y., Zheng, Y., Zhu, Y., Ma, Y., Tang, Y., Zha, Y., Yan, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Xu, Z., Xie, Z., Zhang, Z., Hao, Z., Ma, Z., Yan, Z., Wu, Z., Gu, Z., Zhu, Z., Liu, Z., Li, Z., Xie, Z., Song, Z., Pan, Z., Huang, Z., Xu, Z., Zhang, Z., and Zhang, Z · 2025
Closest in time.
REINFORCE++: A simple and efficient approach for aligning large language models
Hu, J · 2025
Closest in time.
Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model
Hu, J., Zhang, Y., Han, Q., Jiang, D., Zhang, X., and Shum, H · 2025
Closest in time.
Codei/o: Condensing reasoning patterns via code input-output prediction
Li, J., Guo, D., Yang, D., Xu, R., Wu, Y., and He, J · 2025
Closest in time.
Code-r1: Reproducing r1 for code with reliable rewards
Liu, J. and Zhang, L · 2025
Closest in time.
Complexipy: An extremely fast python library to calculate the cognitive complexity of python files, written in rust, 2025
Lopez, R. H. Q · 2025
Closest in time.
There are no new ideas in ai… only new datasets
Morris, J · 2025
Closest in time.
Openai o3-mini, January 2025a
OpenAI · 2025
Closest in time.
Introducing openai o3 and o4-mini, April 2025b
OpenAI · 2025
Closest in time.
Ren, Z. Z., Shao, Z., Song, J., Xin, H., Wang, H., Zhao, W., Zhang, L., Fu, Z., Zhu, Q., Yang, D., Wu, Z. F., Gou, Z., Ma, S., Tang, H., Liu, Y., Gao, W., Guo, D., and Ruan, C · 2025
Closest in time.
Hybridflow: A flexible and efficient RLHF framework
Sheng, G., Zhang, C., Ye, Z., Wu, X., Zhang, W., Zhang, R., Peng, Y., Lin, H., and Wu, C · 2025
Closest in time.
The era of experience
Silver, D. and Sutton, R. S · 2025
Closest in time.
Kimi k1.5: Scaling reinforcement learning with llms
Team, K., Du, A., Gao, B., Xing, B., Jiang, C., Chen, C., Li, C., Xiao, C., Du, C., Liao, C., Tang, C., Wang, C., Zhang, D., Yuan, E., Lu, E., Tang, F., Sung, F., Wei, G., Lai, G., Guo, H., Zhu, H., Ding, H., Hu, H., Yang, H., Zhang, H., Yao, H., Zhao, H., Lu, H., Li, H., Yu, H., Gao, H., Zheng, H., Yuan, H., Chen, J., Guo, J., Su, J., Wang, J., Zhao, J., Zhang, J., Liu, J., Yan, J., Wu, J., Shi, L., Ye, L., Yu, L., Dong, M., Zhang, N., Ma, N., Pan, Q., Gong, Q., Liu, S., Ma, S., Wei, S., Cao, S., Huang, S., Jiang, T., Gao, W., Xiong, W., He, W., Huang, W., Wu, W., He, W., Wei, X., Jia, X., Wu, X., Xu, X., Zu, X., Zhou, X., Pan, X., Charles, Y., Li, Y., Hu, Y., Liu, Y., Chen, Y., Wang, Y., Liu, Y., Qin, Y., Liu, Y., Yang, Y., Bao, Y., Du, Y., Wu, Y., Wang, Y., Zhou, Z., Wang, Z., Li, Z., Zhu, Z., Zhang, Z., Wang, Z., Yang, Z., Huang, Z., Huang, Z., Xu, Z., and Yang, Z · 2025
Closest in time.
Model surgery: Modulating LLM‘s behavior via simple parameter editing
Wang, H., Yue, Y., Lu, R., Shi, J., Zhao, A., Wang, S., Song, S., and Huang, G · 2025
Closest in time.
Logic-rl: Unleashing LLM reasoning with rule-based reinforcement learning
Xie, T., Gao, Z., Ren, Q., Luo, H., Hong, Y., Dai, B., Zhou, J., Qiu, K., Wu, Z., and Luo, C · 2025
Closest in time.
Genius: A generalizable and purely unsupervised self-training framework for advanced reasoning, 2025
Xu, F., Yan, H., Ma, C., Zhao, H., Sun, Q., Cheng, K., He, J., Liu, J., and Wu, Z · 2025
Closest in time.
DAPO: an open-source LLM reinforcement learning system at scale
Yu, Q., Zhang, Z., Zhu, R., Yuan, Y., Zuo, X., Yue, Y., Fan, T., Liu, G., Liu, L., Liu, X., Lin, H., Lin, Z., Ma, B., Sheng, G., Tong, Y., Zhang, C., Zhang, M., Zhang, W., Zhu, H., Zhu, J., Chen, J., Chen, J., Wang, C., Yu, H., Dai, W., Song, Y., Wei, X., Zhou, H., Liu, J., Ma, W., Zhang, Y., Yan, L., Qiao, M., Wu, Y., and Wang, M · 2025
Closest in time.
Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks
Yuan, Y., Yu, Q., Zuo, X., Zhu, R., Xu, W., Chen, J., Wang, C., Fan, T., Du, Z., Wei, X., et al · 2025
Closest in time.
Yue, Y., Chen, Z., Lu, R., Zhao, A., Wang, Z., Yue, Y., Song, S., and Huang, G · 2025
Closest in time.
Diver-ct: Diversity-enhanced red teaming large language model assistants with relaxing constraints
Zhao, A., Xu, Q., Lin, M., Wang, S., Liu, Y., Zheng, Z., and Huang, G · 2025
Closest in time.
Self-referencing agents for unsupervised reinforcement learning
Zhao, A., Zhu, E., Lu, R., Lin, M., Liu, Y., and Huang, G · 2025
Closest in time.
Ttrl: Test-time reinforcement learning, 2025
Zuo, Y., Zhang, K., Qu, S., Sheng, L., Zhu, X., Qi, B., Sun, Y., Cui, G., Ding, N., and Zhou, B · 2025
Closest in time.