Fetching the paper…
Reading the bibliography…
Chain-of-thought (CoT) reasoning in large language models (LLMs) can be formalized as a latent variable problem, where the model needs to generate intermediate reasoning steps.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E. (1952) · 1952
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J. (1991) · 1991
Earlier work this paper cites.
Safe and effective importance sampling
Owen, A. and Zhou, Y. (2000) · 2000
Earlier work this paper cites.
Slice sampling
Neal, R. M. (2003) · 2003
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C. M. and Nasrabadi, N. M. (2006) · 2006
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence _rate for finite training sets
Roux, N., Schmidt, M., and Bach, F. (2012) · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T. (2013) · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P., Welling, M., et al. (2013) · 2013
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S. (2014) · 2014
Earlier work this paper cites.
Accelerating minibatch stochastic gradient descent using stratified sampling
Zhao, P. and Zhang, T. (2014) · 2014
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search
Anthony, T., Tian, Z., and Barber, D. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Earlier work this paper cites.
Stochastic expectation maximization with variance reduction
Chen, J., Zhu, J., Teh, Y. W., and Zhang, T. (2018) · 2018
Earlier work this paper cites.
Buy 4 reinforce samples, get a baseline for free!
Kool, W., van Hoof, H., and Welling, M. (2019) · 2019
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021) · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. (2022) · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., et al. (2022) · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. (2022) · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., Mu, J., and Goodman, N. (2022) · 2022
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R. (2023) · 2023
Cited alongside, same era.
RAFT: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., SHUM, K., and Zhang, T. (2023) · 2023
Cited alongside, same era.
Llama 3 model card
AI@Meta (2024) · 2024
Later among the works it cites.
Rlhf workflow: From reward modeling to online rlhf
Dong, H., Xiong, W., Pang, B., Wang, H., Zhao, H., Zhou, Y., Jiang, N., Sahoo, D., Xiong, C., and Zhang, T. (2024) · 2024
Later among the works it cites.
He, C., Luo, R., Bai, Y., Hu, S., Thai, Z. L., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., et al. (2024) · 2024
Later among the works it cites.
Unveiling the statistical foundations of chain-of-thought prompting methods
Hu, X., Zhang, F., Chen, S., and Yang, Z. (2024) · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforced self-training (rest) for language modeling
Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al. (2023) · 2023
Cited alongside, same era.
Training chain-of-thought via latent-variable inference
Hoffman, M. D., Phan, D., Dohan, D., Douglas, S., Le, T. A., Parisi, A., Sountsov, P., Sutton, C., Vikram, S., and Saurous, R. A. (2023) · 2023
Cited alongside, same era.
Remax: A simple, effective, and efficient reinforcement learning method for aligning large language models
Li, Z., Xu, T., Zhang, Y., Yu, Y., Sun, R., and Luo, Z.-Q. (2023) · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. (2023) · 2023
Cited alongside, same era.
Beyond human data: Scaling self-training for problem-solving with language models
Singh, A., Co-Reyes, J. D., Agarwal, R., Anand, A., Patil, P., Liu, P. J., Harrison, J., Lee, J., Xu, K., Parisi, A., et al. (2023) · 2023
Cited alongside, same era.
Joint prompt optimization of stacked llms using variational inference
Sordoni, A., Yuan, E., Côté, M.-A., Pereira, M., Trischler, A., Xiao, Z., Hosseini, A., Niedtner, F., and Le Roux, N. (2023) · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Cited alongside, same era.
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al. (2024) · 2024
Later among the works it cites.
Numinamath
LI, J., Beeching, E., Tunstall, L., Lipkin, B., Soletskyi, R., Huang, S. C., Rasul, K., Yu, L., Jiang, A., Shen, Z., Qin, Z., Dong, B., Zhou, L., Fleureau, Y., Lample, G., and Polu, S. (2024) · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., Li, Y., Wu, Y., and Guo, D. (2024) · 2024
Later among the works it cites.
Hybridflow: A flexible and efficient rlhf framework
Sheng, G., Zhang, C., Ye, Z., Wu, X., Zhang, W., Zhang, R., Peng, Y., Lin, H., and Wu, C. (2024) · 2024
Later among the works it cites.
Generalized preference optimization: A unified approach to offline alignment
Tang, Y., Guo, Z. D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P. H., Valko, M., Pires, B. Á., and Piot, B. (2024) · 2024
Later among the works it cites.
Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving
Tong, Y., Zhang, X., Wang, R., Wu, R., and He, J. (2024) · 2024
Later among the works it cites.
Dpo meets ppo: Reinforced token optimization for rlhf
Zhong, H., Feng, G., Xiong, W., Zhao, L., He, D., Bian, J., and Wang, L. (2024) · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., Xue, B., Wang, B., Wu, B., Feng, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., Dai, D., Chen, D., Ji, D., Li, E., Lin, F., Dai, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Bao, H., Xu, H., Wang, H., Ding, H., Xin, H., Gao, H., Qu, H., Li, H., Guo, J., Li, J., Wang, J., Chen, J., Yuan, J., Qiu, J., Li, J., Cai, J. L., Ni, J., Liang, J., Chen, J., Dong, K., Hu, K., Gao, K., Guan, K., Huang, K., Yu, K., Wang, L., Zhang, L., Zhao, L., Wang, L., Zhang, L., Xu, L., Xia, L., Zhang, M., Zhang, M., Tang, M., Li, M., Wang, M., Li, M., Tian, N., Huang, P., Zhang, P., Wang, Q., Chen, Q., Du, Q., Ge, R., Zhang, R., Pan, R., Wang, R., Chen, R. J., Jin, R. L., Chen, R., Lu, S., Zhou, S., Chen, S., Ye, S., Wang, S., Yu, S., Zhou, S., Pan, S., Li, S. S., Zhou, S., Wu, S., Ye, S., Yun, T., Pei, T., Sun, T., Wang, T., Zeng, W., Zhao, W., Liu, W., Liang, W., Gao, W., Yu, W., Zhang, W., Xiao, W. L., An, W., Liu, X., Wang, X., Chen, X., Nie, X., Cheng, X., Liu, X., Xie, X., Liu, X., Yang, X., Li, X., Su, X., Lin, X., Li, X. Q., Jin, X., Shen, X., Chen, X., Sun, X., Wang, X., Song, X., Zhou, X., Wang, X., Shan, X., Li, Y. K., Wang, Y. Q., Wei, Y. X., Zhang, Y., Xu, Y., Li, Y., Zhao, Y., Sun, Y., Wang, Y., Yu, Y., Zhang, Y., Shi, Y., Xiong, Y., He, Y., Piao, Y., Wang, Y., Tan, Y., Ma, Y., Liu, Y., Guo, Y., Ou, Y., Wang, Y., Gong, Y., Zou, Y., He, Y., Xiong, Y., Luo, Y., You, Y., Liu, Y., Zhou, Y., Zhu, Y. X., Xu, Y., Huang, Y., Li, Y., Zheng, Y., Zhu, Y., Ma, Y., Tang, Y., Zha, Y., Yan, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Xu, Z., Xie, Z., Zhang, Z., Hao, Z., Ma, Z., Yan, Z., Wu, Z., Gu, Z., Zhu, Z., Liu, Z., Li, Z., Xie, Z., Song, Z., Pan, Z., Huang, Z., Xu, Z., Zhang, Z., and Zhang, Z. (2025) · 2025
Closest in time.
Reinforce++: A simple and efficient approach for aligning large language models
Hu, J. (2025) · 2025
Closest in time.
Efficient reinforcement finetuning via adaptive curriculum learning
Shi, T., Wu, Y., Song, L., Zhou, T., and Zhao, J. (2025) · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Team, K., Du, A., Gao, B., Xing, B., Jiang, C., Chen, C., Li, C., Xiao, C., Du, C., Liao, C., et al. (2025) · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale
Yu, Q., Zhang, Z., Zhu, R., Yuan, Y., Zuo, X., Yue, Y., Fan, T., Liu, G., Liu, L., Liu, X., et al. (2025) · 2025
Closest in time.
Brite: Bootstrapping reinforced thinking process to enhance language model reasoning
Zhong, H., Yin, Y., Zhang, S., Xu, X., Liu, Y., Zuo, Y., Liu, Z., Liu, B., Zheng, S., Guo, H., et al. (2025) · 2025
Closest in time.