Fetching the paper…
Reading the bibliography…
Current approaches for training Process Reward Models (PRMs) often involve breaking down responses into multiple reasoning steps using rule-based techniques, such as using predefined placeholder tokens or setting the reasoning step's length into a fixed size.
Efficient selectivity and backup operators in monte-carlo tree search
Coulom, R · 2006
Earlier work this paper cites.
Thinking, fast and slow
Kahneman, D · 2011
Earlier work this paper cites.
Solving general arithmetic word problems, 2016
Roy, S. and Roth, D · 2016
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R · 2019
Earlier work this paper cites.
Gedi: Generative discriminator guided sequence generation
Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F · 2020
Earlier work this paper cites.
PPL-MCTS: Constrained Textual Generation Through Discriminator-Guided Decoding
Chaffin, A., Claveau, V., and Kijak, E · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Solving math word problems with process- and outcome-based feedback, 2022
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Earlier work this paper cites.
Kcts: knowledge-constrained tree search decoding with token-level hallucination detection
Choi, S., Fang, T., Wang, Z., and Song, Y · 2023
Earlier work this paper cites.
Alphazero-like tree-search can guide large language model decoding and training
Feng, X., Wan, Z., Wen, M., McAleer, S. M., Wen, Y., Zhang, W., and Wang, J · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model, 2023
Hao, S., Gu, Y., Ma, H., Hong, J. J., Wang, Z., Wang, D. Z., and Hu, Z · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Earlier work this paper cites.
Let’s verify step by step, 2023
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Earlier work this paper cites.
Critique ability of large language models, 2023
Luo, L., Lin, Z., Liu, Y., Shu, L., Zhu, Y., Shang, J., and Meng, L · 2023
Earlier work this paper cites.
Let’s reward step by step: Step-level reward model as the navigators for reasoning
Ma, Q., Zhou, H., Liu, T., Yuan, J., Liu, P., You, Y., and Yang, H · 2023
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models, 2023
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Cited alongside, same era.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J. T., Li, Z., Weller, A., and Liu, W · 2023
Cited alongside, same era.
Planning with large language models for code generation
Zhang, S., Chen, Z., Shen, Y., Ding, M., Tenenbaum, J. B., and Gan, C · 2023
Cited alongside, same era.
Step-level value preference optimization for mathematical reasoning
Chen, G., Liao, M., Li, C., and Fan, K · 2024
Cited alongside, same era.
Decoding secret memorization in code llms through token-level characterization
Nie, Y., Wang, C., Wang, K., Xu, G., Xu, G., and Wang, H · 2024
Later among the works it cites.
O1 replication journey: A strategic progress report – part 1, 2024
Qin, Y., Li, X., Zou, H., Liu, Y., Xia, S., Huang, Z., Ye, Y., Yuan, W., Liu, H., Li, Y., and Liu, P · 2024
Later among the works it cites.
Bond: Aligning llms with best-of-n distillation, 2024
Sessa, P. G., Dadashi, R., Hussenot, L., Ferret, J., Vieillard, N., Ramé, A., Shariari, B., Perrin, S., Friesen, A., Cideron, G., Girgin, S., Stanczyk, P., Michi, A., Sinopalnikov, D., Ramos, S., Héliou, A., Severyn, A., Hoffman, M., Momchev, N., and Bachem, O · 2024
Later among the works it cites.
Rewarding progress: Scaling automated process verifiers for llm reasoning, 2024
Setlur, A., Nagpal, C., Fisch, A., Geng, X., Eisenstein, J., Agarwal, R., Agarwal, A., Berant, J., and Kumar, A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gao, B., Cai, Z., Xu, R., Wang, P., Zheng, C., Lin, R., Lu, K., Liu, D., Zhou, C., Xiao, W., Hu, J., Liu, T., and Chang, B · 2024
Cited alongside, same era.
The llama 3 herd of models, 2024
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., and et.al · 2024
Cited alongside, same era.
Guo, D., Zhu, Q., Yang, D., Xie, Z., Dong, K., Zhang, W., Chen, G., Bi, X., Wu, Y., Li, Y., Luo, F., Xiong, Y., and Liang, W · 2024
Cited alongside, same era.
Using logprobs, 2024
Hills, J. and Anadkat, S · 2024
Cited alongside, same era.
Large language models cannot self-correct reasoning yet, 2024
Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., and Zhou, D · 2024
Cited alongside, same era.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Jain, N., Han, K., Gu, A., Li, W.-D., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., and Stoica, I · 2024
Cited alongside, same era.
Step-dpo: Step-wise preference optimization for long-chain reasoning of llms, 2024
Lai, X., Tian, Z., Chen, Y., Yang, S., Peng, X., and Jia, J · 2024
Cited alongside, same era.
Lee, J. H., Yang, J. Y., Heo, B., Han, D., and Yoo, K. M · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y. K., Wu, Y., and Guo, D · 2024
Later among the works it cites.
Policy filtration in rlhf to fine-tune llm for code generation
Shen, W. and Zhang, C · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
Llms cannot find reasoning errors, but can correct them given the error location, 2024
Tyen, G., Mansoor, H., Cărbune, V., Chen, P., and Mak, T · 2024
Later among the works it cites.
Evaluating mathematical reasoning beyond accuracy, 2024
Xia, S., Li, X., Liu, Y., Wu, T., and Liu, P · 2024
Later among the works it cites.
Safedecoding: Defending against jailbreak attacks via safety-aware decoding
Xu, Z., Jiang, F., Niu, L., Jia, J., Lin, B. Y., and Poovendran, R · 2024
Later among the works it cites.
Free process rewards without process labels, 2024
Yuan, L., Li, W., Chen, H., Cui, G., Ding, N., Zhang, K., Zhou, B., Liu, Z., and Peng, H · 2024
Later among the works it cites.
Entropy-regularized process reward model, 2024
Zhang, H., Wang, P., Diao, S., Lin, Y., Pan, R., Dong, H., Zhang, D., Molchanov, P., and Zhang, T · 2024
Later among the works it cites.
Stepwise self-consistent mathematical reasoning with large language models, 2024
Zhao, Z., Rong, Y., Guo, D., Gözlüklü, E., Gülboy, E., and Kasneci, E · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., and et.al · 2025
Closest in time.
Kimi k1.5: Scaling reinforcement learning with llms, 2025
Team, K., Du, A., Gao, B., Xing, B., Jiang, C., Chen, C., Li, C., Xiao, C., Du, C., Liao, C., Tang, C., Wang, C., Zhang, D., Yuan, E., Lu, E., Tang, F., Sung, F., Wei, G., Lai, G., Guo, H., Zhu, H., Ding, H., Hu, H., Yang, H., Zhang, H., Yao, H., Zhao, H., Lu, H., Li, H., Yu, H., Gao, H., Zheng, H., Yuan, H., Chen, J., Guo, J., Su, J., Wang, J., Zhao, J., Zhang, J., Liu, J., Yan, J., Wu, J., Shi, L., Ye, L., Yu, L., Dong, M., Zhang, N., Ma, N., Pan, Q., Gong, Q., Liu, S., Ma, S., Wei, S., Cao, S., Huang, S., Jiang, T., Gao, W., Xiong, W., He, W., Huang, W., Wu, W., He, W., Wei, X., Jia, X., Wu, X., Xu, X., Zu, X., Zhou, X., Pan, X., Charles, Y., Li, Y., Hu, Y., Liu, Y., Chen, Y., Wang, Y., Liu, Y., Qin, Y., Liu, Y., Yang, Y., Bao, Y., Du, Y., Wu, Y., Wang, Y., Zhou, Z., Wang, Z., Li, Z., Zhu, Z., Zhang, Z., Wang, Z., Yang, Z., Huang, Z., Huang, Z., Xu, Z., and Yang, Z · 2025
Closest in time.