Fetching the paper…
Reading the bibliography…
Difficult problems, which often result in long reasoning traces, are widely recognized as key factors for enhancing the performance of reasoning models.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T.B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., Amodei, D., 2020 · 2001
Earlier work this paper cites.
Language models are few-shot learners
Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D.M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., teusz Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., Amodei, D., 2020 · 2005
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.E., Vinyals, O., Dean, J., 2015 · 2015
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D.X., Steinhardt, J., 2021 · 2021
Earlier work this paper cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L.A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J.W., Vinyals, O., Sifre, L., 2022 · 2022
Earlier work this paper cites.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E.H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., Fedus, W., 2022 · 2022
Earlier work this paper cites.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., Cobbe, K., 2023 · 2023
Earlier work this paper cites.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Luo, H., Sun, Q., Xu, C., Zhao, P., Lou, J.G., Tao, C., Geng, X., Lin, Q., Chen, S., Zhang, D., 2023 · 2023
Earlier work this paper cites.
Gpqa: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B.L., Stickland, A.C., Petty, J., Pang, R.Y., Dirani, J., Michael, J., Bowman, S.R., 2023 · 2023
Earlier work this paper cites.
Deepseek llm: Scaling open-source language models with longtermism
Bi, D.A.X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., Gao, H., Gao, K., Gao, W., Ge, R., Guan, K., Guo, D., Guo, J., Hao, G., Hao, Z., He, Y., Hu, W.H., Huang, P., Li, E., Li, G., Li, J., Li, Y., Li, Y.K., Liang, W., Lin, F., Liu, A., Liu), B.L.B., Liu, W., Liu, X., Liu, X., Liu, Y., Lu, H., Lu, S., Luo, F., Ma, S., Nie, X., Pei, T., Piao, Y., Qiu, J., Qu, H., Ren, T., Ren, Z., Ruan, C., Sha, Z., Shao, Z., Song, J.M., Su, X., Sun, J., Sun, Y., Tang, M., Wang, B.L., Wang, P., Wang, S., Wang, Y., Wang, Y., Wu, T., Wu, Y., Xie, X., Xie, Z., Xie, Z., Xiong, Y., Xu, H., Xu, R.X., Xu, Y., Yang, D., mei You, Y., Yu, S., yuan Yu, X., Zhang, B., Zhang, H., Zhang, L., Zhang, L., Zhang, M., Zhang, M., Zhang, W., Zhang, Y., Zhao, C., Zhao, Y., Zhou, S., Zhou, S., Zhu, Q., Zou, Y., 2024 · 2024
Cited alongside, same era.
Unleashing reasoning capability of llms via scalable question synthesis from scratch
Ding, Y., Shi, X., Liang, X., Li, J., Zhu, Q., Zhang, M., 2024 · 2024
Cited alongside, same era.
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al., 2024 · 2024
Cited alongside, same era.
Yang, Q.A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., Lu, K., Bao, K., Yang, K., Yu, L., Li, M., Xue, M., Zhang, P., Zhu, Q., Men, R., Lin, R., Li, T., Xia, T., Ren, X., Ren, X., Fan, Y., Su, Y., Zhang, Y.C., Wan, Y., Liu, Y., Cui, Z., Zhang, Z., Qiu, Z., Quan, S., Wang, Z., 2024 · 2024
Later among the works it cites.
Ballon, M., Algaba, A., Ginis, V., 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al., 2025 · 2025
Closest in time.
Gemini 2.0 is now available to everyone
Kavukcuoglu, K., 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Min, Y., Chen, Z., Jiang, J., Chen, J., Deng, J., Hu, Y., Tang, Y., Wang, J., Cheng, X., Song, H., Zhao, W.X., Liu, Z., Wang, Z., Wen, J., 2024 · 2024
Cited alongside, same era.
Learning to reason with llms
OpenAI, 2024 · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J.M., Zhang, M., Li, Y.K., Wu, Y., Guo, D., 2024 · 2024
Cited alongside, same era.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., Kumar, A., 2024 · 2024
Cited alongside, same era.
Drt: Deep reasoning translation via long chain-of-thought
Wang, J., Meng, F., Liang, Y., Zhou, J., 2024 · 2024
Cited alongside, same era.
S*: Test time scaling for code generation
Li, D., Cao, S., Cao, C., Li, X., Tan, S., Keutzer, K., Xing, J., Gonzalez, J., Stoica, I., 2025a
Cited in the paper.
Llms can easily learn to reason from demonstrations structure, not content, is what matters!
Li, D., Cao, S., Griggs, T., Liu, S., Mo, X., Tang, E., Hegde, S., Hakhamaneshi, K., Patil, S.G., Zaharia, M., Gonzalez, J., Stoica, I., 2025b
Cited in the paper.
From system 1 to system 2: A survey of reasoning large language models
Li, Z., Zhang, D., Zhang, M.L., Zhang, J., Liu, Z., Yao, Y., Xu, H., Zheng, J., Wang, P.J., Chen, X., Zhang, Y., Yin, F., Dong, J., Guo, Z., Song, L., Liu, C.L., 2025c
Cited in the paper.
Open Thoughts
Team, O., 2025a
Cited in the paper.
Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl
Luo, M., Tan, S., Wong, J., Shi, X., Tang, W.Y., Roongta, M., Cai, C., Luo, J., Zhang, T., Li, L.E., Popa, R.A., Stoica, I., 2025 · 2025
Closest in time.
Muennighoff, N., Yang, Z., Shi, W., Li, X.L., Li, F.F., Hajishirzi, H., Zettlemoyer, L.S., Liang, P., Candes, E.J., Hashimoto, T., 2025 · 2025
Closest in time.
Limo: Less is more for reasoning
Ye, Y., Huang, Z., Xiao, Y., Chern, E., Xia, S., Liu, P., 2025 · 2025
Closest in time.
Sift: Grounding llm reasoning in contexts via stickers
Zeng, Z., Huang, X., Li, B., Deng, Z., 2025 · 2025
Closest in time.