Fetching the paper…
Reading the bibliography…
Test-time scaling, which is also often referred to as slow-thinking, has been demonstrated to enhance multi-step reasoning in large language models (LLMs).
The power of word clusters for text classification
Slonim, N., Tishby, N., et al · 2001
Earlier work this paper cites.
Measuring statistical dependence with hilbert-schmidt norms
Gretton, A., Bousquet, O., Smola, A., and Schölkopf, B · 2005
Earlier work this paper cites.
Fano inequality
Fano, R. M · 2008
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Information-theoretic analysis of generalization capability of learning algorithms
Xu, A. and Raginsky, M · 2017
Earlier work this paper cites.
Multi-task image clustering through correlation propagation
Hu, S., Yan, X., and Ye, Y · 2019
Earlier work this paper cites.
How much does your data exploration overfit? controlling bias via information usage
Russo, D. and Zou, J · 2019
Earlier work this paper cites.
West, P., Holtzman, A., Buys, J., and Choi, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
The hsic bottleneck: Deep learning without back-propagation
Ma, W.-D. K., Lewis, J., and Kleijn, W. B · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Saparov, A. and He, H · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., Mu, J., and Goodman, N · 2022
Cited alongside, same era.
Alphazero-like tree-search can guide large language model decoding and training
Feng, X., Wan, Z., Wen, M., McAleer, S. M., Wen, Y., Zhang, W., and Wang, J · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
Hao, S., Gu, Y., Ma, H., Hong, J., Wang, Z., Wang, D., and Hu, Z · 2023
Cited alongside, same era.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Cited alongside, same era.
Skywork-reward: Bag of tricks for reward modeling in llms
Liu, C. Y., Zeng, L., Liu, J., Yan, R., He, J., Wang, C., Yan, S., Liu, Y., and Zhou, Y · 2024
Later among the works it cites.
Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning systems
Min, Y., Chen, Z., Jiang, J., Chen, J., Deng, J., Hu, Y., Tang, Y., Wang, J., Cheng, X., Song, H., et al · 2024
Later among the works it cites.
Skywork-o1 open series
o1 Team, S · 2024
Later among the works it cites.
Learning to reason with llms, 2024
OpenAI · 2024
Later among the works it cites.
Towards tracing trustworthiness dynamics: Revisiting pre-training period of large language models
Qian, C., Zhang, J., Yao, W., Liu, D., Yin, Z., Qiao, Y., Liu, Y., and Shao, J · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuan, Z., Yuan, H., Li, C., Dong, G., Lu, K., Tan, C., Zhou, C., and Zhou, J · 2023
Cited alongside, same era.
The pitfalls of next-token prediction
Bachmann, G. and Nagarajan, V · 2024
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling
Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., and Mirhoseini, A · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Gan, Z. and Liu, Y · 2024
Cited alongside, same era.
Position: The platonic representation hypothesis
Huh, M., Cheung, B., Wang, T., and Isola, P · 2024
Cited alongside, same era.
Technical report: Enhancing llm reasoning with reward-guided tree search
Jiang, J., Chen, Z., Min, Y., Chen, J., Cheng, X., Wang, J., Tang, Y., Sun, H., Deng, J., Zhao, W. X., et al · 2024
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
Understanding chain-of-thought in llms through information theory
Ton, J.-F., Taufiq, M. F., and Liu, Y · 2024
Later among the works it cites.
Helpsteer2-preference: Complementing ratings with preferences, 2024
Wang, Z., Bukharin, A., Delalleau, O., Egert, D., Shen, G., Zeng, J., Kuchaiev, O., and Dong, Y · 2024
Later among the works it cites.
Wu, Y., Sun, Z., Li, S., Welleck, S., and Yang, Y · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Later among the works it cites.
Rest-mcts*: Llm self-training via process reward guided tree search
Zhang, D., Zhoubian, S., Hu, Z., Yue, Y., Dong, Y., and Tang, J · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.