Fetching the paper…
Reading the bibliography…
The integration of slow-thinking mechanisms into large language models (LLMs) offers a promising way toward achieving Level 2 AGI Reasoners, as exemplified by systems like OpenAI's o1.
Speech and language processing: an introduction to natural language processing, computational linguistics, and speech recognition, 2000
Teller, V · 2000
Earlier work this paper cites.
Backtracking search algorithms
Van Beek, P · 2006
Earlier work this paper cites.
Sequence transduction with recurrent neural networks, 2012
Graves, A · 2012
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Generative language modeling for automated theorem proving
Polu, S. and Sutskever, I · 2020
Earlier work this paper cites.
Making large language models better reasoners with step-aware verifier
Li, Y., Lin, Z., Zhang, S., Fu, Q., Chen, B., Lou, J.-G., and Chen, W · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2022
Earlier work this paper cites.
Chain of thought imitation with procedure cloning
Yang, M. S., Schuurmans, D., Abbeel, P., and Nachum, O · 2022
Earlier work this paper cites.
Learning from mistakes makes llm better reasoner
An, S., Ma, Z., Lin, Z., Zheng, N., Lou, J.-G., and Chen, W · 2023
Earlier work this paper cites.
Kcts: Knowledge-constrained tree search decoding with token-level hallucination detection
Choi, S., Fang, T., Wang, Z., and Song, Y · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Earlier work this paper cites.
Sequencematch: Imitation learning for autoregressive sequence modelling with backtracking
Cundy, C. and Ermon, S · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Kto: Model alignment as prospect theoretic optimization
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D · 2024
Cited alongside, same era.
Stream of search (sos): Learning to search in language
Gandhi, K., Lee, D., Grand, G., Liu, M., Cheng, W., Sharma, A., and Goodman, N. D · 2024
Cited alongside, same era.
Beyond a*: Better planning with transformers via search dynamics bootstrapping
Lehnert, L., Sukhbaatar, S., Su, D., Zheng, Q., McVay, P., Rabbat, M., and Tian, Y · 2024
Cited alongside, same era.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
Dualformer: Controllable fast and slow thinking by learning with randomized reasoning traces
Su, D., Sukhbaatar, S., Rabbat, M., Tian, Y., and Zheng, Q · 2024
Later among the works it cites.
Can llms learn from previous mistakes? investigating llms’ errors to boost for reasoning
Tong, Y., Li, D., Wang, S., Wang, Y., Teng, F., and Shang, J · 2024
Later among the works it cites.
Alphazero-like tree-search can guide large language model decoding and training
Wan, Z., Feng, X., Wen, M., McAleer, S. M., Wen, Y., Zhang, W., and Wang, J · 2024
Later among the works it cites.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Let’s verify step by step
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2024
Cited alongside, same era.
Guided stream of search: Learning to better search with language models via optimal path guidance
Moon, S., Park, B., and Song, H. O · 2024
Cited alongside, same era.
Learning to reason with large language models, September 2024
OpenAI · 2024
Cited alongside, same era.
O1 replication journey: A strategic progress report–part 1
Qin, Y., Li, X., Zou, H., Liu, Y., Xia, S., Huang, Z., Ye, Y., Yuan, W., Liu, H., Li, Y., et al · 2024
Cited alongside, same era.
Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Qwen · 2024
Cited alongside, same era.
Odin: Disentangled reward mitigates hacking in rlhf
Chen, L., Zhu, C., Soselia, D., Chen, J., Zhou, T., Goldstein, T., Huang, H., Shoeybi, M., and Catanzaro, B
Cited in the paper.
Do not think that much for 2+ 3=? on the overthinking of o1-like llms
Chen, X., Xu, J., Liang, T., He, Z., Pang, J., Yu, D., Song, L., Liu, Q., Zhou, M., Zhang, Z., et al
Cited in the paper.
Wang, P., Li, L., Shao, Z., Xu, R., Dai, D., Li, Y., Chen, D., Wu, Y., and Sui, Z · 2024
Later among the works it cites.
Inference scaling laws: An empirical analysis of compute-optimal inference for problem-solving with language models
Wu, Y., Sun, Z., Li, S., Welleck, S., and Yang, Y · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Later among the works it cites.
Physics of language models: Part 2.2, how to learn from mistakes on grade-school math problems
Ye, T., Xu, Z., Li, Y., and Allen-Zhu, Z · 2024
Later among the works it cites.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.