Fetching the paper…
Reading the bibliography…
The reasoning capabilities of advanced large language models (LLMs) like o1 have revolutionized artificial intelligence applications.
Efficient selectivity and backup operators in monte-carlo tree search
Coulom, R · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, L. and Szepesvári, C · 2006
Earlier work this paper cites.
A survey of actor-critic reinforcement learning: Standard and natural policy gradients
Grondman, I., Busoniu, L., Lopes, G. A., and Babuska, R · 2012
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Complexity-based prompting for multi-step reasoning
Fu, Y., Peng, H., Sabharwal, A., Clark, P., and Khot, T · 2022
Earlier work this paper cites.
Solving math word problems with process-and outcome-based feedback
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2022
Earlier work this paper cites.
Tora: A tool-integrated reasoning agent for mathematical problem solving
Gou, Z., Shao, Z., Gong, Y., Shen, Y., Yang, Y., Huang, M., Duan, N., and Chen, W · 2023
Earlier work this paper cites.
Towards mitigating llm hallucination via self reflection
Ji, Z., Yu, T., Xu, Y., Lee, N., Ishii, E., and Fung, P · 2023
Earlier work this paper cites.
Challenges and applications of large language models
Kaddour, J., Harris, J., Mozes, M., Bradley, H., Raileanu, R., and McHardy, R · 2023
Earlier work this paper cites.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Earlier work this paper cites.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Luo, H., Sun, Q., Xu, C., Zhao, P., Lou, J., Tao, C., Geng, X., Lin, Q., Chen, S., and Zhang, D · 2023
Earlier work this paper cites.
Scaling data-constrained language models
Muennighoff, N., Rush, A., Barak, B., Le Scao, T., Tazi, N., Piktus, A., Pyysalo, S., Wolf, T., and Raffel, C. A · 2023
Earlier work this paper cites.
OpenAI · 2023
Earlier work this paper cites.
Pan, S., Lialin, V., Muckatira, S., and Rumshisky, A · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
Earlier work this paper cites.
Monte carlo tree search: A review of recent modifications and applications
Świechowski, M., Godlewski, K., Sawicki, B., and Mańdziuk, J · 2023
Earlier work this paper cites.
Skywork: A more open bilingual foundation model
Wei, T., Zhao, L., Zhang, L., Zhu, B., Wang, L., Yang, H., Li, B., Cheng, C., Lü, W., Hu, R., et al · 2023
Cited alongside, same era.
Fine-grained human feedback gives better rewards for language model training
Wu, Z., Hu, Y., Shi, W., Dziri, N., Suhr, A., Ammanabrolu, P., Smith, N. A., Ostendorf, M., and Hajishirzi, H · 2023
Cited alongside, same era.
Mammoth: Building math generalist models through hybrid instruction tuning
Yue, X., Qu, X., Zhang, G., Fu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2023
Cited alongside, same era.
Zhang, Y., Zhang, F., Yang, Z., and Wang, Z · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Llm critics help catch llm bugs
McAleese, N., Pokorny, R. M., Uribe, J. F. C., Nitishinskaya, E., Trebacz, M., and Leike, J · 2024
Later among the works it cites.
Inf-o1 ( π 0 \pi_{0} ): Initiating the journey to the infinity of llm reasoning, 2024
o1 Team, I · 2024
Later among the works it cites.
Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Qwen, T · 2024
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2024
Later among the works it cites.
Skywork-o1 open series
Skywork · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Cited alongside, same era.
Yi: Open foundation models by 01.ai, 2024
AI, ., :, Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., Yu, K., Liu, P., Liu, Q., Yue, S., Yang, S., Yang, S., Yu, T., Xie, W., Huang, W., Hu, X., Ren, X., Niu, X., Nie, P., Xu, Y., Liu, Y., Wang, Y., Cai, Y., Gu, Z., Liu, Z., and Dai, Z · 2024
Cited alongside, same era.
Graph of thoughts: Solving elaborate problems with large language models
Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Podstawski, M., Gianinazzi, L., Gajda, J., Lehmann, T., Niewiadomski, H., Nyczyk, P., et al · 2024
Cited alongside, same era.
Towards scalable automated alignment of llms: A survey
Cao, B., Lu, K., Lu, X., Chen, J., Ren, M., Xiang, H., Liu, P., Lu, Y., He, B., Han, X., et al · 2024
Cited alongside, same era.
A survey on in-context learning
Dong, Q., Li, L., Dai, D., Zheng, C., Ma, J., Li, R., Xia, H., Xu, J., Wu, Z., Chang, B., et al · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
He, C., Luo, R., Bai, Y., Hu, S., Thai, Z. L., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., Liu, J., Qi, L., Liu, Z., and Sun, M · 2024
Cited alongside, same era.
Inf’s open-source large language models
INF-Team · 2024
Cited alongside, same era.
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Team, Q · 2024
Later among the works it cites.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Wang, P., Li, L., Shao, Z., Xu, R., Dai, D., Li, Y., Chen, D., Wu, Y., and Sui, Z · 2024
Later among the works it cites.
Re-reading improves reasoning in large language models
Xu, X., Tao, C., Shen, T., Xu, C., Xu, H., Long, G., Lou, J.-G., and Ma, S · 2024
Later among the works it cites.
Qwen2.5-math technical report: Toward mathematical expert model via self-improvement
Yang, A., Zhang, B., Hui, B., Gao, B., Yu, B., Li, C., Liu, D., Tu, J., Zhou, J., Lin, J., Lu, K., Xue, M., Lin, R., Liu, T., Ren, X., and Zhang, Z · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Later among the works it cites.
Large language models as commonsense knowledge for large-scale task planning
Zhao, Z., Lee, W. S., and Hsu, D · 2024
Later among the works it cites.
Processbench: Identifying process errors in mathematical reasoning
Zheng, C., Zhang, Z., Zhang, B., Lin, R., Lu, K., Yu, B., Liu, D., Zhou, J., and Lin, J · 2024
Later among the works it cites.
Is your model really a good math reasoner? evaluating mathematical reasoning with checklist
Zhou, Z., Liu, S., Ning, M., Liu, W., Wang, J., Wong, D. F., Huang, X., Wang, Q., and Huang, K · 2024
Later among the works it cites.
Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence
Zhu, Q., Guo, D., Shao, Z., Yang, D., Wang, P., Xu, R., Wu, Y., Li, Y., Gao, H., Ma, S., et al · 2024
Later among the works it cites.
The lessons of developing process reward models in mathematical reasoning
Zhang, Z., Zheng, C., Wu, Y., Zhang, B., Lin, R., Yu, B., Liu, D., Zhou, J., and Lin, J · 2025
Closest in time.