Fetching the paper…
Reading the bibliography…
Recent advances in automated theorem proving (ATP) through LLMs have highlighted the potential of formal reasoning with Lean 4 codes.
Isabelle: A generic theorem prover
L. C. Paulson · 1994
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
R. Coulom · 2006
Earlier work this paper cites.
The Lean theorem prover (system description)
L. De Moura, S. Kong, J. Avigad, F. Van Doorn, and J. von Raumer · 2015
Earlier work this paper cites.
Generative language modeling for automated theorem proving
S. Polu and I. Sutskever · 2020
Earlier work this paper cites.
The lean 4 theorem prover and programming language
L. d. Moura and S. Ullrich · 2021
Earlier work this paper cites.
Hypertree proof search for neural theorem proving
G. Lample, T. Lacroix, M.-A. Lachaux, A. Rodriguez, A. Hayat, T. Lavril, G. Ebner, and X. Martinet · 2022
Earlier work this paper cites.
Formal mathematics statement curriculum learning
S. Polu, J. M. Han, K. Zheng, M. Baksys, I. Babuschkin, and I. Sutskever · 2022
Earlier work this paper cites.
Llemma: An open language model for mathematics
Z. Azerbayev, H. Schoelkopf, K. Paster, M. D. Santos, S. McAleer, A. Q. Jiang, J. Deng, S. Biderman, and S. Welleck · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2023
Earlier work this paper cites.
Leandojo: Theorem proving with retrieval-augmented language models
K. Yang, A. Swope, A. Gu, R. Chalamala, P. Song, S. Yu, S. Godil, R. J. Prenger, and A. Anandkumar · 2023
Earlier work this paper cites.
Alphaproof and Alphageometry, July 2024
DeepMind · 2024
Earlier work this paper cites.
Qwen2. 5-coder technical report
B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Dang, et al · 2024
Cited alongside, same era.
Lean-star: Learning to interleave thinking and proving
H. Lin, Z. Sun, Y. Yang, and S. Welleck · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al · 2024
Cited alongside, same era.
Qwen2.5: A party of foundation models, September 2024
Q. Team · 2024
Cited alongside, same era.
Solving olympiad geometry without human demonstrations
T. H. Trinh, Y. Wu, Q. V. Le, H. He, and T. Luong · 2024
Cited alongside, same era.
Deepseek-R1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Goedel-prover: A frontier model for open-source automated theorem proving
Y. Lin, S. Tang, B. Lyu, J. Wu, H. Lin, K. Yang, J. Li, M. Xia, D. Chen, S. Arora, et al · 2025
Closest in time.
Understanding r1-zero-like training: A critical perspective
Z. Liu, C. Chen, W. Li, P. Qi, T. Pang, C. Du, W. S. Lee, and M. Lin · 2025
Closest in time.
The math library of lean 4, 2025
mathlib4 · 2025
Closest in time.
Qwq-32b: Embracing the power of reinforcement learning, March 2025
Q. Team · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Theoremllama: Transforming general-purpose llms into lean4 experts
R. Wang, J. Zhang, Y. Jia, R. Pan, S. Diao, R. Pi, and T. Zhang · 2024
Cited alongside, same era.
Lean workbook: A large-scale lean problem set formalized from natural language math problems
H. Ying, Z. Wu, Y. Geng, J. Wang, D. Lin, and K. Chen · 2024
Cited alongside, same era.
Minif2f: a cross-system benchmark for formal olympiad-level mathematics
K. Zheng, J. M. Han, and S. Polu · 2024
Cited alongside, same era.
Claude 3.7 Sonnet System card
Anthropic · 2025
Cited alongside, same era.
Stp: Self-play llm theorem provers with iterative conjecturing and proving
K. Dong and T. Ma · 2025
Cited alongside, same era.
Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective stars
K. Gandhi, A. Chakravarthy, A. Singh, N. Lile, and N. D. Goodman · 2025
Cited alongside, same era.
Art of problem solving
AoPS
Cited in the paper.
Z. Wan, Y. Li, Y. Song, H. Wang, L. Yang, M. Schmidt, J. Wang, W. Zhang, S. Hu, and Y. Wen · 2025
Closest in time.
Ma-lot: Multi-agent lean-based long chain-of-thought reasoning enhances formal theorem proving
R. Wang, R. Pan, Y. Li, J. Zhang, Y. Jia, S. Diao, R. Pi, J. Hu, and T. Zhang · 2025
Closest in time.
Bfs-prover: Scalable best-first tree search for llm-based automatic theorem proving
R. Xin, C. Xi, J. Yang, F. Chen, H. Wu, X. Xiao, Y. Sun, S. Zheng, and K. Shen · 2025
Closest in time.
Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild
W. Zeng, Y. Huang, Q. Liu, W. Liu, K. He, Z. Ma, and J. He · 2025
Closest in time.
Promptcot: Synthesizing olympiad-level problems for mathematical reasoning in large language models
X. Zhao, W. Wu, J. Guan, and L. Kong · 2025
Closest in time.