Fetching the paper…
Reading the bibliography…
Reasoning LLMs such as OpenAI o1, o3 and DeepSeek R1 have made significant progress in mathematics and coding, yet find challenging advanced tasks such as International Mathematical Olympiad (IMO) combinatorics problems, Abstraction and Reasoning Corpus (ARC) puzzles, and Humanity's Last Exam (HLE) questions.
Problem-Solving Strategies (Problem Books in Mathematics)
Engel, A · 1997
Earlier work this paper cites.
Combinatorics: A Problem Oriented Approach
Marcus, D. A · 1999
Earlier work this paper cites.
Isabelle/HOL: a proof assistant for higher-order logic
Nipkow, T., Wenzel, M., and Paulson, L. C · 2002
Earlier work this paper cites.
An Introduction to Game Theory
Osborne, M. J · 2003
Earlier work this paper cites.
How to Prove It: A Structured Approach
Velleman, D. J · 2006
Earlier work this paper cites.
The Art and Craft of Problem Solving
Zeitz, P · 2007
Earlier work this paper cites.
Z3: An efficient SMT solver
De Moura, L. and Bjørner, N · 2008
Earlier work this paper cites.
R* search
Likhachev, M. and Stentz, A · 2008
Earlier work this paper cites.
Combinatorics
Zhao, Y · 2008
Earlier work this paper cites.
Mathematical Olympiad Challenges
Andreescu, T. and Razvan, G · 2009
Earlier work this paper cites.
Problems from the Book
Andreescu, T. and Dospinescu, G · 2010
Earlier work this paper cites.
The IMO Compendium: A Collection of Problems Suggested for The International Mathematical Olympiads
Djukić, D., Janković, Vladimir Matić, I., and Petrović, N · 2011
Earlier work this paper cites.
Straight from the Book
Andreescu, T. and Dospinescu, G · 2012
Earlier work this paper cites.
Mathematical Olympiad Treasures
Andreescu, T. and Enescu, B · 2012
Earlier work this paper cites.
Dynamic Programming and Optimal Control
Bertsekas, D. P · 2012
Earlier work this paper cites.
Expected uses of probability
Chen, E · 2014
Earlier work this paper cites.
Combinatorics: A Very Short Introduction
Wilson, R. J · 2016
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
On the measure of intelligence
Chollet, F · 2019
Earlier work this paper cites.
Learning to prove theorems via interacting with proof assistants
Yang, K. and Deng, J · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
GShard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Earlier work this paper cites.
IsarStep: A benchmark for high-level mathematical reasoning
Li, W., Yu, L., Wu, Y., and Paulson, L. C · 2020
Earlier work this paper cites.
Generative language modeling for automated theorem proving
Polu, S. and Sutskever, I · 2020
Earlier work this paper cites.
Test-time training with self-supervision for generalization under distribution shifts
Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A., and Hardt, M · 2020
Earlier work this paper cites.
Maintaining a library of formal mathematics
van Doorn, F., Ebner, G., and Lewis, R. Y · 2020
Earlier work this paper cites.
Exploration of neural machine translation in autoformalization of mathematics in mizar
Wang, Q., Brown, C., Kaliszyk, C., and Urban, J · 2020
Earlier work this paper cites.
INT: An inequality benchmark for evaluating generalization in theorem proving
Wu, Y., Jiang, A. Q., Ba, J., and Grosse, R · 2020
Earlier work this paper cites.
From the author’s side: Thoughts on problem writing
Chen, E · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Advancing Mathematics by guiding human intuition with AI
Davies, A., Veličković, P., Buesing, L., Blackwell, S., Zheng, D., Tomašev, N., Tanburn, R., Battaglia, P., Blundell, C., Juhász, A., et al · 2021
Cited alongside, same era.
Scaling scaling laws with board games
Jones, A. L · 2021
Cited alongside, same era.
The Lean 4 theorem prover and programming language
Moura, L. d. and Ullrich, S · 2021
Cited alongside, same era.
MiniF2F: A cross-system benchmark for formal Olympiad-level Mathematics
Zheng, K., Han, J. M., and Polu, S · 2021
Cited alongside, same era.
Caballero, E., Gupta, K., Rish, I., and Krueger, D · 2022
Cited alongside, same era.
LeanAgent: Lifelong learning for formal theorem proving
Kumarappan, A., Tiwari, M., Song, P., George, R. J., Xiao, C., and Anandkumar, A · 2024
Later among the works it cites.
BAIT: Benchmarking (embedding) architectures for interactive theorem-proving
Lamont, S., Norrish, M., Dezfouli, A., Walder, C., and Montague, P · 2024
Later among the works it cites.
H-ARC: A robust estimate of human performance on the abstraction and reasoning corpus benchmark
LeGris, S., Vong, W. K., Lake, B. M., and Gureckis, T. M · 2024
Later among the works it cites.
FVEL: Interactive formal verification environment with large language models via theorem proving
Lin, X., Cao, Q., Huang, Y., Wang, H., Lu, J., Liu, Z., Song, L., and Liang, X · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Discovering faster matrix multiplication algorithms with reinforcement learning
Fawzi, A., Balog, M., Huang, A., Hubert, T., Romera-Paredes, B., Barekatain, M., Novikov, A., R Ruiz, F. J., Schrittwieser, J., Swirszcz, G., et al · 2022
Cited alongside, same era.
Switch Transformers: Scaling to Trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Cited alongside, same era.
HyperTree proof search for neural theorem proving
Lample, G., Lacroix, T., Lachaux, M.-A., Rodriguez, A., Hayat, A., Lavril, T., Ebner, G., and Martinet, X · 2022
Cited alongside, same era.
Formal mathematics statement curriculum learning
Polu, S., Han, J. M., Zheng, K., Baksys, M., Babuschkin, I., and Sutskever, I · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2022
Cited alongside, same era.
Autoformalization with large language models
Wu, Y., Jiang, A. Q., Li, W., Rabe, M., Staats, C., Jamnik, M., and Szegedy, C · 2022
Cited alongside, same era.
Lu, J., Wan, Y., Liu, Z., Huang, Y., Xiong, J., Liu, C., Shen, J., Jin, H., Zhang, J., Wang, H., et al · 2024
Later among the works it cites.
End-to-end neuro-symbolic reinforcement learning with textual explanations
Luo, L., Zhang, G., Xu, H., Yang, Y., Fang, C., and Li, Q · 2024
Later among the works it cites.
Official 2023 IMO Shortlist problems
Matsumoto, Y., Yamauchi, A., Ito, T., Kodama, H., Minegishi, R., Shimizu, G., Kitamura, T., Takaya, Y., Kim, D., García, E. L., Maret, A., Vaderlind, P., Ando, T., Guo, I., Kós, G., and Bealing, S · 2024
Later among the works it cites.
Artificial Intelligence for Science (AI4S): Frontiers and Perspectives Based on Parallel Intelligence
Miao, Q · 2024
Later among the works it cites.
Learning to reason with LLMs
OpenAI · 2024
Later among the works it cites.
Phenomenal yet puzzling: Testing inductive reasoning capabilities of language models with hypothesis refinement
Qiu, L., Jiang, L., Lu, X., Sclar, M., Pyatkin, V., Bhagavatula, C., Wang, B., Kim, Y., Choi, Y., Dziri, N., et al · 2024
Later among the works it cites.
Mathematical discoveries from program search with large language models
Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J., Ellenberg, J. S., Wang, P., Fawzi, O., et al · 2024
Later among the works it cites.
Towards large language models as copilots for theorem proving in Lean
Song, P., Yang, K., and Anandkumar, A · 2024
Later among the works it cites.
Inference Scaling 𝙵 \mathtt{F} Laws: The Limits of LLM Resampling with Imperfect Verifiers
Stroebl, B., Kapoor, S., and Narayanan, A · 2024
Later among the works it cites.
Mathstodon
Tao, T · 2024
Later among the works it cites.
The coq proof assistant (8.19), 2024
The Coq Development Team · 2024
Later among the works it cites.
Official 2024 IMO problems with solutions
Thomas, A., Ai, Y., Ng, A., Kós, G., Guo, I., Carlotti, A., Aaronson, J., Bealing, S., Agisilaou, A., Cranch, J., Myers, J., Yau, H., Ivan, M.-R., Ren, M., and García, E. L · 2024
Later among the works it cites.
Solving Olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T · 2024
Later among the works it cites.
PutnamBench: Evaluating neural theorem-provers on the Putnam mathematical competition
Tsoukalas, G., Lee, J., Jennings, J., Xin, J., Ding, M., Jennings, M., Thakur, A., and Chaudhuri, S · 2024
Later among the works it cites.
Official 2024 usamo problems with solutions
Unites States of America Mathematical Olympiad · 2024
Later among the works it cites.
Proving Olympiad algebraic inequalities without human demonstrations
Wei, C., Sun, M., and Wang, W · 2024
Later among the works it cites.
Wu, Z., Huang, S., Zhou, Z., Ying, H., Wang, J., Lin, D., and Chen, K · 2024
Later among the works it cites.
Monte carlo tree search boosts reasoning via iterative preference learning
Xie, Y., Goyal, A., Zheng, W., Kan, M.-Y., Lillicrap, T. P., Kawaguchi, K., and Shieh, M · 2024
Later among the works it cites.
DeepSeek-Prover: Advancing theorem proving in LLMs through large-scale synthetic data
Xin, H., Guo, D., Shao, Z., Ren, Z., Zhu, Q., Liu, B., Ruan, C., Li, W., and Liang, X · 2024
Later among the works it cites.
LeanDojo: Theorem proving with retrieval-augmented language models
Yang, K., Swope, A., Gu, A., Chalamala, R., Song, P., Yu, S., Godil, S., Prenger, R. J., and Anandkumar, A · 2024
Later among the works it cites.
Quiet-Star: Language models can teach themselves to think before speaking
Zelikman, E., Harik, G., Shao, Y., Jayasiri, V., Haber, N., and Goodman, N. D · 2024
Later among the works it cites.
Gold-medalist performance in solving Olympiad geometry with AlphaGeometry2
Chervonyi, Y., Trinh, T. H., Olšák, M., Yang, X., Nguyen, H., Menegali, M., Jung, J., Verma, V., Le, Q. V., and Luong, T · 2025
Closest in time.
Competitive programming with large reasoning models
El-Kishky, A., Wei, A., Saraiva, A., Minaev, B., Selsam, D., Dohan, D., Song, F., Lightman, H., Clavera, I., Pachocki, J., et al · 2025
Closest in time.
Evaluation platform
Gentrace · 2025
Closest in time.
DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
Phan, L., Gatti, A., Han, Z., Li, N., Hu, J., Zhang, H., Shi, S., Choi, M., Agrawal, A., Chopra, A., Khoja, A., Kim, R., Ren, R., Hausenloy, J., Zhang, O., Mazeika, M., Yue, S., Wang, A., and Hendrycks, D · 2025
Closest in time.