Fetching the paper…
Reading the bibliography…
While Large Language Models (LLMs) demonstrate impressive performance in mathematics, existing math benchmarks come with significant limitations.
Every planar map is four colorable , volume 98
Appel, K. I. and Haken, W · 1989
Earlier work this paper cites.
Neural combinatorial optimization with reinforcement learning
Bello, I., Pham, H., Le, Q. V., Norouzi, M., and Bengio, S · 2017
Earlier work this paper cites.
Exact combinatorial optimization with graph convolutional neural networks
Gasse, M., Chételat, D., Ferroni, N., Charlin, L., and Lodi, A · 2019
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Advancing mathematics by guiding human intuition with AI
Davies, A., Velickovic, P., Buesing, L., Blackwell, S., Zheng, D., Tomasev, N., Tanburn, R., Battaglia, P. W., Blundell, C., Juhász, A., Lackenby, M., Williamson, G., Hassabis, D., and Kohli, P · 2021
Earlier work this paper cites.
The lean 4 theorem prover and programming language
de Moura, L. and Ullrich, S · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Constructions in combinatorics via neural networks
Wagner, A. Z · 2021
Earlier work this paper cites.
minif2f: a cross-system benchmark for formal olympiad-level mathematics
Zheng, K., Han, J. M., and Polu, S · 2022
Earlier work this paper cites.
FIMO: A challenge formal dataset for automated theorem proving
Liu, C., Shen, J., Xin, H., Liu, Z., Yuan, Y., Wang, H., Ju, W., Zheng, C., Yin, Y., Li, L., Zhang, M., and Liu, Q · 2023
Earlier work this paper cites.
Claude 3.5 sonnet model card addendum
Anthropic, A · 2024
Earlier work this paper cites.
Artificial intelligence and machine learning generated conjectures with txgraffiti
Davila, R · 2024
Cited alongside, same era.
Constat: Performance-based contamination detection in large language models
Dekoninck, J., Müller, M. N., and Vechev, M. T · 2024
Cited alongside, same era.
Nphardeval: Dynamic benchmark on reasoning ability of large language models via complexity classes
Fan, L., Hua, W., Li, L., Ling, H., and Zhang, Y · 2024
Cited alongside, same era.
Omni-math: A universal olympiad level mathematic benchmark for large language models
Gao, B., Song, F., Yang, Z., Cai, Z., Miao, Y., Dong, Q., Li, L., Ma, C., Chen, L., Xu, R., Tang, Z., Wang, B., Zan, D., Quan, S., Zhang, G., Sha, L., Zhang, Y., Ren, X., Liu, T., and Chang, B · 2024
Cited alongside, same era.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in AI
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
Mirzadeh, S., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., and Farajtabar, M · 2024
Later among the works it cites.
Puzzlebench: Can llms solve challenging first-order combinatorial reasoning problems?
Mittal, C., Kartik, K., Mausam, and Singla, P · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T. P., Alayrac, J., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., Antonoglou, I., Anil, R., Borgeaud, S., Dai, A. M., Millican, K., Dyer, E., Glaese, M., Sottiaux, T., Lee, B., Viola, F., Reynolds, M., Xu, Y., Molloy, J., Chen, J., Isard, M., Barham, P., Hennigan, T., McIlroy, R., Johnson, M., Schalkwyk, J., Collins, E., Rutherford, E., Moreira, E., Ayoub, K., Goel, M., Meyer, C., Thornton, G., Yang, Z., Michalewski, H., Abbas, Z., Schucher, N., Anand, A., Ives, R., Keeling, J., Lenc, K., Haykal, S., Shakeri, S., Shyam, P., Chowdhery, A., Ring, R., Spencer, S., Sezener, E., and et al · 2024
Later among the works it cites.
Putnambench: Evaluating neural theorem-provers on the putnam mathematical competition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Glazer, E., Erdil, E., Besiroglu, T., Chicharro, D., Chen, E., Gunning, A., Olsson, C. F., Denain, J., Ho, A., de Oliveira Santos, E., Järviniemi, O., Barnett, M., Sandler, R., Vrzala, M., Sevilla, J., Ren, Q., Pratt, E., Levine, L., Barkley, G., Stewart, N., Grechuk, B., Grechuk, T., Enugandla, S. V., and Wildon, M · 2024
Cited alongside, same era.
Logicgame: Benchmarking rule-based reasoning abilities of large language models
Gui, J., Liu, Y., Cheng, J., Gu, X., Liu, X., Wang, H., Dong, Y., Tang, J., and Huang, M · 2024
Cited alongside, same era.
Putnam-AXIOM: A functional and static benchmark for measuring higher level mathematical reasoning
Gulati, A., Miranda, B., Chen, E., Xia, E., Fronsdal, K., de Moraes Dumont, B., and Koyejo, S · 2024
Cited alongside, same era.
Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems
He, C., Luo, R., Bai, Y., Hu, S., Thai, Z. L., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., Liu, J., Qi, L., Liu, Z., and Sun, M · 2024
Cited alongside, same era.
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al · 2024
Cited alongside, same era.
Kor-bench: Benchmarking language models on knowledge-orthogonal reasoning tasks
Ma, K., Du, X., Wang, Y., Zhang, H., Wen, Z., Qu, X., Yang, J., Liu, J., Liu, M., Yue, X., Huang, W., and Zhang, G · 2024
Cited alongside, same era.
Ueber eine elementare frage der mannigfaltigketislehre
Cantor, G
Cited in the paper.
Tsoukalas, G., Lee, J., Jennings, J., Xin, J., Ding, M., Jennings, M., Thakur, A., and Chaudhuri, S · 2024
Later among the works it cites.
Deepseek-prover: Advancing theorem proving in llms through large-scale synthetic data
Xin, H., Guo, D., Shao, Z., Ren, Z., Zhu, Q., Liu, B., Ruan, C., Li, W., and Liang, X · 2024
Later among the works it cites.
Utmath: Math evaluation with unit test via reasoning-to-coding thoughts
Yang, B., Yang, Q., and Liu, R · 2024
Later among the works it cites.
HARP: A challenging human-annotated math reasoning benchmark
Yue, A. S., Madaan, L., Moskovitz, T., Strouse, D., and Singh, A. K · 2024
Later among the works it cites.
URL https://www.fastht.ml/
Fasthtml, 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
Humanity’s last exam
Phan, L., Gatti, A., Han, Z., Li, N., Hu, J., Zhang, H., Shi, S., Choi, M., Agrawal, A., Chopra, A., Khoja, A., Kim, R., Hausenloy, J., Zhang, O., et al · 2025
Closest in time.