Fetching the paper…
Reading the bibliography…
Advances in Large Language Models (LLMs) have sparked interest in their ability to solve Olympiad-level math problems.
Mawps: A math word problem repository
Koncel-Kedziorski, R., Roy, S., Amini, A., Kushman, N., and Hajishirzi, H · 2016
Earlier work this paper cites.
Sympy: symbolic computing in python
Meurer, A., Smith, C. P., Paprocki, M., Čertík, O., Kirpichev, S. B., Rocklin, M., Kumar, A., Ivanov, S., Moore, J. K., Singh, S., Rathnayake, T., Vig, S., Granger, B. E., Muller, R. P., Bonazzi, F., Gupta, H., Vats, S., Johansson, F., Pedregosa, F., Curry, M. J., Terrel, A. R., Roučka, v., Saboo, A., Fernando, I., Kulal, S., Cimrman, R., and Scopatz, A · 2017
Earlier work this paper cites.
Imo grand challenge
Selsam, D., de Moura, L., Buzzard, K., Barton, R., Liang, P., Loos, S., and Wiedijk, F · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
A diverse corpus for evaluating and developing english math word problem solvers
Miao, S.-Y., Liang, C.-C., and Su, K.-Y · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Are NLP models really able to solve simple math word problems?
Patel, A., Bhattamishra, S., and Goyal, N · 2021
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A. J., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V. V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y., Neyshabur, B., Gur-Ari, G., and Misra, V · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
The aimo prize
AIMO · 2023
Earlier work this paper cites.
2023 amc 12a, and 12b problems
AOPS · 2023
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Earlier work this paper cites.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Rethinking benchmark and contamination for language models with rephrased samples
Yang, S., Chiang, W.-L., Zheng, L., Gonzalez, J. E., and Stoica, I · 2023
Cited alongside, same era.
2024 aime community page
AOPS · 2024
Cited alongside, same era.
Llemma: An open language model for mathematics
Azerbayev, Z., Schoelkopf, H., Paster, K., Santos, M. D., McAleer, S. M., Jiang, A. Q., Deng, J., Biderman, S., and Welleck, S · 2024
Cited alongside, same era.
Unleashing reasoning capability of llms via scalable question synthesis from scratch
Ding, Y., Shi, X., Liang, X., Li, J., Zhu, Q., and Zhang, M · 2024
Cited alongside, same era.
Ai-assisted generation of difficult math questions
Shah, V., Yu, D., Lyu, K., Park, S., Yu, J., He, Y., Ke, N. R., Mozer, M., Bengio, Y., Arora, S., et al · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., Li, Y., Wu, Y., and Guo, D · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., et al · 2024
Later among the works it cites.
Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Team, Q · 2024
Later among the works it cites.
Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Omni-math: A universal olympiad level mathematic benchmark for large language models
Gao, B., Song, F., Yang, Z., Cai, Z., Miao, Y., Dong, Q., Li, L., Ma, C., Chen, L., Xu, R., et al · 2024
Cited alongside, same era.
OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems
He, C., Luo, R., Bai, Y., Hu, S., Thai, Z., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., Liu, J., Qi, L., Liu, Z., and Sun, M · 2024
Cited alongside, same era.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Jain, N., Han, K., Gu, A., Li, W.-D., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., and Stoica, I · 2024
Cited alongside, same era.
Numinamath
Li, J., Beeching, E., Tunstall, L., Lipkin, B., Soletskyi, R., Costa Huang, S., Rasul, K., Yu, L., Jiang, A., Shen, Z., Qin, Z., Dong, B., Zhou, L., Fleureau, Y., Lample, G., and Polu, S · 2024
Cited alongside, same era.
Mathstral blog
Mistral · 2024
Cited alongside, same era.
Orca-math: Unlocking the potential of slms in grade school math
Mitra, A., Khanpour, H., Rosset, C., and Awadallah, A · 2024
Cited alongside, same era.
Tong, Y., Zhang, X., Wang, R., Wu, R., and He, J · 2024
Later among the works it cites.
Openmathinstruct-1: A 1.8 million math instruction tuning dataset
Toshniwal, S., Moshkov, I., Narenthiran, S., Gitman, D., Jia, F., and Gitman, I · 2024
Later among the works it cites.
Solving olympiad geometry without human demonstrations
Trinh, T., Wu, Y., Le, Q., He, H., and Luong, T · 2024
Later among the works it cites.
Internlm-math: Open math large language models toward verifiable reasoning
Ying, H., Zhang, S., Li, L., Zhou, Z., Shao, Y., Fei, Z., Ma, Y., Hong, J., Liu, K., Wang, Z., et al · 2024
Later among the works it cites.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., YU, J., Liu, Z., Zhang, Y., Kwok, J., Li, Z., Weller, A., and Liu, W · 2024
Later among the works it cites.
Mammoth2: Scaling instructions from the web
Yue, X., Zheng, T., Zhang, G., and Chen, W · 2024
Later among the works it cites.
Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence
Zhu, Q., Guo, D., Shao, Z., Yang, D., Wang, P., Xu, R., Wu, Y., Li, Y., Gao, H., Ma, S., et al · 2024
Later among the works it cites.
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Zhuo, T. Y., Vu, M. C., Chim, J., Hu, H., Yu, W., Widyasari, R., Yusuf, I. N. B., Zhan, H., He, J., Paul, I., et al · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.