Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated remarkable capabilities in solving complex reasoning tasks, particularly in mathematics.
Scientific Autobiography and Other Papers
Planck, M · 1949
Earlier work this paper cites.
Expert and novice performance in solving physics problems
Larkin, J. H., McDermott, J., Simon, D. P., and Simon, H. A · 1980
Earlier work this paper cites.
The Physics Problem Solver
Mendelson, E. and Zelinski, D. E · 1984
Earlier work this paper cites.
Applications of Artificial Intelligence to Physics
Klahr, P. and Waterman, D. A · 1986
Earlier work this paper cites.
A Brief History of Time: From the Big Bang to Black Holes
Hawking, S · 1988
Earlier work this paper cites.
Teaching problem solving through cooperative grouping. part 1: Group versus individual problem solving
Heller, P., Keith, R., and Anderson, S · 1992
Earlier work this paper cites.
Resource letter on physics education research
McDermott, L. C. and Redish, E. F · 1999
Earlier work this paper cites.
Physics for Scientists and Engineers
Giancoli, D. C · 2000
Earlier work this paper cites.
Teaching Physics with the Physics Suite
Redish, E. F · 2003
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Welbl, J., Liu, N. F., and Gardner, M · 2017
Earlier work this paper cites.
Phyre: A new benchmark for physical reasoning
Bakhtin, A., van der Maaten, L., Johnson, J., Gustafson, L., and Girshick, R · 2019
Earlier work this paper cites.
Reasoning about physical commonsense in natural language, 2019
Bisk, Y., Zellers, R., Le Bras, R., Gao, J., and Choi, Y · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning
Chen, J., Tang, J., Qin, J., Liang, X., Liu, L., Xing, E. P., and Lin, L · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Cao, J. and Xiao, J · 2022
Earlier work this paper cites.
Chen, W., Ma, X., Wang, X., and Cohen, W. W · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., et al · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Have llms advanced enough? a challenging problem solving benchmark for large language models
Arora, D., Singh, H. G., et al · 2023
Cited alongside, same era.
Llemma: An open language model for mathematics
Azerbayev, Z., Schoelkopf, H., Paster, K., Santos, M. D., McAleer, S., Jiang, A. Q., Deng, J., Biderman, S., and Welleck, S · 2023
Cited alongside, same era.
Theoremqa: A theorem-driven question answering dataset
Chen, W., Yin, M., Ku, M., Lu, P., Wan, Y., Ma, X., Xu, J., Wang, X., and Xia, T · 2023
Cited alongside, same era.
Using large language model to solve and explain physics word problems approaching human level
Ding, J., Cen, Y., and Wei, X · 2023
Liu, H., Zheng, Z., Qiao, Y., Duan, H., Fei, Z., Zhou, F., Zhang, W., Zhang, S., Lin, D., and Chen, K · 2024
Later among the works it cites.
Sciagent: Tool-augmented language models for scientific reasoning
Ma, Y., Gou, Z., Hao, J., Xu, R., Wang, S., Pan, L., Yang, Y., Cao, Y., Sun, A., Awadalla, H., et al · 2024
Later among the works it cites.
Mathstral
Mistral · 2024
Later among the works it cites.
The future of ai: Trends and predictions
Mistral · 2024
Later among the works it cites.
Ministral model card, 2024
MistralAI · 2024
Later among the works it cites.
Feabench: Evaluating language models on real world physics reasoning ability
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
cmmlu: Measuring massive multitask language understanding in chinese
Li, H., Zhang, Y., Koto, F., Yang, Y., Zhao, H., Gong, Y., Duan, N., and Baldwin, T · 2023
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J · 2023
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Cited alongside, same era.
Evaluating the performance of large language models on gaokao benchmark
Zhang, X., Li, C., Zong, Y., Ying, Z., He, L., and Qiu, X · 2023
Cited alongside, same era.
Agieval: A human-centric benchmark for evaluating foundation models
Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N · 2023
Cited alongside, same era.
Yi: Open foundation models by 01.ai, 2024
01-AI, :, Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., Yu, K., Liu, P., Liu, Q., Yue, S., Yang, S., Yang, S., Yu, T., Xie, W., Huang, W., Hu, X., Ren, X., Niu, X., Nie, P., Xu, Y., Liu, Y., Wang, Y., Cai, Y., Gu, Z., Liu, Z., and Dai, Z · 2024
Cited alongside, same era.
Numinamath 7b cot
Beeching, E., Huang, S. C., Jiang, A., Li, J., Lipkin, B., Qina, Z., Rasul, K., Shen, Z., Soletskyi, R., and Tunstall, L · 2024
Cited alongside, same era.
Mudur, N., Cui, H., Venugopalan, S., Raccuglia, P., Brenner, M., and Norgaard, P. C · 2024
Later among the works it cites.
Learning to reason with llms
OpenAI · 2024
Later among the works it cites.
Pang, X., Hong, R., Zhou, Z., Lv, F., Yang, X., Liang, Z., Han, B., and Zhang, C · 2024
Later among the works it cites.
Varbench: Robust language model benchmarking through dynamic variable perturbation
Qian, K., Wan, S., Tang, C., Wang, Y., Zhang, X., Chen, M., and Yu, Z · 2024
Later among the works it cites.
Qwq-32b-preview
QwQ-Team · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., Li, Y., Wu, Y., and Guo, D · 2024
Later among the works it cites.
Skywork-o1 model card, 2024
Skywork · 2024
Later among the works it cites.
Functional benchmarks for robust evaluation of reasoning performance, and the reasoning gap
Srivastava, S., PV, A., Menon, S., Sukumar, A., Philipose, A., Prince, S., Thomas, S., et al · 2024
Later among the works it cites.
Mathscale: Scaling instruction tuning for mathematical reasoning
Tang, Z., Zhang, X., Wan, B., and Wei, F · 2024
Later among the works it cites.
Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving
Tong, Y., Zhang, X., Wang, R., Wu, R., and He, J · 2024
Later among the works it cites.
Openmathinstruct-2: Accelerating ai for math with massive open-source instruction data
Toshniwal, S., Du, W., Moshkov, I., Kisacanin, B., Ayrapetyan, A., and Gitman, I · 2024
Later among the works it cites.
Livebench: A challenging, contamination-free llm benchmark
White, C., Dooley, S., Roberts, M., Pal, A., Feuer, B., Jain, S., Shwartz-Ziv, R., Jain, N., Saifullah, K., Naidu, S., et al · 2024
Later among the works it cites.
A careful examination of large language model performance on grade school arithmetic
Zhang, H., Da, J., Lee, D., Robinson, V., Wu, C., Song, W., Zhao, T., Raja, P., Slack, D., Lyu, Q., et al · 2024
Later among the works it cites.
Processbench: Identifying process errors in mathematical reasoning
Zheng, C., Zhang, Z., Zhang, B., Lin, R., Lu, K., Yu, B., Liu, D., Zhou, J., and Lin, J · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., Xue, B., Wang, B., Wu, B., Feng, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., Dai, D., Chen, D., Ji, D., Li, E., Lin, F., Dai, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Bao, H., Xu, H., Wang, H., Ding, H., Xin, H., Gao, H., Qu, H., Li, H., Guo, J., Li, J., Wang, J., Chen, J., Yuan, J., Qiu, J., Li, J., Cai, J. L., Ni, J., Liang, J., Chen, J., Dong, K., Hu, K., Gao, K., Guan, K., Huang, K., Yu, K., Wang, L., Zhang, L., Zhao, L., Wang, L., Zhang, L., Xu, L., Xia, L., Zhang, M., Zhang, M., Tang, M., Li, M., Wang, M., Li, M., Tian, N., Huang, P., Zhang, P., Wang, Q., Chen, Q., Du, Q., Ge, R., Zhang, R., Pan, R., Wang, R., Chen, R. J., Jin, R. L., Chen, R., Lu, S., Zhou, S., Chen, S., Ye, S., Wang, S., Yu, S., Zhou, S., Pan, S., Li, S. S., Zhou, S., Wu, S., Ye, S., Yun, T., Pei, T., Sun, T., Wang, T., Zeng, W., Zhao, W., Liu, W., Liang, W., Gao, W., Yu, W., Zhang, W., Xiao, W. L., An, W., Liu, X., Wang, X., Chen, X., Nie, X., Cheng, X., Liu, X., Xie, X., Liu, X., Yang, X., Li, X., Su, X., Lin, X., Li, X. Q., Jin, X., Shen, X., Chen, X., Sun, X., Wang, X., Song, X., Zhou, X., Wang, X., Shan, X., Li, Y. K., Wang, Y. Q., Wei, Y. X., Zhang, Y., Xu, Y., Li, Y., Zhao, Y., Sun, Y., Wang, Y., Yu, Y., Zhang, Y., Shi, Y., Xiong, Y., He, Y., Piao, Y., Wang, Y., Tan, Y., Ma, Y., Liu, Y., Guo, Y., Ou, Y., Wang, Y., Gong, Y., Zou, Y., He, Y., Xiong, Y., Luo, Y., You, Y., Liu, Y., Zhou, Y., Zhu, Y. X., Xu, Y., Huang, Y., Li, Y., Zheng, Y., Zhu, Y., Ma, Y., Tang, Y., Zha, Y., Yan, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Xu, Z., Xie, Z., Zhang, Z., Hao, Z., Ma, Z., Yan, Z., Wu, Z., Gu, Z., Zhu, Z., Liu, Z., Li, Z., Xie, Z., Song, Z., Pan, Z., Huang, Z., Xu, Z., Zhang, Z., and Zhang, Z · 2025
Closest in time.
Xu, X., Zhang, J., Chen, T., Chao, Z., Hu, J., and Yang, C · 2025
Closest in time.