Fetching the paper…
Reading the bibliography…
In this paper, we study the ability of large language models to learn specific mathematical rules such as distributivity or simplifying equations.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. (2018) · 2018
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 2019
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2020) · 2020
Earlier work this paper cites.
Deep learning for symbolic mathematics
Lample, G. and Charton, F. (2020) · 2020
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y., Neyshabur, B., Gur-Ari, G., and Misra, V. (2022) · 2022
Earlier work this paper cites.
Galactica: A large language model for science
Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., and Stojnic, R. (2022) · 2022
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y. (2023) · 2023
Earlier work this paper cites.
A framework for few-shot language model evaluation
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. (2023) · 2023
Earlier work this paper cites.
Mathprompter: Mathematical reasoning using large language models
Imani, S., Du, L., and Shrivastava, H. (2023) · 2023
Earlier work this paper cites.
Length generalization in arithmetic transformers
Jelassi, S., d’Ascoli, S., Domingo-Enrich, C., Wu, Y., Li, Y., and Charton, F. (2023) · 2023
Cited alongside, same era.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Luo, H., Sun, Q., Xu, C., Zhao, P., Lou, J., Tao, C., Geng, X., Lin, Q., Chen, S., and Zhang, D. (2023) · 2023
Cited alongside, same era.
Orca: Progressive Learning from Complex Explanation Traces of GPT-4
Mukherjee, S., Mitra, A., Jawahar, G., Agarwal, S., Palangi, H., and Awadallah, A. (2023) · 2023
Cited alongside, same era.
Llama 2: Open foundation and Fine-Tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. (2023) · 2023
Llemma: An open language model for mathematics
Azerbayeva, Z., Schoelkopf, H., Paster, K., Santos, M. D., McAleer, S. M., Albert Q. Jiang, J. D., Biderman, S., and Welleck, S. (2024) · 2024
Closest in time.
An empirical study of data ability boundary in llms’ math reasoning
Chen, Z., Chen, Y., Han, J., Huang, Z., Qi, J., and Zhou, Y. (2024) · 2024
Closest in time.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2024) · 2024
Closest in time.
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
Gou, Z., Shao, Z., Gong, Y., Shen, Y., Yang, Y., Huang, M., Duan, N., and Chen, W. (2024) · 2024
Closest in time.
LORA+: efficient low rank adaptation of large models
Hayou, S., Ghosh, N., and Yu, B. (2024) · 2024
Closest in time.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
Wang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R. K.-W., , and Lim, E.-P. (2023) · 2023
Cited alongside, same era.
Leandojo: Theorem proving with retrieval-augmented language models
Yang, K., Swope, A. M., Gu, A., Chalamala, R., Song, P., Yu, S., Godil, S., Prenger, R., and Anandkumar, A. (2023) · 2023
Cited alongside, same era.
Scaling relationship on learning mathematical reasoning with large language models
Yuan, Z., Yuan, H., Li, C., Dong, G., Tan, C., , and Zhou, C. (2023) · 2023
Cited alongside, same era.
Progressive-hint prompting improves reasoning in large language models
Zheng, C., Liu, Z., Xie, E., Li, Z., and Li, Y. (2023) · 2023
Cited alongside, same era.
Large language models for mathematical reasoning: Progresses and challenges
Ahn, J., Verma, R., Lou, R., Liu, D., Zhang, R., and Yin, W. (2024) · 2024
Cited alongside, same era.
Mirzadeh, I., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., and Farajtabar, M. (2024) · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., and Guo, D. (2024) · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trinh, T., Wu, Y., Le, Q., He, H., and Luong, T. (2024) · 2024
Closest in time.
AI Mathematical Olympiad - Progress Prize 1
XTX Investments (2024) · 2024
Closest in time.