Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs), combined with program-based solving techniques, are increasingly demonstrating proficiency in mathematical reasoning.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Are NLP models really able to solve simple math word problems?
Patel, A., Bhattamishra, S., and Goyal, N · 2021
Earlier work this paper cites.
Lime: Learning inductive bias for primitives of mathematical reasoning
Wu, Y., Rabe, M. N., Li, W., Ba, J., Grosse, R. B., and Szegedy, C · 2021
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A. J., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V. V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y., Neyshabur, B., Gur-Ari, G., and Misra, V · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., brian ichter, Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Earlier work this paper cites.
Palm 2 technical report, 2023
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., and et al., E. C · 2023
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y · 2023
Earlier work this paper cites.
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks
Chen, W., Ma, X., Wang, X., and Cohen, W. W · 2023
Earlier work this paper cites.
Qlora: Efficient finetuning of quantized llms, 2023
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Earlier work this paper cites.
Faith and fate: Limits of transformers on compositionality, 2023
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., West, P., Bhagavatula, C., Bras, R. L., Hwang, J. D., Sanyal, S., Welleck, S., Ren, X., Ettinger, A., Harchaoui, Z., and Choi, Y · 2023
Cited alongside, same era.
Specializing smaller language models towards multi-step reasoning
Fu, Y., Peng, H., Ou, L., Sabharwal, A., and Khot, T · 2023
Cited alongside, same era.
PAL: Program-aided language models
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G · 2023
Cited alongside, same era.
Large language models are reasoning teachers
Ho, N., Schmid, L., and Yun, S.-Y · 2023
Cited alongside, same era.
Mistral 7b, 2023
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Cited alongside, same era.
Leveraging training data in few-shot prompting for numerical reasoning
Distilling reasoning capabilities into smaller language models
Shridhar, K., Stolfo, A., and Sachan, M · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., and et al., S. B · 2023
Closest in time.
Self-consistency improves chain of thought reasoning in language models, 2023
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2023
Closest in time.
Large language models are better reasoners with self-verification, 2023
Weng, Y., Zhu, M., Xia, F., Li, B., He, S., Liu, K., and Zhao, J · 2023
Closest in time.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J. T., Li, Z., Weller, A., and Liu, W · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jie, Z. and Lu, W · 2023
Cited alongside, same era.
Let’s verify step by step, 2023
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Cited alongside, same era.
Teaching small language models to reason
Magister, L. C., Mallinson, J., Adamek, J., Malmi, E., and Severyn, A · 2023
Cited alongside, same era.
Selfcheck: Using llms to zero-shot check their own step-by-step reasoning, 2023
Miao, N., Teh, Y. W., and Rainforth, T · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Code llama: Open foundation models for code, 2023
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C. C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G · 2023
Cited alongside, same era.
Mammoth: Building math generalist models through hybrid instruction tuning, 2023
Yue, X., Qu, X., Zhang, G., Fu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2023
Closest in time.
Automatic model selection with large language models for reasoning, 2023
Zhao, X., Xie, Y., Kawaguchi, K., He, J., and Xie, Q · 2023
Closest in time.
Progressive-hint prompting improves reasoning in large language models, 2023
Zheng, C., Liu, Z., Xie, E., Li, Z., and Li, Y · 2023
Closest in time.
Solving challenging math word problems using gpt-4 code interpreter with code-based self-verification, 2023
Zhou, A., Wang, K., Lu, Z., Shi, W., Luo, S., Qin, Z., Lu, S., Jia, A., Song, L., Zhan, M., and Li, H · 2023
Closest in time.
Pad: Program-aided distillation specializes large models in reasoning, 2023
Zhu, X., Qi, B., Zhang, K., Long, X., and Zhou, B · 2023
Closest in time.
ToRA: A tool-integrated reasoning agent for mathematical problem solving
Gou, Z., Shao, Z., Gong, Y., yelong shen, Yang, Y., Huang, M., Duan, N., and Chen, W · 2024
Closest in time.