Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have been applied to Math Word Problems (MWPs) with transformative impacts, revolutionizing how these complex problems are approached and solved in various domains including educational settings.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert, 2020
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2020
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset, 2021
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Have llms advanced enough? a challenging problem solving benchmark for large language models, 2023
D. Arora, H. G. Singh, and Mausam · 2023
Earlier work this paper cites.
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks
W. Chen, X. Ma, X. Wang, and W. W. Cohen · 2023
Earlier work this paper cites.
Fill in the blank: Exploring and enhancing llm capabilities for backward reasoning in math word problems, 2023
A. Deb, N. Oza, S. Singla, D. Khandelwal, D. Garg, and P. Singla · 2023
Earlier work this paper cites.
Large language models in education: Vision and opportunities, 2023
W. Gan, Z. Qi, J. Wu, and J. C.-W. Lin · 2023
Earlier work this paper cites.
Solving math word problems by combining language models with symbolic solvers, 2023
J. He-Yueya, G. Poesia, R. E. Wang, and N. D. Goodman · 2023
Cited alongside, same era.
Mathprompter: Mathematical reasoning using large language models, 2023
S. Imani, L. Du, and H. Shrivastava · 2023
Cited alongside, same era.
ARB: Advanced Reasoning Benchmark for Large Language Models, July 2023
T. Sawada, D. Paleka, A. Havrilla, P. Tadepalli, P. Vidas, A. Kranias, J. J. Nay, K. Gupta, and A. Komatsuzaki · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Cited alongside, same era.
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback, Jan. 2024
Y. Dubois, X. Li, R. Taori, T. Zhang, I. Gulrajani, J. Ba, C. Guestrin, P. Liang, and T. B. Hashimoto · 2024
Closest in time.
Time travel in llms: Tracing data contamination in large language models, 2024
S. Golchin and M. Surdeanu · 2024
Closest in time.
Mixtral of experts, 2024
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2024
Closest in time.
Is Self-Repair a Silver Bullet for Code Generation?, Feb. 2024
T. X. Olausson, J. P. Inala, C. Wang, J. Gao, and A. Solar-Lezama · 2024
Closest in time.
What makes math word problems challenging for llms?, 2024
K. A. Srivatsa and E. Kochmar · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-consistency improves chain of thought reasoning in language models, 2023
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2023
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models, 2023
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou · 2023
Cited alongside, same era.
Scaling relationship on learning mathematical reasoning with large language models, 2023
Z. Yuan, H. Yuan, C. Li, G. Dong, K. Lu, C. Tan, C. Zhou, and J. Zhou · 2023
Cited alongside, same era.
Mammoth: Building math generalist models through hybrid instruction tuning, 2023
X. Yue, X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su, and W. Chen · 2023
Cited alongside, same era.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Dec. 2023
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Cited alongside, same era.
URL https://openai.com/index/hello-gpt-4o/
Hello GPT-4o, a
Cited in the paper.
URL https://www.anthropic.com/news/claude-3-family
Introducing the next generation of Claude \ Anthropic, b
Cited in the paper.
URL https://openai.com/index/chatgpt/
Introducing ChatGPT, c
Cited in the paper.
Closest in time.
Evaluating mathematical reasoning beyond accuracy, 2024
S. Xia, X. Li, Y. Liu, T. Wu, and P. Liu · 2024
Closest in time.
Can llms solve longer math word problems better?, 2024
X. Xu, T. Xiao, Z. Chao, Z. Huang, C. Yang, and Y. Wang · 2024
Closest in time.
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models, May 2024
L. Yu, W. Jiang, H. Shi, J. Yu, Z. Liu, Y. Zhang, J. T. Kwok, Z. Li, A. Weller, and W. Liu · 2024
Closest in time.