Fetching the paper…
Reading the bibliography…
With the rapid advancement of Artificial Intelligence (AI), Large Language Models (LLMs) have significantly impacted a wide array of domains, including healthcare, engineering, science, education, and mathematical reasoning.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Gpt-3: What’s it good for?
R. Dale · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Pre-trained language models for interactive decision-making
S. Li, X. Puig, C. Paxton, Y. Du, C. Wang, L. Fan, T. Chen, D.-A. Huang, E. Akyürek, A. Anandkumar, et al · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Ad-autogpt: an autonomous gpt for alzheimer’s disease infodemiology
H. Dai, Y. Li, Z. Liu, L. Zhao, Z. Wu, S. Song, Y. Shen, D. Zhu, X. Li, S. Li, et al · 2023
Earlier work this paper cites.
Chatgpt for good? on opportunities and challenges of large language models for education
E. Kasneci, K. Seßler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, G. Groh, S. Günnemann, E. Hüllermeier, et al · 2023
Earlier work this paper cites.
Mathvista: Evaluating math reasoning in visual contexts with gpt-4v, bard, and other large multimodal models
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao · 2023
Earlier work this paper cites.
Recent advances in natural language processing via large pre-trained language models: A survey
B. Min, H. Ross, E. Sulem, A. P. B. Veyseh, T. H. Nguyen, O. Sainz, E. Agirre, I. Heintz, and D. Roth · 2023
Earlier work this paper cites.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2023
Earlier work this paper cites.
Chatabl: Abductive learning via natural language interaction with chatgpt
T. Zhong, Y. Wei, L. Yang, Z. Wu, Z. Liu, X. Wei, W. Li, J. Yao, C. Ma, X. Li, et al · 2023
Earlier work this paper cites.
A. Akella · 2024
Earlier work this paper cites.
Prompt design and engineering: Introduction and advanced methods
X. Amatriain · 2024
Earlier work this paper cites.
Deepseek llm: Scaling open-source language models with longtermism
X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu, et al · 2024
Earlier work this paper cites.
Prompting change: Exploring prompt engineering in large language model ai and its potential to transform education
W. Cain · 2024
Earlier work this paper cites.
A survey on evaluation of large language models
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, et al · 2024
Earlier work this paper cites.
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models
D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y. Wu, et al · 2024
Earlier work this paper cites.
Prompt problems: A new programming exercise for the generative ai era
P. Denny, J. Leinonen, J. Prather, A. Luxton-Reilly, T. Amarouche, B. A. Becker, and B. N. Reeves · 2024
Earlier work this paper cites.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Earlier work this paper cites.
From llm to nmt: Advancing low-resource machine translation with claude
M. Enis and M. Hopkins · 2024
Earlier work this paper cites.
A. P. Gema, J. O. J. Leang, G. Hong, A. Devoto, A. C. M. Mancino, R. Saxena, X. He, Y. Zhao, X. Du, M. R. G. Madani, et al · 2024
Earlier work this paper cites.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence
D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. Li, et al · 2024
Earlier work this paper cites.
The dawn of gui agent: A preliminary case study with claude 3.5 computer use
S. Hu, M. Ouyang, D. Gao, and M. Z. Shou · 2024
Cited alongside, same era.
A survey on evaluation of multimodal large language models
J. Huang and J. Zhang · 2024
Cited alongside, same era.
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al · 2024
Cited alongside, same era.
Google gemini as a next generation ai educational tool: a review of emerging educational technology
M. Imran and N. Almusharraf · 2024
Cited alongside, same era.
Assessing the strengths and weaknesses of large language models
S. Lappin · 2024
Cited alongside, same era.
Instance-adaptive zero-shot chain-of-thought prompting
X. Yuan, C. Shen, S. Yan, X. Zhang, L. Xie, W. Wang, R. Guan, Y. Wang, and J. Ye · 2024
Later among the works it cites.
A careful examination of large language model performance on grade school arithmetic
H. Zhang, J. Da, D. Lee, V. Robinson, C. Wu, W. Song, T. Zhao, P. Raja, C. Zhuang, D. Slack, et al · 2024
Later among the works it cites.
Evaluation of openai o1: Opportunities and challenges of agi
T. Zhong, Z. Liu, Y. Pan, Y. Zhang, Y. Zhou, S. Liang, Z. Wu, Y. Lyu, P. Shu, X. Yu, et al · 2024
Later among the works it cites.
Is your model really a good math reasoner? evaluating mathematical reasoning with checklist
Z. Zhou, S. Liu, M. Ning, W. Liu, J. Wang, D. F. Wong, X. Huang, Q. Wang, and K. Huang · 2024
Later among the works it cites.
Math-puma: Progressive upward multimodal alignment to enhance mathematical reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mathematics and machine creativity: A survey on bridging mathematics with ai
S. Liang, W. Zhang, and T. Zhong · 2024
Cited alongside, same era.
Small language models: Survey, measurements, and insights
Z. Lu, X. Li, D. Cai, R. Yi, F. Liu, X. Zhang, N. D. Lane, and M. Xu · 2024
Cited alongside, same era.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
I. Mirzadeh, K. Alizadeh, H. Shahrokhi, O. Tuzel, S. Bengio, and M. Farajtabar · 2024
Cited alongside, same era.
A. Myrzakhan, S. M. Bsharat, and Z. Shen · 2024
Cited alongside, same era.
Turning up the heat: Min-p sampling for creative and coherent llm outputs
M. Nguyen, A. Baker, C. Neo, A. Roush, A. Kirsch, and R. Shwartz-Ziv · 2024
Cited alongside, same era.
Multimath: Bridging visual and mathematical reasoning for large language models
S. Peng, D. Fu, L. Gao, X. Zhong, H. Fu, and Z. Tang · 2024
Cited alongside, same era.
A systematic survey of prompt engineering in large language models: Techniques and applications
P. Sahoo, A. K. Singh, S. Saha, V. Jain, S. Mondal, and A. Chadha · 2024
Cited alongside, same era.
W. Zhuang, X. Huang, X. Zhang, and J. Zeng · 2024
Later among the works it cites.
Llama 4: Open foundation language models
M. AI · 2025
Closest in time.
Early external safety testing of openai’s o3-mini: Insights from the pre-deployment evaluation
A. Arrieta, M. Ugarte, P. Valle, J. A. Parejo, and S. Segura · 2025
Closest in time.
Large language models and mathematical reasoning failures
J. Boye and B. Moell · 2025
Closest in time.
Auggpt: Leveraging chatgpt for text data augmentation
H. Dai, Z. Liu, W. Liao, X. Huang, Y. Cao, Z. Wu, L. Zhao, S. Xu, F. Zeng, W. Liu, et al · 2025
Closest in time.
Questioning the survey responses of large language models
R. Dominguez-Olmedo, M. Hardt, and C. Mendler-Dünner · 2025
Closest in time.
Meta-analysis and evaluation of large models’ mathematical abilities based on dikwp semantic mathematics
Y. Duan · 2025
Closest in time.
E. Evstafev · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Localized zeroth-order prompt optimization
W. Hu, Y. Shu, Z. Yu, Z. Wu, X. Lin, Z. Dai, S.-K. Ng, and B. K. H. Low · 2025
Closest in time.
deepeval, Apr. 2025
J. Ip and K. Vongthongsri · 2025
Closest in time.
Q. Jiang, Z. Gao, and G. E. Karniadakis · 2025
Closest in time.
From system 1 to system 2: A survey of reasoning large language models
Z.-Z. Li, D. Zhang, M.-L. Zhang, J. Zhang, Z. Liu, Y. Yao, H. Xu, J. Zheng, P.-J. Wang, X. Chen, et al · 2025
Closest in time.
Brief analysis of deepseek r1 and it’s implications for generative ai
S. Mercer, S. Spillard, and D. P. Martin · 2025
Closest in time.
A survey of deepseek models
F. Neha and D. Bhati · 2025
Closest in time.
Chain of thoughtlessness? an analysis of cot in planning
K. Stechly, K. Valmeekam, and S. Kambhampati · 2025
Closest in time.
Z. R. Tam, C.-K. Wu, C.-Y. Lin, and Y.-N. Chen · 2025
Closest in time.
Qwen-2.5 outperforms other large language models in the chinese national nursing licensing examination: Retrospective cross-sectional comparative study
S. Zhu, W. Hu, Z. Yang, J. Yan, and F. Zhang · 2025
Closest in time.