Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are increasingly relied upon to solve complex mathematical word problems.
Language models are few-shot learners
Tom B Brown. 2020 · 2005
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021 · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021 · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Earlier work this paper cites.
Hallucinations in chatgpt: An unreliable tool for learning
Zakia Ahmad, Wahid Kaiser, and Sifatur Rahim. 2023 · 2023
Earlier work this paper cites.
Artificial hallucinations in chatgpt: implications in scientific writing
Hussam Alkaissi and Samy I McFarlane. 2023 · 2023
Earlier work this paper cites.
Role of chat gpt in public health
Som S Biswas. 2023 · 2023
Earlier work this paper cites.
Walid Hariri. 2023 · 2023
Earlier work this paper cites.
The role of chatgpt in scientific communication: writing better scientific review articles
Jingshan Huang and Ming Tan. 2023 · 2023
Earlier work this paper cites.
Mathprompter: Mathematical reasoning using large language models
Shima Imani, Liang Du, and Harsh Shrivastava. 2023 · 2023
Earlier work this paper cites.
A survey of gpt-3 family large language models including chatgpt and gpt-4
Katikapalli Subramanyam Kalyan. 2023 · 2023
Cited alongside, same era.
Better zero-shot reasoning with role-play prompting
Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, Xin Zhou, Enzhi Wang, and Xiaohang Dong. 2023 · 2023
Cited alongside, same era.
The dark side of chatgpt: Legal and ethical challenges from stochastic parrots and hallucination
Zihao Li. 2023 · 2023
Cited alongside, same era.
Ryan Lingo. 2023 · 2023
Cited alongside, same era.
Chatgpt (mar 14 version) [large language model]
OpenAI. 2023 · 2023
Cited alongside, same era.
Artifacts or abduction: How do llms answer multiple-choice questions without the question?
Nishant Balepur, Abhilasha Ravichander, and Rachel Rudinger. 2024 · 2024
Closest in time.
Prompting change: exploring prompt engineering in large language model ai and its potential to transform education
William Cain. 2024 · 2024
Closest in time.
Efficient prompting methods for large language models: A survey
Kaiyan Chang, Songcheng Xu, Chenglong Wang, Yingfeng Luo, Tong Xiao, and Jingbo Zhu. 2024 · 2024
Closest in time.
Detecting hallucinations in large language models using semantic entropy
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. 2024 · 2024
Closest in time.
Mathematical capabilities of chatgpt
Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Petersen, and Julius Berner. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the risk of misinformation pollution with large language models
Yikang Pan, Liangming Pan, Wenhu Chen, Preslav Nakov, Min-Yen Kan, and William Yang Wang. 2023 · 2023
Cited alongside, same era.
An independent evaluation of chatgpt on mathematical word problems (mwp)
Paulo Shakarian, Abhinav Koyyalamudi, Noel Ngu, and Lakshmivihari Mareedu. 2023 · 2023
Cited alongside, same era.
Chatgpt: A revolutionary tool for teaching and learning mathematics
Yousef Wardat, Mohammad A Tashtoush, Rommel AlAli, and Adeeb M Jarrah. 2023 · 2023
Cited alongside, same era.
Evaluating reading comprehension exercises generated by llms: A showcase of chatgpt in education applications
Changrong Xiao, Sean Xin Xu, Kunpeng Zhang, Yufang Wang, and Lei Xia. 2023 · 2023
Cited alongside, same era.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. 2023 · 2023
Cited alongside, same era.
Large language models for mathematical reasoning: Progresses and challenges
Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. 2024 · 2024
Cited alongside, same era.
How susceptible are llms to influence in prompts?
Sotiris Anagnostidis and Jannis Bulian. 2024 · 2024
Cited alongside, same era.
Unk-vqa: A dataset and a probe into the abstention ability of multi-modal large models
Yangyang Guo, Fangkai Jiao, Zhiqi Shen, Liqiang Nie, and Mohan Kankanhalli. 2024 · 2024
Closest in time.
Chatgpt as a math questioner? evaluating chatgpt on generating pre-university math questions
Phuoc Pham Van Long, Duc Anh Vu, Nhat M. Hoang, Xuan Long Do, and Anh Tuan Luu. 2024 · 2024
Closest in time.
Large language models are unconscious of unreasonability in math problems
Jingyuan Ma, Damai Dai, and Zhifang Sui. 2024 · 2024
Closest in time.
Openai api documentation
OpenAI. 2024 · 2024
Closest in time.
Benchmarking hallucination in large language models based on unanswerable math word problem
Yuhong Sun, Zhangyue Yin, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Hui Zhao. 2024 · 2024
Closest in time.
When to trust llms: Aligning confidence with response quality
Shuchang Tao, Liuyi Yao, Hanxing Ding, Yuexiang Xie, Qi Cao, Fei Sun, Jinyang Gao, Huawei Shen, and Bolin Ding. 2024 · 2024
Closest in time.
Mathattack: Attacking large language models towards math solving ability
Zihao Zhou, Qiufeng Wang, Mingyu Jin, Jie Yao, Jianan Ye, Wei Liu, Wei Wang, Xiaowei Huang, and Kaizhu Huang. 2024 · 2024
Closest in time.