Fetching the paper…
Reading the bibliography…
In LLM evaluations, reasoning is often distinguished from recall/memorization by performing numerical variations to math-oriented questions.
Understanding foundation models: Are we back in 1924?
Alan F. Smeaton. 2024 · 1924
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, and Pranav Shyam et al. 2020 · 2005
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021a · 2009
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Impact of pretraining term frequencies on few-shot numerical reasoning
Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022 · 2022
Earlier work this paper cites.
MEGA: Multilingual Evaluation of Generative AI
Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Maxamed Axmed, Kalika Bali, and Sunayana Sitaram. 2023 · 2023
Earlier work this paper cites.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang (Lorraine) Li, Liwei Jiang, Bill Yuchen Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena Hwang, Soumya Sanyal, Xiang Ren, Allyson Ettinger, Zaid Harchaoui, and Yejin Choi. 2023 · 2023
Earlier work this paper cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, and Florian others Bressand. 2023 · 2023
Earlier work this paper cites.
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2023 · 2023
Earlier work this paper cites.
NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark
Oscar Sainz, Jon Campos, Iker García-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre. 2023 · 2023
Earlier work this paper cites.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, and Lachaux et al. 2023 · 2023
Earlier work this paper cites.
On the robustness of chatgpt: An adversarial and out-of-distribution perspective
Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, Binxin Jiao, Yue Zhang, and Xing Xie. 2023 · 2023
Earlier work this paper cites.
M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models
Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. 2023 · 2023
Earlier work this paper cites.
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023 · 2023
Cited alongside, same era.
PromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Yue Zhang, Neil Zhenqiang Gong, and Xing Xie. 2023 · 2023
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, and Aleman et al. 2024 · 2024
Cited alongside, same era.
The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic. 2024 · 2024
Cited alongside, same era.
Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs
Introducing meta llama 3: The most capable openly available llm to date
Meta. 2024 · 2024
Later among the works it cites.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar. 2024 · 2024
Later among the works it cites.
Marianna Nezhurina, Lucia Cipolina-Kun, Mehdi Cherti, and Jenia Jitsev. 2024 · 2024
Later among the works it cites.
Arithmetic without algorithms: Language models solve math with a bag of heuristics
Yaniv Nikankin, Anja Reusch, Aaron Mueller, and Yonatan Belinkov. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simone Balloccu, Patrícia Schmidtová, Mateusz Lango, and Ondrej Dusek. 2024 · 2024
Cited alongside, same era.
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024 · 2024
Cited alongside, same era.
Language models show human-like content effects on reasoning tasks
Ishita Dasgupta, Andrew K. Lampinen, Stephanie C. Y. Chan, Hannah R. Sheahan, Antonia Creswell, Dharshan Kumaran, James L. McClelland, and Felix Hill. 2024 · 2024
Cited alongside, same era.
Yihong Dong, Xue Jiang, Huanyu Liu, Zhi Jin, Bin Gu, Mengfei Yang, and Ge Li. 2024 · 2024
Cited alongside, same era.
Gemma 2: Improving open language models at a practical size
Gemma-Team. 2024 · 2024
Cited alongside, same era.
Data contamination quiz: A tool to detect and estimate contamination in large language models
Shahriar Golchin and Mihai Surdeanu. 2024 · 2024
Cited alongside, same era.
Evaluating llms’ mathematical and coding competency through ontology-guided interventions
Pengfei Hong, Navonil Majumder, Deepanway Ghosal, Somak Aditya, Rada Mihalcea, and Soujanya Poria. 2024 · 2024
Cited alongside, same era.
Not all llm reasoners are created equal
Arian Hosseini, Alessandro Sordoni, Daniel Toyama, Aaron Courville, and Rishabh Agarwal. 2024 · 2024
Cited alongside, same era.
Akshara Prabhakar, Thomas L. Griffiths, and R. Thomas McCoy. 2024 · 2024
Later among the works it cites.
Vinay Samuel, Yue Zhou, and Henry Peng Zou. 2024 · 2024
Later among the works it cites.
Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap
Saurabh Srivastava, Annarose M. B, Anto P V, Shashank Menon, Ajay Sukumar, Adwaith Samod T, Alan Philipose, Stevin Prince, and Sooraj Thomas. 2024 · 2024
Later among the works it cites.
Mmlu-pro+: Evaluating higher-order reasoning and shortcut learning in llms
Saeid Asgari Taghanaki, Aliasgahr Khani, and Amir Khasahmadi. 2024 · 2024
Later among the works it cites.
Llms still can’t plan; can lrms? a preliminary evaluation of openai’s o1 on planbench
Karthik Valmeekam, Kaya Stechly, and Subbarao Kambhampati. 2024 · 2024
Later among the works it cites.
Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, and Yoon Kim. 2024 · 2024
Later among the works it cites.
Do large language models understand logic or just mimick context?
Junbing Yan, Chengyu Wang, Jun Huang, and Wei Zhang. 2024 · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeekAI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, and Runxin Xu et al. 2025 · 2025
Closest in time.
Math-perturb: Benchmarking llms’ math reasoning abilities against hard perturbations
Kaixuan Huang, Jiacheng Guo, Zihao Li, Xiang Ji, Jiawei Ge, Wenzhe Li, Yingqing Guo, Tianle Cai, Hui Yuan, Runzhe Wang, Yue Wu, Ming Yin, Shange Tang, Yangsibo Huang, Chi Jin, Xinyun Chen, Chiyuan Zhang, and Mengdi Wang. 2025 · 2025
Closest in time.
Openai o3-mini
OpenAI. 2025 · 2025
Closest in time.