Fetching the paper…
Reading the bibliography…
This paper poses two critical issues in evaluating base models (without post-training): (1) Unstable evaluation during training: in the early stages of pre-training, the models lack the capability to answer questions as required, leading to unstable evaluation results.
A new measure of rank correlation
M. Kendall. 1938 · 1938
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Language models are few-shot learners
Nick Ryder Tom B. Brown, Benjamin Mann and et al. 2020 · 2005
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021a · 2009
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan. 2018 · 2018
Earlier work this paper cites.
Evaluating large language models trained on code
Jerry Tworek Mark Chen, Heewoo Jun, and et al. 2021 · 2021
Earlier work this paper cites.
How does gpt obtain its ability? tracing emergent abilities of language models to their sources
Hao Fu, Yao; Peng and Tushar Khot. 2022 · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022 · 2022
Earlier work this paper cites.
Language models are multilingual chain-of-thought reasoners
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2022 · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. 2022 · 2022
Earlier work this paper cites.
Evaluating large language models: A comprehensive survey
Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Supryadi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, and Deyi Xiong. 2023 · 2023
Earlier work this paper cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Earlier work this paper cites.
Cmmlu: Measuring massive multitask language understanding in chinese
Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. 2023 · 2023
Cited alongside, same era.
Opencompass: A universal evaluation platform for foundation models
OpenCompass. 2023 · 2023
Cited alongside, same era.
Abhimanyu Dubey Aaron Grattafiori and et al. Abhinav Jauhri. 2024 · 2024
Cited alongside, same era.
Rishabh Agarwal, Avi Singh, Lei M. Zhang, Bernd Bohnet, Luis Rosias, Stephanie Chan, Biao Zhang, Ankesh Anand, Zaheer Abbas, Azade Nova, John D. Co-Reyes, Eric Chu, Feryal Behbahani, Aleksandra Faust, and Hugo Larochelle. 2024 · 2024
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference
Large language models: A survey
Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024 · 2024
Later among the works it cites.
OpenAI. 2024 · 2024
Later among the works it cites.
Multi-logieval: Towards evaluating multi-step logical reasoning ability of large language models
Nisarg Patel, Mohith Kulkarni, Mihir Parmar, Aashna Budhiraja, Mutsumi Nakamura, Neeraj Varshney, and Chitta Baral. 2024 · 2024
Later among the works it cites.
The prompt report: A systematic survey of prompting techniques
Nishant Balepur Sander Schulhoff, Michael Ilie and et al. 2024 · 2024
Later among the works it cites.
Musr: Testing the limits of chain-of-thought with multistep soft reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024 · 2024
Cited alongside, same era.
DeepSeek-AI. 2024 · 2024
Cited alongside, same era.
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Tianyu Liu, Baobao Chang, Xu Sun, Lei Li, and Zhifang Sui. 2024 · 2024
Cited alongside, same era.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2024 · 2024
Cited alongside, same era.
Gemma 2: Improving open language models at a practical size
GemmaTeam. 2024 · 2024
Cited alongside, same era.
Chatglm: A family of large language models from glm-130b to glm-4 all tools
GLMTeam. 2024 · 2024
Cited alongside, same era.
Mario: Math reasoning with code interpreter output – a reproducible pipeline
Minpeng Liao, Wei Luo, Chengxi Li, Jing Wu, and Kai Fan. 2024 · 2024
Cited alongside, same era.
Kor-bench: Benchmarking language models on knowledge-orthogonal reasoning tasks
Kaijing Ma, Xinrun Du, Yunran Wang, Haoran Zhang, Zhoufutu Wen, Xingwei Qu, Jian Yang, Jiaheng Liu, Minghao Liu, Xiang Yue, Wenhao Huang, and Ge Zhang. 2024 · 2024
Cited alongside, same era.
Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett. 2024 · 2024
Later among the works it cites.
Mathscale: Scaling instruction tuning for mathematical reasoning
Zhengyang Tang, Xingxing Zhang, Benyou Wang, and Furu Wei. 2024 · 2024
Later among the works it cites.
Chain-of-thought reasoning without prompting
Xuezhi Wang and Denny Zhou. 2024 · 2024
Later among the works it cites.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al. 2024 · 2024
Later among the works it cites.
Measuring short-form factuality in large language models
Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus. 2024 · 2024
Later among the works it cites.
Every flop counts: Scaling a 300b mixture-of-experts ling llm without premium gpus
LingTeam. 2025 · 2025
Closest in time.
QwenTeam. 2025 · 2025
Closest in time.