Fetching the paper…
Reading the bibliography…
Existing benchmarks for large language models (LLMs) increasingly struggle to differentiate between top-performing models, underscoring the need for more challenging evaluation frameworks.
Taxonomy of educational objectives. Vol. 1: Cognitive domain
Benjamin S Bloom et al · 1956
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2018
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown et al · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel et al · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender et al · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi et al. Bommasani · 2021
Cited alongside, same era.
What will it take to fix benchmarking in natural language understanding?
Samuel R Bowman and George E Dahl · 2021
Cited alongside, same era.
Abstraction and analogy-making in artificial intelligence
Melanie Mitchell · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Cited alongside, same era.
Predictability and surprise in large generative models
Deep Ganguli et al · 2022
Open-ended questions in the wild: A case study of large language model evaluation
Aideen Rodriguez et al · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Mohammad Shoeybi, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo et al. Touvron · 2023
Later among the works it cites.
Solving math word problems with process-and outcome-based feedback
Jonas Wiedenhoff, Jie Chen, Nafise Sadat Moosavi, and Iryna Gurevych · 2023
Later among the works it cites.
Ifeval: Instruction following evaluation for large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Qwen technical report, 2023
Jinze Bai et al · 2023
Cited alongside, same era.
Fracgpt: Reasoning and hybrid recovery of fraction arithmetic
Zizhun Li, Zitong Yu, Qi Yu, and Zhijie Chen · 2023
Cited alongside, same era.
Zihao Yue, Tianyu Lin, Yuxiao Shen, Jialu Yang, Yujie Xu, Zhihong Cheng, and Dawn Song · 2023
Later among the works it cites.
Musr: Multi-task learning for ultra-fine entity typing and semantic role labeling
Bowei Zou and Mohit Bansal · 2023
Later among the works it cites.
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen · 2024
Closest in time.