Fetching the paper…
Reading the bibliography…
ARC Challenge appears more difficult than ARC Easy for modern LLMs primarily due to an evaluation setup that prevents direct comparison of answer choices rather than inherent complexity.
PIQA: Reasoning about Physical Commonsense in Natural Language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019 · 1911
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown et al. 2020 · 2005
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Andrew S. Gordon, Zornitsa Kozareva, and Melissa Roemmele. 2011 · 2011
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016 · 2016
Earlier work this paper cites.
RACE: Large-scale ReAding Comprehension Dataset From Examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Earlier work this paper cites.
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Socialiqa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
HellaSwag: Can a Machine Really Finish Your Sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Multiple Choice Normalization in LM Evaluation
Leo Gao. 2021 · 2021
Cited alongside, same era.
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Later among the works it cites.
RWKV: Reinventing RNNs for the Transformer Era
Bo Peng et al. 2023 · 2023
Later among the works it cites.
Yi: Open Foundation Models by 01.AI
01. AI et al. 2024 · 2024
Closest in time.
A framework for few-shot language model evaluation
Leo Gao et al. 2024 · 2024
Closest in time.
Aaron Grattafiori et al. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural theory-of-mind? on the limits of social intelligence in large LMs
Maarten Sap, Ronan Le Bras, Daniel Fried, and Yejin Choi. 2022 · 2022
Cited alongside, same era.
Jinze Bai et al. 2023 · 2023
Cited alongside, same era.
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
DeepSeek AI et al. 2024a
Cited in the paper.
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
DeepSeek AI et al. 2024b
Cited in the paper.
Gemma 2: Improving Open Language Models at a Practical Size
Gemma Team et al. 2024a
Cited in the paper.
Gemma: Open Models Based on Gemini Research and Technology
Gemma Team et al. 2024b
Cited in the paper.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a
Cited in the paper.
Closest in time.
Albert Q. Jiang et al. 2024 · 2024
Closest in time.
Cheaper, Better, Faster, Stronger
Mistral AI. 2024 · 2024
Closest in time.