Fetching the paper…
Reading the bibliography…
With LLMs shifting their role from statistical modeling of language to serving as general-purpose AI agents, how should LLM evaluations change? Arguably, a key ability of an AI agent is to flexibly combine, as needed, the basic skills it has learned.
Taxonomy of educational objectives. The classification of educational goals. Handbook 1: Cognitive domain
B. S. Bloom, M. B. Engelhart, E. J. Furst, W. H. Hill, and D. R. Krathwohl · 1956
Earlier work this paper cites.
A Taxonomy for Learning, Teaching, and Assessing. A Revision of Bloom’s Taxonomy of Educational Objectives
Lorin W. Anderson and David R. Krathwohl (eds.) · 2001
Earlier work this paper cites.
Nltk: The natural language toolkit
Edward Loper and Steven Bird · 2002
Earlier work this paper cites.
Mindreading. an integrated account of pretence, self-awareness, and understanding other’s minds
Shaun Nichols and Stephen P Stich · 2003
Earlier work this paper cites.
Pragmatics: a multidisciplinary perspective
Louise Cummings · 2005
Earlier work this paper cites.
Emerging perspectives on learning, teaching, and technology. , chapter Bloom’s taxonomy: Original and revised
Mary Forehand · 2005
Earlier work this paper cites.
Google ngram viewer
Google · 2012
Earlier work this paper cites.
Automated student model improvement
Kenneth R Koedinger, Elizabeth A McLaughlin, and John C Stamper · 2012
Earlier work this paper cites.
General and efficient cognitive model discovery using a simulated student
Nan Li, Eliane Stampfer, William Cohen, and Kenneth Koedinger · 2013
Earlier work this paper cites.
The art of reasoning. an introduction to critical and logical thinking
David Kelly · 2014
Earlier work this paper cites.
Commonsense reasoning: an event calculus based approach
Erik T Mueller · 2014
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation, September 2021
Redpajama: An open source recipe to reproduce llama training dataset, 2023
Together Computer · 2023
Closest in time.
An astonishing regularity in student learning rate
Kenneth R Koedinger, Paulo F Carvalho, Ran Liu, and Elizabeth A McLaughlin · 2023
Closest in time.
Alpacaeval: An automatic evaluator of instruction-following models
Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Estimating contamination via perplexity: Quantifying memorisation in language model evaluation, 2023
Yucheng Li · 2023
Closest in time.
Mistral 7b, 9 2023
Mistral AI Team · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2021
Cited alongside, same era.
Accelerating reinforcement learning with learned skill priors
Karl Pertsch, Youngwoon Lee, and Joseph Lim · 2021
Cited alongside, same era.
The falcon series of language models: Towards open frontier models
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Maitha Alhammadi, Mazzotta Daniele, Daniel Heslow, Julien Launay, Quentin Malartic, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo · 2023
Cited alongside, same era.
A theory for emergence of complex skills in language models
Sanjeev Arora and Anirudh Goyal · 2023
Cited alongside, same era.
Qwen technical report
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu · 2023
Cited alongside, same era.
Open llm leaderboard
Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf · 2023
Cited alongside, same era.
Skill-it! a data-driven skills framework for understanding and training language models
Mayee F Chen, Nicholas Roberts, Kush Bhatia, Jue Wang, Ce Zhang, Frederic Sala, and Christopher Ré · 2023
Cited alongside, same era.
Large language models are not zero-shot communicators
Laura Ruis, Akbir Khan, Stella Biderman, Sara Hooker, Tim Rocktäschel, and Edward Grefenstette · 2023
Closest in time.
Spread your wings: Falcon 180b is here, 9 2023
Philipp Schmid, Omar Sanseviero, Pedro Cuenca, Leandro von Werra, and Julien Launay · 2023
Closest in time.
TigerBot: A cutting-edge foundation for your very own llm
TigerResearch · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Xwin-lm, 9 2023
Xwin-LM Team · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Closest in time.