Fetching the paper…
Reading the bibliography…
The versatility of large language models (LLMs) led to the creation of diverse benchmarks that thoroughly test a variety of language models' abilities.
Statistical theories of mental test scores
Lord, F., Novick, M., and Birnbaum, A · 1968
Earlier work this paper cites.
The feasibility of using item response theory as a psychometric model for the gre aptitude test
Kingston, N. M. and Dorans, N. J · 1982
Earlier work this paper cites.
Using item response theory to equate scholastic aptitude test scores
Petersen, N. S. et al · 1982
Earlier work this paper cites.
Consistency and asymptotic normality of the maximum likelihood estimator in generalized linear models
Fahrmeir, L. and Kaufmann, H · 1985
Earlier work this paper cites.
Minimal-mse linear combinations of variance estimators of the sample mean
Song, W. T · 1988
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction , volume 2
Hastie, T., Tibshirani, R., Friedman, J. H., and Friedman, J. H · 2009
Earlier work this paper cites.
Active evaluation of classifiers on large datasets
Katariya, N., Iyer, A., and Sarawagi, S · 2012
Earlier work this paper cites.
Item response theory: What it is and how you can use the irt procedure to apply it
An, X. and Yung, Y.-F · 2014
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Bojar, O., Buck, C., Federmann, C., Haddow, B., Koehn, P., Leveling, J., Monz, C., Pecina, P., Post, M., Saint-Amand, H., et al · 2014
Earlier work this paper cites.
Item response theory
Cai, L., Choi, K., Hansen, M., and Harrell, L · 2016
Earlier work this paper cites.
Building an evaluation scale using item response theory
Lalor, J. P., Wu, H., and Yu, H · 2016
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
Kočiskỳ, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Handbook of item response theory: Three volume set
Van der Linden, W. J · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Item response theory models in the measurement theory
Brzezińska, J · 2020
Earlier work this paper cites.
Active Learning for BERT: An Empirical Study
Ein-Dor, L., Halfon, A., Gera, A., Shnarch, E., Dankin, L., Choshen, L., Danilevsky, M., Aharonov, R., Katz, Y., and Slonim, N · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Cited alongside, same era.
Adversarial nli: A new benchmark for natural language understanding
Nie, Y., Williams, A., Dinan, E., Bansal, M., Weston, J., and Kiela, D · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Ground-truth labels matter: A deeper look into input-label demonstrations
Yoo, K. M., Kim, J., Kim, H. J., Cho, H., Jo, H., Lee, S.-W., Lee, S.-g., and Kim, T · 2022
Later among the works it cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Later among the works it cites.
Open llm leaderboard
Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T · 2023
Later among the works it cites.
py-irt: A scalable item response theory library for python
Lalor, J. P. and Rodriguez, P · 2023
Later among the works it cites.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Active bayesian assessment of black-box classifiers
Ji, D., Logan, R. L., Smyth, P., and Steyvers, M · 2021
Cited alongside, same era.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., and Szolovits, P · 2021
Cited alongside, same era.
Active testing: Sample-efficient model evaluation
Kossen, J., Farquhar, S., Gal, Y., and Rainforth, T · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2021
Cited alongside, same era.
Evaluation examples are not equally informative: How should that change NLP leaderboards?
Rodriguez, P., Barrow, J., Hoyle, A. M., Lalor, J. P., Jia, R., and Boyd-Graber, J · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
Comparing test sets with item response theory
Vania, C., Htut, P. M., Huang, W., Mungra, D., Pang, R. Y., Phang, J., Liu, H., Cho, K., and Bowman, S. R · 2021
Cited alongside, same era.
Liu, Z., Qiao, A., Neiswanger, W., Wang, H., Tan, B., Tao, T., Li, J., Wang, Y., Sun, S., Pangarkar, O., et al · 2023
Later among the works it cites.
Effective sample size, dimensionality, and generalization in covariate shift adaptation
Maia Polo, F. and Vicente, R · 2023
Later among the works it cites.
State of what art? a call for multi-prompt llm evaluation
Mizrahi, M., Kaplan, G., Malkin, D., Dror, R., Shahaf, D., and Stanovsky, G · 2023
Later among the works it cites.
Efficient benchmarking (of language models)
Perlitz, Y., Bandel, E., Gera, A., Arviv, O., Ein-Dor, L., Shnarch, E., Slonim, N., Shmueli-Scheuer, M., and Choshen, L · 2023
Later among the works it cites.
Sclar, M., Choi, Y., Tsvetkov, Y., and Suhr, A · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Anchor points: Benchmarking models with much fewer examples
Vivek, R., Ethayarajh, K., Yang, D., and Kiela, D · 2023
Later among the works it cites.
Larger language models do in-context learning differently
Wei, J., Wei, J., Tay, Y., Tran, D., Webson, A., Lu, Y., Chen, X., Liu, H., Huang, D., Zhou, D., et al · 2023
Later among the works it cites.
How predictable are large language model capabilities? a case study on big-bench
Ye, Q., Fu, H. Y., Ren, X., and Jia, R · 2023
Later among the works it cites.
Efficiently measuring the cognitive ability of llms: An adaptive testing perspective
Zhuang, Y., Liu, Q., Ning, Y., Huang, W., Lv, R., Huang, Z., Zhao, G., Zhang, Z., Mao, Q., Wang, S., et al · 2023
Later among the works it cites.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Guha, N., Nyarko, J., Ho, D., Ré, C., Chilton, A., Chohlas-Wood, A., Peters, A., Waldon, B., Rockmore, D., Zambrano, D., et al · 2024
Closest in time.
Mind your format: Towards consistent evaluation of in-context learning improvements
Voronov, A., Wolf, L., and Ryabinin, M · 2024
Closest in time.