Fetching the paper…
Reading the bibliography…
As Language Model (LM) capabilities advance, evaluating and supervising them at scale is getting harder for humans.
Reliability of content analysis: The case of nominal scale coding
Scott, W. A · 1955
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Cohen, J · 1960
Earlier work this paper cites.
The measurement of interrater agreement
Fleiss, J. L., Levin, B., Paik, M. C., et al · 1981
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Reliability in content analysis: Some common misconceptions and recommendations
Krippendorff, K · 2004
Earlier work this paper cites.
Do recruiters prefer applicants with similar skills? evidence from a randomized natural experiment
Bagues, M. and Perez-Villadoniga, M. J · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., WU, Y., Tamar, A., Harb, J., Pieter Abbeel, O., and Mordatch, I · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Welbl, J., Liu, N. F., and Gardner, M · 2017
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Khashabi, D., Chaturvedi, S., Roth, M., Upadhyay, S., and Roth, D · 2018
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
Huang, L., Le Bras, R., Bhagavatula, C., and Choi, Y · 2019
Earlier work this paper cites.
Similarity of neural network representations revisited
Kornblith, S., Norouzi, M., Lee, H., and Hinton, G · 2019
Earlier work this paper cites.
WiC: the word-in-context dataset for evaluating context-sensitive meaning representations
Pilehvar, M. T. and Camacho-Collados, J · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Social IQa: Commonsense reasoning about social interactions
Sap, M., Rashkin, H., Chen, D., Le Bras, R., and Choi, Y · 2019
Earlier work this paper cites.
DREAM: A challenge data set and models for dialogue-based reading comprehension
Sun, K., Yu, D., Chen, J., Yu, D., Choi, Y., and Cardie, C · 2019
Earlier work this paper cites.
QuaRTz: An open-domain dataset of qualitative relationship questions
Tafjord, O., Gardner, M., Lin, K., and Clark, P · 2019
Earlier work this paper cites.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2019
Earlier work this paper cites.
PAWS: Paraphrase adversaries from word scrambling
Zhang, Y., Baldridge, J., and He, L · 2019
Earlier work this paper cites.
“going on a vacation” takes longer than “going for a walk”: A study of temporal commonsense understanding
Zhou, B., Khashabi, D., Ning, Q., and Roth, D · 2019
Earlier work this paper cites.
Beyond accuracy: quantifying trial-by-trial behaviour of cnns and humans by measuring error consistency
Geirhos, R., Meding, K., and Wichmann, F. A · 2020
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Adversarial NLI: A new benchmark for natural language understanding
Nie, Y., Williams, A., Dinan, E., Bansal, M., Weston, J., and Kiela, D · 2020
Earlier work this paper cites.
Getting closer to ai complete question answering: A set of prerequisite real tasks
Rogers, A., Kovaleva, O., Downey, M., and Rumshisky, A · 2020
Earlier work this paper cites.
Min-mid-max scaling, limits of agreement, and agreement score, 2020
Safak, V · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Revisiting model stitching to compare neural representations
Bansal, Y., Nakkiran, P., and Barak, B · 2021
Earlier work this paper cites.
The matthews correlation coefficient (mcc) is more informative than cohen’s kappa and brier score in binary classification assessment
Chicco, D., Warrens, M. J., and Jurman, G · 2021
Earlier work this paper cites.
Partial success in closing the gap between human and machine vision
Geirhos, R., Narayanappa, K., Mitzkus, B., Thieringer, T., Bethge, M., Wichmann, F. A., and Brendel, W · 2021
Earlier work this paper cites.
Algorithmic monoculture and social welfare
Kleinberg, J. and Raghavan, M · 2021
Earlier work this paper cites.
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Kocmi, T., Federmann, C., Grundkiewicz, R., Junczys-Dowmunt, M., Matsushita, H., and Menezes, A · 2021
Earlier work this paper cites.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
Pillutla, K., Swayamdipta, S., Zellers, R., Thickstun, J., Welleck, S., Choi, Y., and Harchaoui, Z · 2021
Earlier work this paper cites.
LMdiff: A visual diff tool to compare language models
Strobelt, H., Hoover, B., Satyanaryan, A., and Gehrmann, S · 2021
Cited alongside, same era.
Measuring progress on scalable oversight for large language models, 2022
Bowman, S. R., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner, S., Lukošiūtė, K., Askell, A., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Olah, C., Amodei, D., Amodei, D., Drain, D., Li, D., Tran-Johnson, E., Kernion, J., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lovitt, L., Elhage, N., Schiefer, N., Joseph, N., Mercado, N., DasSarma, N., Larson, R., McCandlish, S., Kundu, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Telleen-Lawton, T., Brown, T., Henighan, T., Hume, T., Bai, Y., Hatfield-Dodds, Z., Mann, B., and Kaplan, J · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2024
Later among the works it cites.
Vision superalignment: Weak-to-strong generalization for vision foundation models, 2024
Guo, J., Chen, H., Wang, C., Han, K., Xu, C., and Wang, Y · 2024
Later among the works it cites.
Position: Open-endedness is essential for artificial superhuman intelligence
Hughes, E., Dennis, M. D., Parker-Holder, J., Behbahani, F., Mavalankar, A., Shi, Y., Schaul, T., and Rocktäschel, T · 2024
Later among the works it cites.
Position: the platonic representation hypothesis
Huh, M., Cheung, B., Wang, T., and Isola, P · 2024
Later among the works it cites.
Llm comparator: Visual analytics for side-by-side evaluation of large language models
Kahng, M., Tenney, I., Pushkarna, M., Liu, M. X., Wexler, J., Reif, E., Kallarackal, K., Chang, M., Terry, M., and Dixon, L · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhong, R., Snell, C., Klein, D., and Steinhardt, J · 2022
Cited alongside, same era.
Holistic evaluation of language models
Bommasani, R., Liang, P., and Lee, T · 2023
Cited alongside, same era.
Which prompts make the difference? data prioritization for efficient human llm evaluation, 2023
Boubdir, M., Kim, E., Ermis, B., Fadaee, M., and Hooker, S · 2023
Cited alongside, same era.
Rethink reporting of evaluation results in ai
Burnell, R., Schellaert, W., Burden, J., Ullman, T. D., Martinez-Plumed, F., Tenenbaum, J. B., Rutar, D., Cheke, L. G., Sohl-Dickstein, J., Mitchell, M., Kiela, D., Shanahan, M., Voorhees, E. M., Cohn, A. G., Leibo, J. Z., and Hernandez-Orallo, J · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois, Y., Li, C. X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P. S., and Hashimoto, T. B · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 2023
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2023
Cited alongside, same era.
Chatgpt outperforms crowd workers for text-annotation tasks
Gilardi, F., Alizadeh, M., and Kubli, M · 2023
Cited alongside, same era.
Kenton, Z., Siegel, N. Y., Kramár, J., Brown-Cohen, J., Albanie, S., Bulian, J., Agarwal, R., Lindner, D., Tang, Y., Goodman, N. D., and Shah, R · 2024
Later among the works it cites.
Benchmarking cognitive biases in large language models as evaluators
Koo, R., Lee, M., Raheja, V., Park, J. I., Kim, Z. M., and Kang, D · 2024
Later among the works it cites.
Let’s verify step by step
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2024
Later among the works it cites.
LLMs as narcissistic evaluators: When ego inflates evaluation scores
Liu, Y., Moosavi, N., and Lin, C · 2024
Later among the works it cites.
Aidanbench: Stress-testing language model creativity on open-ended questions
McLaughlin, A., Campbell, J., Uppuluri, A., and Yang, Y · 2024
Later among the works it cites.
Phi-4 technical report
Microsoft Research · 2024
Later among the works it cites.
Ministral 8b instruct model card
Mistral AI · 2024
Later among the works it cites.
Open-llm-leaderboard: From multi-choice to open-style questions for llms evaluation, benchmark, and arena, 2024
Myrzakhan, A., Bsharat, S. M., and Shen, Z · 2024
Later among the works it cites.
Gpt-4 technical report, 2024
OpenAI · 2024
Later among the works it cites.
LLM evaluators recognize and favor their own generations
Panickssery, A., Bowman, S. R., and Feng, S · 2024
Later among the works it cites.
Fantastic gains and where to find them: On the existence and prospect of general knowledge transfer between any pretrained model
Roth, K., Thede, L., Koepke, A. S., Vinyals, O., Hénaff, O. J., and Akata, Z · 2024
Later among the works it cites.
Experiments in weak-to-strong generalization, 2024
Scherlis, A., Mallen, A., Quirke, L., and Belrose, N · 2024
Later among the works it cites.
Welcome to the falcon 3 family of open models!
Technology Innovation Institute · 2024
Later among the works it cites.
Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges, 2024
Thakur, A. S., Choudhary, K., Ramayapally, V. S., Vaidyanathan, S., and Hupkes, D · 2024
Later among the works it cites.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Wang, Y., Ma, X., Zhang, G., Ni, Y., Chandra, A., Guo, S., Ren, W., Arulraj, A., He, X., Jiang, Z., Li, T., Ku, M., Wang, K., Zhuang, A., Fan, R., Yue, X., and Chen, W · 2024
Later among the works it cites.
Transcendence: Generative models can outperform the experts that train them
Zhang, E., Zhu, V., Saphra, N., Kleiman, A., Edelman, B. L., Tambe, M., Kakade, S., and Malach, E · 2024
Later among the works it cites.
A portfolio approach to research funding
Canton, E · 2025
Closest in time.
Vibecheck: Discover and quantify qualitative differences in large language models
Dunlap, L., Mandal, K., Darrell, T., Steinhardt, J., and Gonzalez, J. E · 2025
Closest in time.
Collapse or thrive? perils and promises of synthetic data in a self-generating world, 2025
Kazdan, J., Schaeffer, R., Dey, A., Gerstgrasser, M., Rafailov, R., Donoho, D. L., and Koyejo, S · 2025
Closest in time.
Similarity of neural network models: A survey of functional and representational measures
Klabunde, M., Schumacher, T., Strohmaier, M., and Lemmerich, F · 2025
Closest in time.
Composable interventions for language models
Kolbeinsson, A., O’Brien, K., Huang, T., Gao, S., Liu, S., Schwarz, J. R., Vaidya, A., Mahmood, F., Zitnik, M., Chen, T., and Hartvigsen, T · 2025
Closest in time.
An adversarial perspective on machine unlearning for AI safety
Łucki, J., Wei, B., Huang, Y., Henderson, P., Tramèr, F., and Rando, J · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen Team · 2025
Closest in time.
How to mitigate overfitting in weak-to-strong generalization?
Shi, J., Cheng, Q., Fei, Z., Zheng, Y., Guo, Q., and Qiu, X · 2025
Closest in time.
Mind the gap: Examining the self-improvement capabilities of large language models
Song, Y., Zhang, H., Eisenach, C., Kakade, S. M., Foster, D. P., and Ghai, U · 2025
Closest in time.
Justice or prejudice? quantifying biases in llm-as-a-judge
Ye, J., Wang, Y., Huang, Y., Chen, D., Zhang, Q., Moniz, N., Gao, T., Geyer, W., Huang, C., Chen, P., Chawla, N. V., and Zhang, X · 2025
Closest in time.
Cheating automatic LLM benchmarks: Null models achieve high win rates
Zheng, X., Pang, T., Du, C., Liu, Q., Jiang, J., and Lin, M · 2025
Closest in time.
Weak-to-strong preference optimization: Stealing reward from weak aligned model
Zhu, W., He, Z., Wang, X., Liu, P., and Wang, R · 2025
Closest in time.