Fetching the paper…
Reading the bibliography…
Most popular benchmarks for comparing LLMs rely on a limited set of prompt templates, which may not fully capture the LLMs' abilities and can affect the reproducibility of results on leaderboards.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 1908
Earlier work this paper cites.
The problem of m rankings
Maurice G Kendall and B Babington Smith · 1939
Earlier work this paper cites.
Probabilistic models for some intelligence and attainment tests
Rasch Georg · 1960
Earlier work this paper cites.
Statistical theories of mental test scores
FM Lord, MR Novick, and Allan Birnbaum · 1968
Earlier work this paper cites.
The linear logistic test model as an instrument in educational research
Gerhard H Fischer · 1973
Earlier work this paper cites.
Consistency and asymptotic normality of the maximum likelihood estimator in generalized linear models
Ludwig Fahrmeir and Heinz Kaufmann · 1985
Earlier work this paper cites.
Explanatory item response models: A generalized linear and nonlinear approach
Paul De Boeck · 2004
Earlier work this paper cites.
Development of a measure of early mathematics achievement using the rasch model: The research-based early maths assessment
Douglas H Clements, Julie H Sarama, and Xiufeng H Liu · 2008
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Item response theory
Li Cai, Kilchan Choi, Mark Hansen, and Lauren Harrell · 2016
Earlier work this paper cites.
Building an evaluation scale using item response theory
John P Lalor, Hao Wu, and Hong Yu · 2016
Earlier work this paper cites.
Effective user interface designs to increase energy-efficient behavior in a rasch-based energy recommender system
Alain Starke, Martijn Willemsen, and Chris Snijders · 2017
Earlier work this paper cites.
Handbook of item response theory: Three volume set
Wim J Van der Linden · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
A probability path
Sidney Resnick · 2019
Earlier work this paper cites.
Item response theory models in the measurement theory
Justyna Brzezińska · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih · 2020
Earlier work this paper cites.
Fixed-budget best-arm identification in structured bandits
Mohammad Javad Azizi, Branislav Kveton, and Mohammad Ghavamzadeh · 2021
Cited alongside, same era.
Evaluation examples are not equally informative: How should that change NLP leaderboards?
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, and Jordan Boyd-Graber · 2021
Cited alongside, same era.
Comparing test sets with item response theory
Clara Vania, Phu Mon Htut, William Huang, Dhara Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, and Samuel R Bowman · 2021
Cited alongside, same era.
Lmentry: A language model benchmark of elementary language tasks
Avia Efrat, Or Honovich, and Omer Levy · 2022
Cited alongside, same era.
Grips: Gradient-free, edit-based instruction search for prompting large language models
Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal · 2023
Later among the works it cites.
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Anchor points: Benchmarking models with much fewer examples
Rajan Vivek, Kawin Ethayarajh, Diyi Yang, and Douwe Kiela · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Cited alongside, same era.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al · 2022
Cited alongside, same era.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al · 2022
Cited alongside, same era.
Open llm leaderboard
Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf · 2023
Cited alongside, same era.
Statistical inference for noisy incomplete binary matrix
Yunxiao Chen, Chengcheng Li, Jing Ouyang, and Gongjun Xu · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 12 2023
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V Le, Denny Zhou, and Xinyun Chen · 2023
Later among the works it cites.
Promptbench: A unified library for evaluation of large language models
Kaijie Zhu, Qinlin Zhao, Hao Chen, Jindong Wang, and Xing Xie · 2023
Later among the works it cites.
Label-efficient model selection for text generation
Shir Ashury-Tahan, Benjamin Sznajder, Leshem Choshen, Liat Ein-Dor, Eyal Shnarch, and Ariel Gera · 2024
Closest in time.
Unitxt: Flexible, shareable and reusable data preparation and evaluation for generative ai, 2024
Elron Bandel, Yotam Perlitz, Elad Venezian, Roni Friedman-Melamed, Ofir Arviv, Matan Orbach, Shachar Don-Yehyia, Dafna Sheinwald, Ariel Gera, Leshem Choshen, Michal Shmueli-Scheuer, and Yoav Katz · 2024
Closest in time.
Navigating the modern evaluation landscape: Considerations in benchmarks and frameworks for large language models (llms)
Leshem Choshen, Ariel Gera, Yotam Perlitz, Michal Shmueli-Scheuer, and Gabriel Stanovsky · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Team Gemma, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Closest in time.
tinybenchmarks: evaluating llms with fewer examples
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
Meta · 2024
Closest in time.
Gpt-4o mini: advancing cost-efficient intelligence
OpenAI · 2024
Closest in time.
Livexiv–a multi-modal live benchmark based on arxiv papers content
Nimrod Shabtay, Felipe Maia Polo, Sivan Doveh, Wei Lin, M Jehanzeb Mirza, Leshem Chosen, Mikhail Yurochkin, Yuekai Sun, Assaf Arbelle, Leonid Karlinsky, et al · 2024
Closest in time.
Best arm identification for prompt learning under a limited budget
Chengshuai Shi, Kun Yang, Jing Yang, and Cong Shen · 2024
Closest in time.
Introducing qwen1. 5—qwenlm. github. io, 2024
Qwen Team · 2024
Closest in time.
Mind your format: Towards consistent evaluation of in-context learning improvements
Anton Voronov, Lena Wolf, and Max Ryabinin · 2024
Closest in time.