Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) vary in their abilities on a range of tasks.
Item factor analysis: Current approaches and future directions
R. J. Wirth and Michael C. Edwards · 1939
Earlier work this paper cites.
The psychology of human differences
Leona E Tyler · 1947
Earlier work this paper cites.
Factor analysis by minimizing residuals (minres)
Harry H. Harman and Wayne H. Jones · 1966
Earlier work this paper cites.
Statistical theory for logistic mental test models with a prior distribution of ability
Allan Birnbaum · 1969
Earlier work this paper cites.
A Monte Carlo investigation of logistic mental test models
Vern William Urry · 1970
Earlier work this paper cites.
Maximum Likelihood from Incomplete Data Via the EM Algorithm
A. P. Dempster, N. M. Laird, and D. B. Rubin · 1977
Earlier work this paper cites.
Introduction to factor analysis: What it is and how to do it
Jae-On Kim and Charles W Mueller · 1978
Earlier work this paper cites.
Applications of item response theory to practical testing problems
Frederic M Lord · 1980
Earlier work this paper cites.
Marginal maximum likelihood estimation of item parameters: Application of an EM algorithm
R. Darrell Bock and Murray Aitkin · 1981
Earlier work this paper cites.
Adaptive eap estimation of ability in a microcomputer environment
R Darrell Bock and Robert J Mislevy · 1982
Earlier work this paper cites.
Generalized additive models: some applications
Trevor Hastie and Robert Tibshirani · 1987
Earlier work this paper cites.
Weighted likelihood estimation of ability in item response theory
Thomas A Warm · 1989
Earlier work this paper cites.
Fundamentals of item response theory , volume 2
Ronald K Hambleton, Hariharan Swaminathan, and H Jane Rogers · 1991
Earlier work this paper cites.
Item response theory for scores on tests including polytomous items with ordered responses
David Thissen, Mary Pommerich, Kathleen Billeaud, and Valerie SL Williams · 1995
Earlier work this paper cites.
High-dimensional Full-information Item Factor Analysis
R. Darrell Bock and Steven Schilling · 1997
Earlier work this paper cites.
Theory of Point Estimation
E.L. Lehmann and George Casella · 1998
Earlier work this paper cites.
A comparison of item exposure control methods in computerized adaptive testing
Javier Revuelta and Vicente Ponsoda · 1998
Earlier work this paper cites.
Item Response Theory
Frank B. Baker and Seock-Ho Kim (eds.) · 2004
Earlier work this paper cites.
A sharing item response theory model for computerized adaptive testing
Daniel O Segall · 2004
Earlier work this paper cites.
Incorporating randomness in the fisher information for improving item-exposure control in CATs
Juan Ramón Barrada, Julio Olea, Vicente Ponsoda, and Francisco José Abad · 2008
Earlier work this paper cites.
Building predictive models in r using the caret package
Kuhn and Max · 2008
Earlier work this paper cites.
Multidimensional Item Response Theory
M.D. Reckase · 2009
Earlier work this paper cites.
A method for the comparison of item selection rules in computerized adaptive testing
Juan Ramón Barrada, Julio Olea, Vicente Ponsoda, and Francisco José Abad · 2010
Earlier work this paper cites.
Item Response Theory
Christine DeMars · 2010
Earlier work this paper cites.
Additive Gaussian Processes
David Duvenaud, Hannes Nickisch, and Carl Edward Rasmussen · 2011
Earlier work this paper cites.
Robust estimation of latent ability in item response models
Christof Schuster and Ke-Hai Yuan · 2011
Earlier work this paper cites.
mirt: A multidimensional item response theory package for the R environment
R. Philip Chalmers · 2012
Earlier work this paper cites.
Random generation of response patterns under computerized adaptive testing with the R package catR
D. Magis and G. Raiche · 2012
Earlier work this paper cites.
Regression: Models, Methods and Applications
Ludwig Fahrmeir, Thomas Kneib, Stefan Lang, and Brian Marx · 2013
Earlier work this paper cites.
Item response theory: Principles and applications
Ronald K Hambleton and Hariharan Swaminathan · 2013
Earlier work this paper cites.
Analysis of Multivariate and High-Dimensional Data
Inge Koch · 2013
Cited alongside, same era.
A note on the item information function of the four-parameter logistic model
David Magis · 2013
Cited alongside, same era.
In All Likelihood: Statistical Modelling and Inference Using Likelihood
Y. Pawitan · 2013
Cited alongside, same era.
Fundamentals of psychology
Michael Eysenck · 2014
Cited alongside, same era.
Psychometrics Behind Computerized Adaptive Testing
Hua-Hua Chang · 2015
Cited alongside, same era.
Item Response Theory
Li Cai, Kilchan Choi, Mark Hansen, and Lauren Harrell · 2016
Cited alongside, same era.
Making sense of item response theory in machine learning
Data Contamination: From Memorization to Exploitation, March 2022
Inbal Magar and Roy Schwartz · 2022
Later among the works it cites.
Probabilistic Machine Learning: An introduction
Kevin P. Murphy · 2022
Later among the works it cites.
Computation for Latent Variable Model Estimation: A Unified Stochastic Proximal Framework
Siliang Zhang and Yunxiao Chen · 2022
Later among the works it cites.
Open llm leaderboard
Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fernando Martínez-Plumed, Ricardo BC Prudêncio, Adolfo Martínez-Usó, and José Hernández-Orallo · 2016
Cited alongside, same era.
Bayesian Prior Choice in IRT Estimation Using MCMC and Variational Bayes
Prathiba Natesan, Ratna Nandakumar, Tom Minka, and Jonathan D. Rubright · 2016
Cited alongside, same era.
An application of item response theory to psychological test development
Cristian Zanon, Claudio S Hutz, Hanwook Yoo, and Ronald K Hambleton · 2016
Cited alongside, same era.
Parameter Recovery in Multidimensional Item Response Theory Models Under Complexity and Nonnormality
Dubravka Svetina, Arturo Valdivia, Stephanie Underhill, Shenghai Dai, and Xiaolin Wang · 2017
Cited alongside, same era.
Generalized Additive Models: An Introduction with R
Simon N. Wood · 2017
Cited alongside, same era.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Chelsea Schoenick, and Oyvind Tafjord · 2018
Cited alongside, same era.
John Burden, Konstantinos Voudouris, Ryan Burnell, Danaja Rutar, Lucy Cheke, and José Hernández-Orallo · 2023
Later among the works it cites.
Rethink reporting of evaluation results in ai
Ryan Burnell, Wout Schellaert, John Burden, Tomer D Ullman, Fernando Martinez-Plumed, Joshua B Tenenbaum, Danaja Rutar, Lucy G Cheke, Jascha Sohl-Dickstein, Melanie Mitchell, et al · 2023
Later among the works it cites.
Who is chatgpt? benchmarking llms’ psychological portrayal using psychobench
Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, and Michael R Lyu · 2023
Later among the works it cites.
Chatgpt for good? on opportunities and challenges of large language models for education
Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, et al · 2023
Later among the works it cites.
Pei Ke, Bosi Wen, Zhuoer Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, et al · 2023
Later among the works it cites.
Automated Annotation of Meta-Features for Predicting Language Model Performance in Natural Language Processing Tasks
Yael Moros Daval · 2023
Later among the works it cites.
Keivalya Pandya and Mehfuza Holia · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2023
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, and co authors · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Later among the works it cites.
psych: Procedures for Psychological, Psychometric, and Personality Research
William Revelle · 2023
Later among the works it cites.
How Predictable Are Large Language Model Capabilities? A Case Study on BIG-bench, October 2023
Qinyuan Ye, Harvey Yiyun Fu, Xiang Ren, and Robin Jia · 2023
Later among the works it cites.
Large language models are human-level prompt engineers
Yifan Zhang, Hongyi Sun, Amish Patel, Guangyi Liu, Chenguang Yin, and Zhe Chen · 2023
Later among the works it cites.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, December 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
Efficiently measuring the cognitive ability of llms: An adaptive testing perspective
Yan Zhuang, Qi Liu, Yuting Ning, Weizhe Huang, Rui Lv, Zhenya Huang, Guanhao Zhao, Zheng Zhang, Qingyang Mao, Shijin Wang, et al · 2023
Later among the works it cites.
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, March 2024
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica · 2024
Closest in time.
Cogbench: a large language model walks into a psychology lab
Julian Coda-Forno, Marcel Binz, Jane X Wang, and Eric Schulz · 2024
Closest in time.
A framework for few-shot language model evaluation, 07 2024
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2024
Closest in time.
Model Editing with Canonical Examples, February 2024
John Hewitt, Sarah Chen, Lanruo Lora Xie, Edward Adams, Percy Liang, and Christopher D. Manning · 2024
Closest in time.
Evaluating and inducing personality in pre-trained language models
Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han, Chi Zhang, and Yixin Zhu · 2024
Closest in time.
From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline, June 2024
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica · 2024
Closest in time.
tinyBenchmarks: Evaluating LLMs with fewer examples, February 2024
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin · 2024
Closest in time.
R: A Language and Environment for Statistical Computing
R Core Team · 2024
Closest in time.
rBayesianOptimization: Bayesian Optimization of Hyperparameters , 2024
Yachen Yan · 2024
Closest in time.
An improved stochastic EM
Siliang Zhang, Yunxiao Chen, and Yang Liu · 2044
Closest in time.