Fetching the paper…
Reading the bibliography…
Warning: Contains harmful model outputs.
Application of computerized adaptive testing to educational problems
Weiss, D. J. and Kingsbury, G. G · 1984
Earlier work this paper cites.
Confirmatory factor analysis and item response theory: two approaches for exploring measurement invariance
Reise, S. P., Widaman, K. F., and Pugh, R. H · 1993
Earlier work this paper cites.
Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning
Messick, S · 1995
Earlier work this paper cites.
Test validity: A matter of consequence
Messick, S · 1998
Earlier work this paper cites.
Introduction to Measurement Theory
Allen, M. and Yen, W · 2001
Earlier work this paper cites.
The nature, importance, and difficulty of machine ethics
Moor, J · 2006
Earlier work this paper cites.
A suggested change in terminology and emphasis regarding validity and education
Lissitz, R. and Samuelsen, K · 2007
Earlier work this paper cites.
Therapist’s Guide to Positive Psychological Interventions
Magyar-Moe, J · 2009
Earlier work this paper cites.
A Practitioner’s Introduction to Equating with Primers on Classical Test Theory and Item Response Theory
Ryan, J. and Brockmann, F · 2009
Earlier work this paper cites.
Rehabilitation Outcome Measures
Stokes, E · 2010
Earlier work this paper cites.
Elements of Adaptive Testing
van der Linden, W. J. and Glas, C. A · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Automatic item generation: theory and practice
Gierl, M. J. and Haladyna, T. M · 2012
Earlier work this paper cites.
Using automatic item generation to create multiple-choice test items
Gierl, M. J., Lai, H., and Turner, S. R · 2012
Earlier work this paper cites.
Core psychiatry
Wright, P., Stern, J., and Phelan, M · 2012
Earlier work this paper cites.
The theory and practice of item response theory
De Ayala, R. J · 2013
Earlier work this paper cites.
Models of translation competitions
Hopkins, M. and May, J · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Semi-supervised learning with deep generative models
Kingma, D. P., Mohamed, S., Jimenez Rezende, D., and Welling, M · 2014
Earlier work this paper cites.
Encyclopedia of Quality of Life and Well-Being Research
Michalos, A. C. (ed.) · 2014
Earlier work this paper cites.
Building an evaluation scale using item response theory
Lalor, J. P., Wu, H., and Yu, H · 2016
Earlier work this paper cites.
Making sense of item response theory in machine learning
Martínez-Plumed, F., Prudêncio, R. B. C., Usó, A. M., and Hernández-Orallo, J · 2016
Earlier work this paper cites.
IRT-based aggregation model of crowdsourced pairwise comparison for evaluating machine translations
Otani, N., Nakazawa, T., Kawahara, D., and Kurohashi, S · 2016
Earlier work this paper cites.
A tutorial on fisher information
Ly, A., Marsman, M., Verhagen, J., Grasman, R. P., and Wagenmakers, E.-J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
A Review of Recent Advances in Adaptive Assessment , pp. 113–142
Vie, J.-J., Popineau, F., Bruillard, É., and Bourda, Y · 2017
Earlier work this paper cites.
Variational autoencoder for semi-supervised text classification
Xu, W., Sun, H., Deng, C., and Tan, Y · 2017
Earlier work this paper cites.
A survey on automatic detection of hate speech in text
Fortuna, P. and Nunes, S · 2018
Earlier work this paper cites.
Dynamic evaluation of neural sequence models
Krause, B., Kahembwe, E., Murray, I., and Renals, S · 2018
Earlier work this paper cites.
Understanding deep learning performance through an examination of test set difficulty: A psychometric case study
Lalor, J. P., Wu, H., Munkhdalai, T., and Yu, H · 2018
Earlier work this paper cites.
Item response theory in ai: Analysing machine learning classifiers at the instance level
Plumed, F., Prudêncio, R., Martínez-Usó, A., and Hernandez-Orallo, J · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Earlier work this paper cites.
Singularity: pattern fuzzing for worst case complexity
Wei, J., Chen, J., Feng, Y., Ferles, K., and Dillig, I · 2018
Earlier work this paper cites.
Interpretable variational autoencoders for cognitive models
Curi, M., Converse, G. A., Hajewski, J., and Oliveira, S · 2019
Earlier work this paper cites.
Openwebtext corpus
Gokaslan, A. and Cohen, V · 2019
Earlier work this paper cites.
Learning latent parameters without human response patterns: Item response theory with artificial crowds
Lalor, J. P., Wu, H., and Yu, H · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
Forbes, M., Hwang, J. D., Shwartz, V., Sap, M., and Choi, Y · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Aligning ai with shared human values
Hendrycks, D., Burns, C., Basart, S., Critch, A., Li, J., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Smart seed selection-based effective black box fuzzing for iiot protocol
Kim, S., Cho, J., Lee, C., and Shon, T · 2020
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nangia, N., Vania, C., Bhalerao, R., and Bowman, S. R · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Item response theory for efficient human evaluation of chatbots
Sedoc, J. and Ungar, L · 2020
Earlier work this paper cites.
Neural cognitive diagnosis for intelligent education systems
Wang, F., Liu, Q., Chen, E., Huang, Z., Chen, Y., Yin, Y., Huang, Z., and Wang, S · 2020
Earlier work this paper cites.
Variational item response theory: Fast, accurate, and expressive, 2020
Wu, M., Davis, R. L., Domingue, B. W., Piech, C., and Goodman, N · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Cited alongside, same era.
RedditBias: A real-world resource for bias evaluation and debiasing of conversational language models
Barikeri, S., Lauscher, A., Vulić, I., and Glavaš, G · 2021
Cited alongside, same era.
Quality meets diversity: A model-agnostic framework for computerized adaptive testing, 2021
Bi, H., Ma, H., Huang, Z., Yin, Y., Liu, Q., Chen, E., Su, Y., and Wang, S · 2021
Cited alongside, same era.
Bold: Dataset and metrics for measuring biases in open-ended language generation
Dhamala, J., Sun, T., Kumar, V., Krishna, S., Pruksachatkun, Y., Chang, K.-W., and Gupta, R · 2021
Cited alongside, same era.
Bobcat: Bilevel optimization-based computerized adaptive testing
Ghosh, A. and Lan, A · 2021
Cited alongside, same era.
Variational temporal irt: Fast, accurate, and explainable inference of dynamic learner proficiency
Kim, Y., Sankaranarayanan, S., Piech, C., and Thille, C · 2023
Later among the works it cites.
Chatgpt: Jack of all trades, master of none
Kocoń, J., Cichecki, I., Kaszyca, O., Kochanek, M., Szydło, D., Baran, J., Bielaniewicz, J., Gruza, M., Janz, A., Kanclerz, K., et al · 2023
Later among the works it cites.
Openassistant conversations-democratizing large language model alignment
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z. R., Stevens, K., Barhoum, A., Nguyen, D., Stanley, O., Nagyfi, R., et al · 2023
Later among the works it cites.
API-bank: A comprehensive benchmark for tool-augmented LLMs
Li, M., Zhao, Y., Yu, B., Song, F., Li, H., Yu, H., Li, Z., Huang, F., and Li, Y · 2023
Later among the works it cites.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deberta: Decoding-enhanced bert with disentangled attention
He, P., Liu, X., Gao, J., and Chen, W · 2021
Cited alongside, same era.
Seed selection for successful fuzzing
Herrera, A., Gunadi, H., Magrath, S., Norrish, M., Payer, M., and Hosking, A. L · 2021
Cited alongside, same era.
Dynabench: Rethinking benchmarking in NLP
Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., Ma, Z., Thrush, T., Riedel, S., Waseem, Z., Stenetorp, P., Jia, R., Bansal, M., Potts, C., and Williams, A · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Cited alongside, same era.
Towards understanding and mitigating social biases in language models
Liang, P. P., Wu, C., Morency, L.-P., and Salakhutdinov, R · 2021
Cited alongside, same era.
Dynaboard: An evaluation-as-a-service platform for holistic next-generation benchmarking
Ma, Z., Ethayarajh, K., Thrush, T., Jain, S., Wu, L., Jia, R., Potts, C., Williams, A., and Kiela, D · 2021
Cited alongside, same era.
StereoSet: Measuring stereotypical bias in pretrained language models
Nadeem, M., Bethke, A., and Reddy, S · 2021
Cited alongside, same era.
We’re afraid language models aren’t modeling ambiguity
Liu, A., Wu, Z., Michael, J., Suhr, A., West, P., Koller, A., Swayamdipta, S., Smith, N. A., and Choi, Y · 2023
Later among the works it cites.
A novel computerized adaptive testing framework with decoupled learning selector
Ma, H., Zeng, Y., Yang, S., Qin, C., Zhang, X., and Zhang, L · 2023
Later among the works it cites.
Inverse scaling: When bigger isn’t better
McKenzie, I. R., Lyzhov, A., Pieler, M., Parrish, A., Mueller, A., Prabhu, A., McLean, E., Kirtland, A., Ross, A., Liu, A., et al · 2023
Later among the works it cites.
AART: AI-assisted red-teaming with diverse data generation for new LLM-powered applications
Radharapu, B., Robinson, K., Aroyo, L., and Lahoti, P · 2023
Later among the works it cites.
What makes it ok to set a fire? iterative self-distillation of contexts and rationales for disambiguating defeasible social and moral situations
Rao, K., Jiang, L., Pyatkin, V., Gu, Y., Tandon, N., Dziri, N., Brahman, F., and Choi, Y · 2023
Later among the works it cites.
Evaluating the moral beliefs encoded in llms
Scherrer, N., Shi, C., Feder, A., and Blei, D · 2023
Later among the works it cites.
On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning
Shaikh, O., Zhang, H., Held, W., Bernstein, M., and Yang, D · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2023
Later among the works it cites.
Safety assessment of chinese large language models, 2023
Sun, H., Zhang, Z., Deng, J., Cheng, J., and Huang, M · 2023
Later among the works it cites.
Self-criticism: Aligning large language models with their understanding of helpfulness, honesty, and harmlessness
Tan, X., Shi, S., Qiu, X., Qu, C., Qi, Z., Xu, Y., and Qi, Y · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2023
Team, G · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
Evaluating evaluation metrics: A framework for analyzing NLG evaluation metrics using measurement theory
Xiao, Z., Zhang, S., Lai, V., and Liao, Q. V · 2023
Later among the works it cites.
Evaluating evaluation metrics: A framework for analyzing nlg evaluation metrics using measurement theory
Xiao, Z., Zhang, S., Lai, V., and Liao, Q. V · 2023
Later among the works it cites.
Unified detoxifying and debiasing in language generation via inference-time adaptive optimization
Yang, Z., Yi, X., Li, P., Liu, Y., and Xie, X · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts, 2023
Yu, J., Lin, X., Yu, Z., and Xing, X · 2023
Later among the works it cites.
Evaluating large language models at evaluating instruction following
Zeng, Z., Yu, J., Gao, T., Meng, Y., Goyal, T., and Chen, D · 2023
Later among the works it cites.
Agieval: A human-centric benchmark for evaluating foundation models, 2023
Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N · 2023
Later among the works it cites.
Efficiently measuring the cognitive ability of llms: An adaptive testing perspective, 2023
Zhuang, Y., Liu, Q., Ning, Y., Huang, W., Lv, R., Huang, Z., Zhao, G., Zhang, Z., Mao, Q., Wang, S., and Chen, E · 2023
Later among the works it cites.
Large language models for mathematical reasoning: Progresses and challenges
Ahn, J., Verma, R., Lou, R., Liu, D., Zhang, R., and Yin, W · 2024
Closest in time.
Measuring implicit bias in explicitly unbiased large language models
Bai, X., Wang, A., Sucholutsky, I., and Griffiths, T. L · 2024
Closest in time.
Masterkey: Automated jailbreaking of large language model chatbots
Deng, G., Liu, Y., Li, Y., Wang, K., Zhang, Y., Li, Z., Wang, H., Zhang, T., and Liu, Y · 2024
Closest in time.
Attacks, defenses and evaluations for llm conversation safety: A survey, 2024
Dong, Z., Zhou, Z., Yang, C., Shao, J., and Qiao, Y · 2024
Closest in time.
NPHardEval: Dynamic benchmark on reasoning ability of large language models via complexity classes
Fan, L., Hua, W., Li, L., Ling, H., and Zhang, Y · 2024
Closest in time.
Bias and fairness in large language models: A survey, 2024
Gallegos, I. O., Rossi, R. A., Barrow, J., Tanjim, M. M., Kim, S., Dernoncourt, F., Yu, T., Zhang, R., and Ahmed, N. K · 2024
Closest in time.
Curiosity-driven red-teaming for large language models
Hong, Z.-W., Shenfeld, I., Wang, T.-H., Chuang, Y.-S., Pareja, A., Glass, J. R., Srivastava, A., and Agrawal, P · 2024
Closest in time.
Flames: Benchmarking value alignment of LLMs in Chinese
Huang, K., Liu, X., Guo, Q., Sun, T., Sun, J., Wang, Y., Zhou, Z., Wang, Y., Teng, Y., Qiu, X., Wang, Y., and Lin, D · 2024
Closest in time.
Item response theory for natural language processing
Lalor, J. P., Rodriguez, P., Sedoc, J., and Hernandez-Orallo, J · 2024
Closest in time.
Rewardbench: Evaluating reward models for language modeling
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., et al · 2024
Closest in time.
Jailbreaking chatgpt via prompt engineering: An empirical study, 2024
Liu, Y., Deng, G., Xu, Z., Li, Y., Zheng, Y., Zhang, Y., Zhao, L., Zhang, T., Wang, K., and Liu, Y · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI · 2024
Closest in time.
tinybenchmarks: evaluating LLMs with fewer examples
Polo, F. M., Weber, L., Choshen, L., Sun, Y., Xu, G., and Yurochkin, M · 2024
Closest in time.
Safetyprompts: a systematic review of open datasets for evaluating and improving large language model safety, 2024
Röttger, P., Pernisi, F., Vidgen, B., and Hovy, D · 2024
Closest in time.
Do-not-answer: Evaluating safeguards in LLMs
Wang, Y., Li, H., Han, X., Nakov, P., and Baldwin, T · 2024
Closest in time.
Learning human-like representations to enable learning human values
Wynn, A., Sucholutsky, I., and Griffiths, T. L · 2024
Closest in time.
Unveiling the generalization power of fine-tuned large language models
Yang, H., Zhang, Y., Xu, J., Lu, H., Heng, P.-A., and Lam, W · 2024
Closest in time.
Cross-task generalization abilities of large language models
Ye, Q · 2024
Closest in time.
KoLA: Carefully benchmarking world knowledge of large language models
Yu, J., Wang, X., Tu, S., Cao, S., Zhang-Li, D., Lv, X., Peng, H., Yao, Z., Zhang, X., Li, H., et al · 2024
Closest in time.
Dyval: Graph-informed dynamic evaluation of large language models
Zhu, K., Chen, J., Wang, J., Gong, N. Z., Yang, D., and Xie, X · 2024
Closest in time.
From static benchmarks to adaptive testing: Psychometrics in ai evaluation, 2024
Zhuang, Y., Liu, Q., Ning, Y., Huang, W., Pardos, Z. A., Kyllonen, P. C., Zu, J., Mao, Q., Lv, R., Huang, Z., Zhao, G., Zhang, Z., Wang, S., and Chen, E · 2024
Closest in time.
Applying modern psychometric techniques to melodic discrimination testing: Item response theory, computerised adaptive testing, and automatic item generation
Harrison, P. M. C., Collins, T., and Müllensiefen, D · 2045
Closest in time.