Fetching the paper…
Reading the bibliography…
We present a new approach to interpreting IRR that is empirical and contextualized.
The proof and measurement of association between two things
C. Spearman. 1904 · 1904
Earlier work this paper cites.
Reliability of content analysis: The case of nominal scale coding
W Scott. 1955 · 1955
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
Weighted kappa: nominal scale agreement provision for scaled disagreement or partial credit
Jacob Cohen. 1968 · 1968
Earlier work this paper cites.
Bivariate agreement coefficients for reliability of data
Klaus Krippendorff. 1970 · 1970
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
J.L. Fleiss. 1971 · 1971
Earlier work this paper cites.
Measures of response agreement for qualitative data: Some generalizations and alternatives
Richard J. Light. 1971 · 1971
Earlier work this paper cites.
The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability
Joseph L Fleiss and Jacob Cohen. 1973 · 1973
Earlier work this paper cites.
The measurement of observer agreement for categorical data
J. Richard Landis and Gary G. Koch. 1977 · 1977
Earlier work this paper cites.
A proposed solution to the base rate problem in the kappa statistic
Edward L. Spitznagel and John E. Helzer. 1985 · 1985
Earlier work this paper cites.
Diversity of decision-making models and the measurement of interrater agreement
John S Uebersax. 1987 · 1987
Earlier work this paper cites.
A generalization of cohen’s kappa agreement measure to interval measurement and multiple raters
Kenneth J. Berry and Paul W. Mielke Jr. 1988 · 1988
Earlier work this paper cites.
On agreement indices for nominal data
Roel Popping. 1988 · 1988
Earlier work this paper cites.
A partial solution to the base rate problem of the k statistic
Gavin W. Stewart and Joseph M. Rey. 1988 · 1988
Earlier work this paper cites.
High agreement but low kappa: Ii. resolving the paradoxes
Domenic V Cicchetti and Alvan R Feinstein. 1990 · 1990
Earlier work this paper cites.
High agreement but low kappa: I. the problems of two paradoxes
Alvan R Feinstein and Domenic V Cicchetti. 1990 · 1990
Earlier work this paper cites.
Bias, prevalence and kappa
Ted Byrt, Bishop Janet, and John Carlin. 1993 · 1993
Cited alongside, same era.
The correction for attenuation
Paul M Muchinsky. 1996 · 1996
Cited alongside, same era.
The reliability of a dialogue structure coding scheme
Jean Carletta, Amy Isard, Stephen Isard, Jacqueline C Kowtko, Gwyneth Doherty-Sneddon, and Anne H. Anderson. 1997 · 1997
Cited alongside, same era.
A measure of agreement for interval or nominal multivariate observations
Harald Janson and Ulf Olsson. 2001 · 2001
Cited alongside, same era.
Reference standards, judges, and comparison subjects: roles for experts in evaluating system performance
George Hripcsak and Adam Wilcox. 2002 · 2002
Cited alongside, same era.
A measure of agreement for interval or nominal multivariate observations by different sets of judges
Harald Janson and Ulf Olsson. 2004 · 2004
The anatomy of a large-scale human computation engine
Shailesh Kochhar, Stefano Mazzocchi, and Praveen Paritosh. 2010 · 2010
Later among the works it cites.
Death to kappa: birth of quantity disagreement and allocation disagreement for accuracy assessment
Robert Gilmore Pontius Jr and Marco Millones. 2011 · 2011
Later among the works it cites.
Computing inter-rater reliability for observational data: An overview and tutorial
Kevin A Hallgren. 2012 · 2012
Later among the works it cites.
Human computation must be reproducible
Praveen Paritosh. 2012 · 2012
Later among the works it cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015 · 2015
Later among the works it cites.
Rigorously collecting commonsense judgments for complex question-answer content
Mehrnoosh Sameki, Aditya Barua, and Praveen Paritosh. 2015 · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Content analysis: An introduction to its methodology
Klaus Krippendorff. 2004 · 2004
Cited alongside, same era.
Interpreting kappa in observational research: Baserate matters
Cornelia Taylor Bruckner and Paul Yoder. 2006 · 2006
Cited alongside, same era.
Using intrinsic and extrinsic metrics to evaluate accuracy and facilitation in computer-assisted coding
Philip Resnik, Michael Niv, Michael Nossal, Gregory Schnitzer, Jean Stoner, Andrew Kapit, and Richard Toren. 2006 · 2006
Cited alongside, same era.
Answering the call for a standard reliability measure for coding data
Andrew F Hayes and Klaus Krippendorff. 2007 · 2007
Cited alongside, same era.
Inter-coder agreement for computational linguistics
Ron Artstein and Massimo Poesio. 2008 · 2008
Cited alongside, same era.
Freebase: A collaboratively created graph database for structuring human knowledge
Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor. 2008 · 2008
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Later among the works it cites.
Self-report captures 27 distinct categories of emotion bridged by continuous gradients
Alan Cowen and Dacher Keltner. 2017 · 2017
Later among the works it cites.
Online controlled experiments and a/b testing
Ron Kohavi and Roger Longbotham. 2017 · 2017
Later among the works it cites.
A formal proof of a paradox associated with cohen’s kappa
Matthijs J Warrens. 2010 · 2017
Later among the works it cites.
Artificial intelligence faces reproducibility crisis
Matthew Hutson. 2018 · 2018
Later among the works it cites.
Crowdsourcing subjective tasks: The case study of understanding toxicity in online discussions
Lora Aroyo, Lucas Dixon, Nithum Thain, Olivia Redfield, and Rachel Rosen. 2019 · 2019
Later among the works it cites.
Do imagenet classifiers generalize to imagenet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. 2019 · 2019
Later among the works it cites.
Identifying statistical bias in dataset replication
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Jacob Steinhardt, and Aleksander Madry. 2020 · 2020
Later among the works it cites.
The ai bookie bet: Will machine learning outgrow human labeling?
Mike Schaekermann, Christopher Homan, Lora Aroyo, Praveen Paritosh, Kurt Bollacker, and Chris Welty. 2020 · 2020
Later among the works it cites.