Fetching the paper…
Reading the bibliography…
For safety, medical AI systems undergo thorough evaluations before deployment, validating their predictions against a ground truth which is assumed to be fixed and certain.
A coefficient of agreement for nominal scales
J. Cohen · 1960
Earlier work this paper cites.
The analysis of permutations
R. L. Plackett · 1975
Earlier work this paper cites.
The measurement of observer agreement for categorical data
J. R. Landis and G. G. Koch · 1977
Earlier work this paper cites.
Maximum likelihood estimation of observer error-rates using the em algorithm
A. P. Dawid and A. M. Skene · 1979
Earlier work this paper cites.
Learning from noisy examples
D. Angluin and P. D. Laird · 1987
Earlier work this paper cites.
High agreement but low kappa: I. the problems of two paradoxes
A. R. Feinstein and D. V. Cicchetti · 1990
Earlier work this paper cites.
Learning in the presence of malicious errors
M. J. Kearns and M. Li · 1993
Earlier work this paper cites.
Inferring ground truth from subjective labelling of venus images
P. Smyth, U. M. Fayyad, M. C. Burl, P. Perona, and P. Baldi · 1994
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
M. J. Kearns · 1998
Earlier work this paper cites.
A weighted kendall’s tau statistic
G. S. Shieh · 1998
Earlier work this paper cites.
Estimating a kernel fisher discriminant in the presence of label noise
N. D. Lawrence and B. Schölkopf · 2001
Earlier work this paper cites.
Comparing top k lists
R. Fagin, R. Kumar, and D. Sivakumar · 2003
Earlier work this paper cites.
Statistical methods for rates and proportions, 3rd edition
J. L. Fleiss, B. Levin, and M. C. Paik · 2003
Earlier work this paper cites.
Mm algorithms for generalized bradley-terry models
D. R. Hunter · 2003
Earlier work this paper cites.
Methods for ranking information retrieval systems without relevance judgments
S. Wu and F. Crestani · 2003
Earlier work this paper cites.
Comparing and aggregating rankings with ties
R. Fagin, R. Kumar, M. Mahdian, D. Sivakumar, and E. Vee · 2004
Earlier work this paper cites.
Learning from ambiguously labeled examples
E. Hüllermeier and J. Beringer · 2005
Earlier work this paper cites.
Rank aggregation for similar items
D. Sculley · 2007
Earlier work this paper cites.
Classification with partial labels
N. Nguyen and R. Caruana · 2008
Earlier work this paper cites.
Exploiting ’subjective’ annotations
D. Reidsma and R. op den Akker · 2008
Earlier work this paper cites.
Get another label? improving data quality and data mining using multiple, noisy labelers
V. S. Sheng, F. J. Provost, and P. G. Ipeirotis · 2008
Earlier work this paper cites.
Cheap and fast - but is it good? evaluating non-expert annotations for natural language tasks
R. Snow, B. O’Connor, D. Jurafsky, and A. Y. Ng · 2008
Earlier work this paper cites.
Utility data annotation with amazon mechanical turk
A. Sorokin and D. A. Forsyth · 2008
Earlier work this paper cites.
Truth discovery with multiple conflicting information providers on the web
X. Yin, J. Han, and P. S. Yu · 2008
Earlier work this paper cites.
Learning with annotation noise
E. Beigman and B. B. Klebanov · 2009
Earlier work this paper cites.
Integrating conflicting data: The role of source dependence
X. L. Dong, L. Berti-Équille, and D. Srivastava · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Efficient bayesian inference for generalized bradley–terry models
F. Caron and A. Doucet · 2010
Earlier work this paper cites.
Generalized distances between rankings
R. Kumar and S. Vassilvitskii · 2010
Earlier work this paper cites.
A similarity measure for indefinite rankings
W. Webber, A. Moffat, and J. Zobel · 2010
Earlier work this paper cites.
Online crowdsourcing: Rating annotators and obtaining cost-effective labels
P. Welinder and P. Perona · 2010
Earlier work this paper cites.
The multidimensional wisdom of crowds
P. Welinder, S. Branson, S. J. Belongie, and P. Perona · 2010
Earlier work this paper cites.
Learning from partial labels
T. Cour, B. Sapp, and B. Taskar · 2011
Earlier work this paper cites.
How to grade a test without knowing the answers - A bayesian graphical model for adaptive crowdsourcing and aptitude testing
Y. Bachrach, T. Graepel, T. Minka, and J. Guiver · 2012
Earlier work this paper cites.
Did it happen? the pragmatic complexity of veridicality assessment
M. de Marneffe, C. D. Manning, and C. Potts · 2012
Earlier work this paper cites.
Bayesian nonparametric plackett-luce models for the analysis of clustered ranked data
Francois, Y. W. Teh, and T. B. Murphy · 2012
Earlier work this paper cites.
Truth finding on the deep web: Is the problem solved?
X. Li, X. L. Dong, K. Lyons, W. Meng, and D. Srivastava · 2012
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
R. D. Luce · 2012
Earlier work this paper cites.
Interrater reliability: the kappa statistic
M. L. McHugh · 2012
Earlier work this paper cites.
AVA: A large-scale database for aesthetic visual analysis
N. Murray, L. Marchesotti, and F. Perronnin · 2012
Cited alongside, same era.
The problem with kappa
D. M. W. Powers · 2012
Cited alongside, same era.
On truth discovery in social sensing: a maximum likelihood estimation approach
D. Wang, L. M. Kaplan, H. K. Le, and T. F. Abdelzaher · 2012
Cited alongside, same era.
A bayesian approach to discovering truth from conflicting sources for data integration
B. Zhao, B. I. P. Rubinstein, J. Gemmell, and J. Han · 2012
Cited alongside, same era.
A consensual linear opinion pool
A. Carvalho and K. Larson · 2013
Cited alongside, same era.
Metrics, statistics, tests
T. Sakai · 2013
Cited alongside, same era.
Big transfer (bit): General visual representation learning
A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, and N. Houlsby · 2020
Later among the works it cites.
A deep learning system for differential diagnosis of skin diseases
Y. Liu, A. Jain, C. Eng, D. H. Way, K. Lee, P. Bui, K. Kanada, G. de Oliveira Marinho, J. Gallegos, S. Gabriele, V. Gupta, N. Singh, V. Natarajan, R. Hofmann-Wellenhof, G. S. Corrado, L. H. Peng, D. R. Webster, D. Ai, S. Huang, Y. Liu, R. C. Dunn, and D. Coz · 2020
Later among the works it cites.
Discrepancy ratio: Evaluating model performance when even experts disagree on the truth
I. Lovchinsky, A. Daks, I. Malkin, P. Samangouei, A. Saeedi, Y. Liu, S. Sankaranarayanan, T. Gafner, B. Sternlieb, P. Maher, and N. Silberman · 2020
Later among the works it cites.
What can we learn from collective human opinions on natural language inference data?
Y. Nie, X. Zhou, and M. Bansal · 2020
Later among the works it cites.
Human-AI Interaction in the Presence of Ambiguity: From Deliberation-based Labeling to Ambiguity-aware AI
M. Schaekermann · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Aroyo and C. Welty · 2014
Cited alongside, same era.
RAPID: rating pictorial aesthetics using deep learning
X. Lu, Z. Lin, H. Jin, J. Yang, and J. Z. Wang · 2014
Cited alongside, same era.
Sleep-spindle detection: crowdsourcing and evaluating performance of experts, non-experts and automated methods
S. C. Warby, S. L. Wendt, P. Welinder, E. G. Munk, O. Carrillo, H. B. Sorensen, P. Jennum, P. E. Peppard, P. Perona, and E. Mignot · 2014
Cited alongside, same era.
Learning from multiple annotators with varying expertise
Y. Yan, R. Rosales, G. Fung, S. Ramanathan, and J. G. Dy · 2014
Cited alongside, same era.
Truth is a lie: Crowd truth and the seven myths of human annotation
L. Aroyo and C. Welty · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li · 2015
Cited alongside, same era.
Disentangling human error from the ground truth in segmentation of medical images
L. Zhang, R. Tanno, M. Xu, C. Jin, J. Jacob, O. Ciccarelli, F. Barkhof, and D. C. Alexander · 2020
Later among the works it cites.
We need to consider disagreement in evaluation
V. Basile, M. Fell, T. Fornaciari, D. Hovy, S. Paun, B. Plank, M. Poesio, and A. Uma · 2021
Later among the works it cites.
Learning from crowds by modeling common confusions
Z. Chu, J. Ma, and H. Wang · 2021
Later among the works it cites.
Improving reference standards for validation of ai-based radiography
G. E. Duggan, J. J. Reicher, Y. Liu, D. Tse, and S. Shetty · 2021
Later among the works it cites.
A survey of race, racism, and anti-racism in NLP
A. Field, S. L. Blodgett, Z. Waseem, and Y. Tsvetkov · 2021
Later among the works it cites.
Iterative quality control strategies for expert medical image labeling
B. Freeman, N. Hammel, S. Phene, A. Huang, R. Ackermann, O. Kanzheleva, M. Hutson, C. Taggart, Q. Duong, and R. Sayres · 2021
Later among the works it cites.
The disagreement deconvolution: Bringing machine learning performance metrics in line with reality
M. L. Gordon, K. Zhou, K. Patel, T. Hashimoto, and M. S. Bernstein · 2021
Later among the works it cites.
Development and assessment of an artificial intelligence-based tool for skin condition diagnosis by primary care physicians and nurse practitioners in teledermatology practices
A. Jain, D. H. Way, V. Gupta, Y. Gao, G. de Oliveira Marinho, J. Hartford, R. Sayres, K. Kanada, C. Eng, K. Nagpal, K. Desalvo, G. S. Corrado, L. H. Peng, D. R. Webster, R. C. Dunn, D. Coz, S. J. Huang, Y. Liu, P. Bui, and Y. Liu · 2021
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
C. G. Northcutt, A. Athalye, and J. Mueller · 2021
Later among the works it cites.
Hatecheck: Functional tests for hate speech detection models
P. Röttger, B. Vidgen, D. Nguyen, Z. Waseem, H. Z. Margetts, and J. B. Pierrehumbert · 2021
Later among the works it cites.
Wise teamwork: Collective confidence calibration predicts the effectiveness of group discussion
I. Silver, B. A. Mellers, and P. E. Tetlock · 2021
Later among the works it cites.
Learning from disagreement: A survey
A. Uma, T. Fornaciari, D. Hovy, S. Paun, B. Plank, and M. Poesio · 2021
Later among the works it cites.
Robust and efficient medical imaging with self-supervision
S. Azizi, L. Culp, J. Freyberg, B. Mustafa, S. Baur, S. Kornblith, T. Chen, P. MacWilliams, S. S. Mahdavi, E. Wulczyn, B. Babenko, M. Wilson, A. Loh, P. C. Chen, Y. Liu, P. Bavishi, S. M. McKinney, J. Winkens, A. G. Roy, Z. Beaver, F. Ryan, J. Krogue, M. Etemadi, U. Telang, Y. Liu, L. Peng, G. S. Corrado, D. R. Webster, D. J. Fleet, G. E. Hinton, N. Houlsby, A. Karthikesalingam, M. Norouzi, and V. Natarajan · 2022
Later among the works it cites.
Stop measuring calibration when humans disagree
J. Baan, W. Aziz, B. Plank, and R. Fernández · 2022
Later among the works it cites.
Fine-tuning language models to find agreement among humans with diverse preferences
M. A. Bakker, M. J. Chadwick, H. Sheahan, M. H. Tessler, L. Campbell-Gillingham, J. Balaguer, N. McAleese, A. Glaese, J. Aslanides, M. M. Botvinick, and C. Summerfield · 2022
Later among the works it cites.
Measuring annotator agreement generally across complex structured, multi-object, and free-text annotation tasks
A. Braylan, O. Alonso, and M. Lease · 2022
Later among the works it cites.
Eliciting and learning with soft labels from every annotator
K. M. Collins, U. Bhatt, and A. Weller · 2022
Later among the works it cites.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
A. M. Davani, M. Díaz, and V. Prabhakaran · 2022
Later among the works it cites.
Jury learning: Integrating dissenting voices into machine learning models
M. L. Gordon, M. S. Lam, J. S. Park, K. Patel, J. T. Hancock, T. Hashimoto, and M. S. Bernstein · 2022
Later among the works it cites.
Robustness to label noise depends on the shape of the noise distribution in feature space
D. Oyen, M. Kucer, N. Hengartner, and H. S. Singh · 2022
Later among the works it cites.
The ’problem’ of human label variation: On ground truth in data, modeling and evaluation
B. Plank · 2022
Later among the works it cites.
In search of ambiguity: A three-stage workflow design to clarify annotation guidelines for crowd workers
V. K. Pradhan, M. Schaekermann, and M. Lease · 2022
Later among the works it cites.
Two contrasting data annotation paradigms for subjective NLP tasks
P. Röttger, B. Vidgen, D. Hovy, and J. B. Pierrehumbert · 2022
Later among the works it cites.
Does your dermatology classifier know what it doesn’t know? detecting the long-tail of unseen conditions
A. G. Roy, J. Ren, S. Azizi, A. Loh, V. Natarajan, B. Mustafa, N. Pawlowski, J. Freyberg, Y. Liu, Z. Beaver, N. Vo, P. Bui, S. Winter, P. MacWilliams, G. S. Corrado, U. Telang, Y. Liu, A. T. Cemgil, A. Karthikesalingam, B. Lakshminarayanan, and J. Winkens · 2022
Later among the works it cites.
Scaling and disagreements: Bias, noise, and ambiguity
A. Uma, D. Almanea, and M. Poesio · 2022
Later among the works it cites.
Pico: Contrastive label disambiguation for partial label learning
H. Wang, R. Xiao, Y. Li, L. Feng, G. Niu, G. Chen, and J. Zhao · 2022
Later among the works it cites.
Learning from label proportions by learning with label noise
J. Zhang, Y. Wang, and C. Scott · 2022
Later among the works it cites.
G. Abercrombie, V. Rieser, and D. Hovy · 2023
Closest in time.
A. Belz, C. Thomson, E. Reiter, G. Abercrombie, J. M. Alonso-Moral, M. Arvan, J. C. K. Cheung, M. Cieliebak, E. Clark, K. van Deemter, T. Dinkar, O. Dusek, S. Eger, Q. Fang, A. Gatt, D. Gkatzia, J. González-Corbelle, D. Hovy, M. Hürlimann, T. Ito, J. D. Kelleher, F. Klubicka, H. Lai, C. van der Lee, E. van Miltenburg, Y. Li, S. Mahamood, M. Mieskes, M. Nissim, N. Parde, O. Plátek, V. Rieser, P. M. Romero, J. R. Tetreault, A. Toral, X. Wan, L. Wanner, L. Watson, and D. Yang · 2023
Closest in time.
Semeval-2023 task 11: Learning with disagreements (lewidi)
E. Leonardelli, A. Uma, G. Abercrombie, D. Almanea, V. Basile, T. Fornaciari, B. Plank, V. Rieser, and M. Poesio · 2023
Closest in time.
The definition of glaucomatous optic neuropathy in artificial intelligence research and clinical applications
F. A. Medeiros, T. Lee, A. A. Jammal, L. A. Al-Aswad, M. B. Eydelman, and J. S. Schuman · 2023
Closest in time.
Why don’t you do it right? analysing annotators’ disagreement in subjective tasks
M. Sandri, E. Leonardelli, S. Tonelli, and E. Jezek · 2023
Closest in time.
ilab at semeval-2023 task 11 le-wi-di: Modelling disagreement or modelling perspectives?
N. Vitsakis, A. Parekh, T. Dinkar, G. Abercrombie, I. Konstas, and V. Rieser · 2023
Closest in time.