Fetching the paper…
Reading the bibliography…
Most Artificial Intelligence applications are based on supervised machine learning (ML), which ultimately grounds on manually annotated data.
Fleiss, J.L.: Measuring nominal scale agreement among many raters. Psychological bulletin 76
1971
Earlier work this paper cites.
Linstone, H.A., Turoff, M.: The delphi method: Techniques and applications. In: The Delphi method: Techniques and applications. Addison-Wesley Publishing (1975)
1975
Earlier work this paper cites.
Jewett, M.A., Bombardier, C., Caron, D., Ryan, M.R., Gray, R.R., Louis, E.L.S., Witchell, S.J., Kumra, S., Psihramis, K.E.: Potential for inter-observer and intra-observer variability in x-ray review to establish stone-free rates after lithotripsy. The Journal of urology 147
1992
Earlier work this paper cites.
Lok, C.E., Morgan, C.D., Ranganathan, N.: The accuracy and interobserver agreement in detecting the ‘gallop sounds’ by cardiac auscultation. Chest 114
1998
Earlier work this paper cites.
Cicchetti, D., Bronen, R., Spencer, S., Haut, S., Berg, A., Oliver, P., Tyrer, P.: Rating scales, scales of measurement, issues of reliability: resolving some critical issues for clinicians and researchers. The Journal of nervous and mental disease 194
2006
Earlier work this paper cites.
Artstein, R., Poesio, M.: Inter-coder agreement for computational linguistics. Computational Linguistics 34
2008
Earlier work this paper cites.
Mehta, S., Granton, J., Lapinsky, S.E., Newton, G., Bandayrel, K., Little, A., Siau, C., Cook, D.J., Ayers, D., Singer, J., et al.: Agreement in electrocardiogram interpretation in patients with septic shock. Read Online: Critical Care Medicine— Society of Critical Care Medicine 39
2011
Earlier work this paper cites.
Noble, J.A.: Minority voices of crowdsourcing: Why we should pay attention to every member of the crowd. In: proceedings of the ACM 2012 conference on computer supported cooperative work companion. pp. 179–182 (2012)
2012
Earlier work this paper cites.
Chen, X., Bennett, P.N., Collins-Thompson, K., Horvitz, E.: Pairwise ranking aggregation in a crowdsourced setting. In: Proceedings of the sixth ACM international conference on Web search and data mining. pp. 193–202 (2013)
2013
Earlier work this paper cites.
Eickhoff, C., de Vries, A.P.: Increasing cheat robustness of crowdsourcing tasks. Information retrieval 16
2013
Earlier work this paper cites.
Balasubramanian, V., Ho, S.S., Vovk, V.: Conformal prediction for reliable machine learning: theory, adaptations and applications. Newnes (2014)
2014
Earlier work this paper cites.
Plank, B., Hovy, D., Søgaard, A.: Learning part-of-speech taggers with inter-annotator agreement loss. In: Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics. pp. 742–751 (2014)
2014
Earlier work this paper cites.
Aroyo, L., Welty, C.: Truth is a lie: Crowd truth and the seven myths of human annotation. AI Magazine 36
2015
Earlier work this paper cites.
Svensson, C.M., Figge, M.T., Hübler, R.: Automated classification of circulating tumor cells and the impact of interobsever variability on classifier training and performance. Journal of Immunology Research 2015
2015
Earlier work this paper cites.
Brandt, F., Conitzer, V., Endriss, U., Lang, J., Procaccia, A.D.: Handbook of computational social choice. Cambridge University Press (2016)
2016
Earlier work this paper cites.
Foreman, B., Mahulikar, A., Tadi, P., Claassen, J., Szaflarski, J., Halford, J.J., Dean, B.C., Kaplan, P.W., Hirsch, L.J., LaRoche, S., et al.: Generalized periodic discharges and ‘triphasic waves’: a blinded evaluation of inter-rater agreement and clinical significance. Clinical Neurophysiology 127
2016
Earlier work this paper cites.
Kahneman, D., Rosenfield, A., Gandhi, L., Blaser, T.: Noise: How to overcome the high, hidden cost of inconsistent decision making. Harvard Business Review 94
2016
Earlier work this paper cites.
Sharmanska, V., Hernández-Lobato, D., Miguel Hernandez-Lobato, J., Quadrianto, N.: Ambiguity helps: Classification with disagreements in crowdsourced annotations. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 2194–2202 (2016)
2016
Earlier work this paper cites.
Chang, J.C., Amershi, S., Kamar, E.: Revolt: Collaborative crowdsourcing for labeling machine learning datasets. In: Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems. pp. 2334–2346 (2017)
2017
Cited alongside, same era.
Jagabathula, S., Subramanian, L., Venkataraman, A.: Identifying unreliable and adversarial workers in crowdsourced labeling tasks. The Journal of Machine Learning Research 18
2017
Cited alongside, same era.
Vaughan, J.W.: Making better use of the crowd: How crowdsourcing can advance machine learning research. J. Mach. Learn. Res. 18
2017
Cited alongside, same era.
Barnard, M.E., Pyden, A., Rice, M.S., Linares, M., Tworoger, S.S., Howitt, B.E., Meserve, E.E., Hecht, J.L.: Inter-pathologist and pathology report agreement for ovarian tumor characteristics in the nurses’ health studies. Gynecologic oncology 150
2018
Cited alongside, same era.
Cabitza, F., Locoro, A., Alderighi, C., Rasoini, R., Compagnone, D., Berjano, P.: The elephant in the record: on the multiplicity of data recording work. Health informatics journal 25
2019
Later among the works it cites.
Heinecke, S., Reyzin, L.: Crowdsourced PAC Learning under Classification Noise. In: Proceedings of the Seventh AAAI Conference on Human Computation and Crowdsourcing. vol. 7, pp. 41–49. AAAI (2019)
2019
Later among the works it cites.
Maul, A., Mari, L., Wilson, M.: Intersubjectivity of measurement across the sciences. Measurement 131
2019
Later among the works it cites.
Peterson, J.C., Battleday, R.M., Griffiths, T.L., Russakovsky, O.: Human uncertainty makes classification more robust. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 9617–9626 (2019)
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bender, E.M., Friedman, B.: Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics 6
2018
Cited alongside, same era.
Dumitrache, A., Aroyo, L., Welty, C.: Capturing ambiguity in crowdsourcing frame disambiguation. In: Proceedings of the AAAI Conference on Human Computation and Crowdsourcing. vol. 6 (2018)
2018
Cited alongside, same era.
Eickhoff, C.: Cognitive biases in crowdsourcing. In: Proceedings of the eleventh ACM international conference on web search and data mining. pp. 162–170 (2018)
2018
Cited alongside, same era.
Guan, M., Gulshan, V., Dai, A., Hinton, G.: Who said what: Modeling individual labelers improves classification. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 32 (2018)
2018
Cited alongside, same era.
Krippendorff, K.: Content analysis: An introduction to its methodology. Sage publications (2018)
2018
Cited alongside, same era.
Li, S.Y., Jiang, Y., Chawla, N.V., Zhou, Z.H.: Multi-label learning from crowds. IEEE Transactions on Knowledge and Data Engineering 31
2018
Cited alongside, same era.
Salminen, J.O., Al-Merekhi, H.A., Dey, P., Jansen, B.J.: Inter-rater agreement for social computing studies. In: 2018 Fifth International Conference on Social Networks Analysis, Management and Security (SNAMS). pp. 80–87. IEEE (2018)
2018
Cited alongside, same era.
Schaekermann, M., Goh, J., Larson, K., Law, E.: Resolvable vs. irresolvable disagreement: A study on worker deliberation in crowd work. Proceedings of the ACM on Human-Computer Interaction 2
2018
Cited alongside, same era.
Sudre, C.H., Anson, B.G., Ingala, S., Lane, C.D., Jimenez, D., Haider, L., Varsavsky, T., Tanno, R., Smith, L., Ourselin, S., et al.: Let’s agree to disagree: Learning highly debatable multirater labelling. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. pp. 665–673. Springer (2019)
2019
Later among the works it cites.
Akhtar, S., Basile, V., Patti, V.: Modeling annotator perspective and polarized opinions to improve hate speech detection. In: Proceedings of the AAAI Conference on Human Computation and Crowdsourcing. vol. 8, pp. 151–154 (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
Cabitza, F., Campagner, A., Sconfienza, L.M.: As if sand were stone. new concepts and metrics to probe the ground on which to build trustable ai. BMC Medical Informatics and Decision Making 20
2020
Later among the works it cites.
Rizos, G., Schuller, B.W.: Average jane, where art thou?–recent avenues in efficient machine learning under subjectivity uncertainty. In: International Conference on Information Processing and Management of Uncertainty in Knowledge-Based Systems. pp. 42–55. Springer (2020)
2020
Later among the works it cites.
Tao, F., Jiang, L., Li, C.: Label similarity-based weighted soft majority voting and pairing for crowdsourcing. Knowledge and Information Systems 62
2020
Later among the works it cites.
Uma, A., Fornaciari, T., Hovy, D., Paun, S., Plank, B., Poesio, M.: A case for soft loss functions. In: Proceedings of the AAAI Conference on Human Computation and Crowdsourcing. vol. 8, pp. 173–177 (2020)
2020
Later among the works it cites.
Basile, V.: It’s the end of the gold standard as we know it. leveraging non-aggregated data for better evaluation and explanation of subjective tasks. In: Baldoni, M., Bandini, S. (eds.) AIxIA 2020: Advances in Artificial Intelligence. XIXth International Conference of the Italian Association for Artificial Intelligence, Virtual Event, November 24-27, 2020, Revised and Selected papers, vol. 12414. Springer Nature Switzerland AG (2021)
2021
Closest in time.
Cabitza, F., Campagner, A.: The need to separate the wheat from the chaff in medical informatics. International Journal of Medical Informatics 152
2021
Closest in time.
Campagner, A., Ciucci, D., Svensson, C.M., Figge, M.T., Cabitza, F.: Ground truthing from multi-rater labeling with three-way decision and possibility theory. Information Sciences 545
2021
Closest in time.
Hildebrandt, M.: The issue of bias. the framing powers of ml. In: editor, T. (ed.) Machine Learning and Society: Impact, Trust, Transparency. MIT Press (2021)
2021
Closest in time.
Kompa, B., Snoek, J., Beam, A.L.: Second opinion needed: communicating uncertainty in medical machine learning. NPJ Digital Medicine 4
2021
Closest in time.