Fetching the paper…
Reading the bibliography…
The rapid entry of machine learning approaches in our daily activities and high-stakes domains demands transparency and scrutiny of their fairness and reliability.
Metrology for AI: From Benchmarks to Instruments
Welty, C.; Paritosh, P.; and Aroyo, L. 2019 · 1911
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Cohen, J. 1960 · 1960
Earlier work this paper cites.
Crowdsourcing stereotypes: Linguistic bias in metadata generated via gwap
Otterbacher, J. 2015 · 1964
Earlier work this paper cites.
Indices of Qualitative Variation
Wilcox, A. R. 1967 · 1967
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
Fleiss, J. L. 1971 · 1971
Earlier work this paper cites.
The measurement of observer agreement for categorical data
Landis, J. R.; and Koch, G. G. 1977 · 1977
Earlier work this paper cites.
Validity in content analysis
Krippendorff, K. 1980 · 1980
Earlier work this paper cites.
Contextual correlates of semantic similarity
Miller, G. A.; and Charles, W. G. 1991 · 1991
Earlier work this paper cites.
Dependence of weighted kappa coefficients on the number of categories
Brenner, H.; and Kliebsch, U. 1996 · 1996
Earlier work this paper cites.
Assessing agreement on classification tasks: the kappa statistic
Carletta, J. 1996 · 1996
Earlier work this paper cites.
Beyond accuracy: What data quality means to data consumers
Wang, R. Y.; and Strong, D. M. 1996 · 1996
Earlier work this paper cites.
Placing search in context: The concept revisited
Finkelstein, L.; Gabrilovich, E.; Matias, Y.; Rivlin, E.; Solan, Z.; Wolfman, G.; and Ruppin, E. 2001 · 2001
Earlier work this paper cites.
Hybrid recommender systems: Survey and experiments
Burke, R. 2002 · 2002
Earlier work this paper cites.
Measuring agreement on set-valued items (MASI) for semantic and pragmatic annotation
Passonneau, R. 2006 · 2006
Earlier work this paper cites.
Inter-coder agreement for computational linguistics
Artstein, R.; and Poesio, M. 2008 · 2008
Earlier work this paper cites.
ISO/IEC 25012: Software Engineering: Software Product Quality Requirements and Evaluation (SQuaRE): Data Quality Model
for Standardization, I. O. 2008 · 2008
Earlier work this paper cites.
Crowdsourcing User Studies with Mechanical Turk
Kittur, A.; Chi, E. H.; and Suh, B. 2008 · 2008
Earlier work this paper cites.
Cheap and fast–but is it good? evaluating non-expert annotations for natural language tasks
Snow, R.; O’connor, B.; Jurafsky, D.; and Ng, A. Y. 2008 · 2008
Earlier work this paper cites.
An effective, low-cost measure of semantic relatedness obtained from Wikipedia links
Witten, I. H.; and Milne, D. N. 2008 · 2008
Earlier work this paper cites.
Sentiment Analysis in the News
Balahur, A.; Steinberger, R.; Kabadjov, M. A.; Zavarella, V.; der Goot, E. V.; Halkia, M.; Pouliquen, B.; and Belyaeva, J. 2010 · 2010
Earlier work this paper cites.
Quality management on amazon mechanical turk
Ipeirotis, P. G.; Provost, F.; and Wang, J. 2010 · 2010
Earlier work this paper cites.
How reliable are annotations via crowdsourcing: a study about inter-annotator agreement for multi-label image annotation
Nowak, S.; and Rüger, S. 2010 · 2010
Earlier work this paper cites.
Online crowdsourcing: rating annotators and obtaining cost-effective labels
Welinder, P.; and Perona, P. 2010 · 2010
Earlier work this paper cites.
Repeatable and reliable search system evaluation using crowdsourcing
Blanco, R.; Halpin, H.; Herzig, D. M.; Mika, P.; Pound, J.; Thompson, H. S.; and Tran Duc, T. 2011 · 2011
Earlier work this paper cites.
Computing Krippendorff’s alpha-reliability
Krippendorff, K. 2011 · 2011
Earlier work this paper cites.
War versus inspirational in forrest gump: Cultural effects in tagging communities
Dong, Z.; Shi, C.; Sen, S.; Terveen, L.; and Riedl, J. 2012 · 2012
Earlier work this paper cites.
Crowdsourcing micro-level multimedia annotations: The challenges of evaluation and interface
Park, S.; Mohammadi, G.; Artstein, R.; and Morency, L.-P. 2012 · 2012
Earlier work this paper cites.
Reactive crowdsourcing
Bozzon, A.; Brambilla, M.; Ceri, S.; and Mauri, A. 2013 · 2013
Earlier work this paper cites.
Learning whom to trust with MACE
Hovy, D.; Berg-Kirkpatrick, T.; Vaswani, A.; and Hovy, E. 2013 · 2013
Earlier work this paper cites.
An evaluation of aggregation techniques in crowdsourcing
Hung, N. Q. V.; Tam, N. T.; Tran, L. N.; and Aberer, K. 2013 · 2013
Earlier work this paper cites.
The future of crowd work
Kittur, A.; Nickerson, J. V.; Bernstein, M.; Gerber, E.; Shaw, A.; Zimmerman, J.; Lease, M.; and Horton, J. 2013 · 2013
Earlier work this paper cites.
Measuring crowd truth: Disagreement metrics combined with worker behavior filters
Soberón, G.; Aroyo, L.; Welty, C.; Inel, O.; Lin, H.; and Overmeen, M. 2013 · 2013
Earlier work this paper cites.
The Three Sides of CrowdTruth
Aroyo, L.; and Welty, C. 2014 · 2014
Earlier work this paper cites.
Measuring gradience in speakers’ grammaticality judgements
Lau, J. H.; Clark, A.; and Lappin, S. 2014 · 2014
Earlier work this paper cites.
Neural word embedding as implicit matrix factorization
Levy, O.; and Goldberg, Y. 2014 · 2014
Earlier work this paper cites.
Toward crowdsourcing micro-level behavior annotations: the challenges of interface, training, and generalization
Park, S.; Shoemark, P.; and Morency, L.-P. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Cited alongside, same era.
Towards interactive, intelligent, and integrated multimedia analytics
Zahálka, J.; and Worring, M. 2014 · 2014
Cited alongside, same era.
Truth is a lie: Crowd truth and the seven myths of human annotation
Aroyo, L.; and Welty, C. 2015 · 2015
Cited alongside, same era.
Incentivizing high quality crowdwork
Ho, C.-J.; Slivkins, A.; Suri, S.; and Vaughan, J. W. 2015 · 2015
Cited alongside, same era.
Information quality research challenge: adapting information quality principles to user-generated content
Lukyanenko, R.; and Parsons, J. 2015 · 2015
Cited alongside, same era.
Turkers, scholars,” arafat” and” peace” cultural communities and algorithmic gold standards
Sen, S.; Giesel, M. E.; Gold, R.; Hillmann, B.; Lesicko, M.; Naden, S.; Russell, J.; Wang, Z.; and Hecht, B. 2015 · 2015
The global landscape of AI ethics guidelines
Jobin, A.; Ienca, M.; and Vayena, E. 2019 · 2019
Later among the works it cites.
Exploiting worker correlation for label aggregation in crowdsourcing
Li, Y.; Rubinstein, B.; and Cohn, T. 2019 · 2019
Later among the works it cites.
Model cards for model reporting
Mitchell, M.; Wu, S.; Zaldivar, A.; Barnes, P.; Vasserman, L.; Hutchinson, B.; Spitzer, E.; Raji, I. D.; and Gebru, T. 2019 · 2019
Later among the works it cites.
Platform-related factors in repeatability and reproducibility of crowdsourcing tasks
Qarout, R.; Checco, A.; Demartini, G.; and Bontcheva, K. 2019 · 2019
Later among the works it cites.
Machine learning in mental health: a scoping review of methods and applications
Shatte, A. B.; Hutchinson, D. M.; and Teague, S. J. 2019 · 2019
Later among the works it cites.
Fairlearn: A toolkit for assessing and improving fairness in AI
Bird, S.; Dudík, M.; Edgar, R.; Horn, B.; Lutz, R.; Milan, V.; Sameki, M.; Wallach, H.; and Walker, K. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Twitter as a Lifeline: Human-annotated Twitter Corpora for NLP of Crisis-related Messages
Imran, M.; Mitra, P.; and Castillo, C. 2016 · 2016
Cited alongside, same era.
Parting crowds: Characterizing divergent interpretations in crowdsourced annotation tasks
Kairam, S.; and Heer, J. 2016 · 2016
Cited alongside, same era.
Much Ado About Time: Exhaustive Annotation of Temporal Data
Sigurdsson, G. A.; Russakovsky, O.; Farhadi, A.; Laptev, I.; and Gupta, A. 2016 · 2016
Cited alongside, same era.
The FAIR Guiding Principles for scientific data management and stewardship
Wilkinson, M. D.; Dumontier, M.; Aalbersberg, I. J.; Appleton, G.; Axton, M.; Baak, A.; Blomberg, N.; Boiten, J.-W.; da Silva Santos, L. B.; Bourne, P. E.; et al. 2016 · 2016
Cited alongside, same era.
Using Spearman’s correlation coefficients for exploratory data analysis on big dataset
Xiao, C.; Ye, J.; Esteves, R. M.; and Rong, C. 2016 · 2016
Cited alongside, same era.
Quality assessment for linked data: A survey
Zaveri, A.; Rula, A.; Maurino, A.; Pietrobon, R.; Lehmann, J.; and Auer, S. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Modeling and aggregation of complex annotations via annotation distances
Braylan, A.; and Lease, M. 2020 · 2020
Later among the works it cites.
Garbage in, garbage out? do machine learning application papers in social computing report where human-labeled training data comes from?
Geiger, R. S.; Yu, K.; Yang, Y.; Dai, M.; Qiu, J.; Tang, R.; and Huang, J. 2020 · 2020
Later among the works it cites.
Eliciting User Preferences for Personalized Explanations for Video Summaries
Inel, O.; Tintarev, N.; and Aroyo, L. 2020 · 2020
Later among the works it cites.
Data desiderata: Reliability and fidelity in high-stakes AI
Kapania, S.; Sambasivan, N.; Olson, K.; Highfill, H.; Akrong, D.; Paritosh, P.; and Aroyo, L. 2020 · 2020
Later among the works it cites.
Annotator rationales for labeling tasks in crowdsourcing
Kutlu, M.; McDonnell, T.; Lease, M.; and Elsayed, T. 2020 · 2020
Later among the works it cites.
Between Subjectivity and Imposition: Power Dynamics in Data Annotation for Computer Vision
Miceli, M.; Schuessler, M.; and Yang, T. 2020 · 2020
Later among the works it cites.
DREC: towards a Datasheet for Reporting Experiments in Crowdsourcing
Ramírez, J.; Baez, M.; Casati, F.; Cernuzzi, L.; and Benatallah, B. 2020 · 2020
Later among the works it cites.
Studying the effects of cognitive biases in evaluation of conversational agents
Santhanam, S.; Karduni, A.; and Shaikh, S. 2020 · 2020
Later among the works it cites.
Toward a perspectivist turn in ground truthing for predictive computing
Basile, V.; Cabitza, F.; Campagner, A.; and Fell, M. 2021 · 2021
Later among the works it cites.
It’s About Time: A View of Crowdsourced Data Before and During the Pandemic
Christoforou, E.; Barlas, P.; and Otterbacher, J. 2021 · 2021
Later among the works it cites.
A checklist to combat cognitive biases in crowdsourcing
Draws, T.; Rieger, A.; Inel, O.; Gadiraju, U.; and Tintarev, N. 2021 · 2021
Later among the works it cites.
Empirical methodology for crowdsourcing ground truth
Dumitrache, A.; Inel, O.; Timmermans, B.; Ortiz, C.; Sips, R.-J.; Aroyo, L.; and Welty, C. 2021 · 2021
Later among the works it cites.
Datasheets for datasets
Gebru, T.; Morgenstern, J.; Vecchione, B.; Vaughan, J. W.; Wallach, H.; Iii, H. D.; and Crawford, K. 2021 · 2021
Later among the works it cites.
Algorithmic hiring in practice: Recruiter and HR Professional’s perspectives on AI use in hiring
Li, L.; Lassiter, T.; Oh, J.; and Lee, M. K. 2021 · 2021
Later among the works it cites.
A survey on bias and fairness in machine learning
Mehrabi, N.; Morstatter, F.; Saxena, N.; Lerman, K.; and Galstyan, A. 2021 · 2021
Later among the works it cites.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
Paullada, A.; Raji, I. D.; Bender, E. M.; Denton, E.; and Hanna, A. 2021 · 2021
Later among the works it cites.
On the state of reporting in crowdsourcing experiments and a checklist to aid current practices
Ramírez, J.; Sayin, B.; Baez, M.; Casati, F.; Cernuzzi, L.; Benatallah, B.; and Demartini, G. 2021 · 2021
Later among the works it cites.
“Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI
Sambasivan, N.; Kapania, S.; Highfill, H.; Akrong, D.; Paritosh, P.; and Aroyo, L. M. 2021 · 2021
Later among the works it cites.
Cross-replication Reliability - An Empirical Approach to Interpreting Inter-rater Reliability
Wong, K.; Paritosh, P. K.; and Aroyo, L. 2021 · 2021
Later among the works it cites.
Re-labeling imagenet: from single to multi-labels, from global to localized labels
Yun, S.; Oh, S. J.; Heo, B.; Han, D.; Choe, J.; and Chun, S. 2021 · 2021
Later among the works it cites.
Quantified Reproducibility Assessment of NLP Results
Belz, A.; Popovic, M.; and Mille, S. 2022 · 2022
Later among the works it cites.
Measuring Annotator Agreement Generally across Complex Structured, Multi-object, and Free-text Annotation Tasks
Braylan, A.; Alonso, O.; and Lease, M. 2022 · 2022
Later among the works it cites.
Crowdworksheets: Accounting for individual and collective identities underlying crowdsourced dataset annotation
Díaz, M.; Kivlichan, I.; Rosen, R.; Baker, D.; Amironesei, R.; Prabhakaran, V.; and Denton, E. 2022 · 2022
Later among the works it cites.
The Effects of Crowd Worker Biases in Fact-Checking Tasks
Draws, T.; La Barbera, D.; Soprano, M.; Roitero, K.; Ceolin, D.; Checco, A.; and Mizzaro, S. 2022 · 2022
Later among the works it cites.
Fine-tuning machine confidence with human relevance for video discovery
Inel, O.; and Aroyo, L. 2022 · 2022
Later among the works it cites.
On reporting scores and agreement for error annotation tasks
Popović, M.; and Belz, A. 2022 · 2022
Later among the works it cites.
Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI
Pushkarna, M.; Zaldivar, A.; and Kjartansson, O. 2022 · 2022
Later among the works it cites.
The Crowd is Made of People: Observations from Large-Scale Crowd Labelling
Thomas, P.; Kazai, G.; White, R.; and Craswell, N. 2022 · 2022
Later among the works it cites.
ISO/IEC 25012-based methodology for managing data quality requirements in the development of information systems: Towards Data Quality by Design
Guerra-García, C.; Nikiforova, A.; Jiménez, S.; Perez-Gonzalez, H. G.; Ramírez-Torres, M.; and Ontañon-García, L. 2023 · 2023
Closest in time.