Fetching the paper…
Reading the bibliography…
A common practice in building NLP datasets, especially using crowd-sourced annotations, involves obtaining multiple annotator judgements on the same data instances, which are then flattened to produce a single "ground truth" label or score, through majority voting, averaging, or adjudication.
An argument for basic emotions
Paul Ekman. 1992 · 1992
Earlier work this paper cites.
Experiments in emotional speech
Julia Hirschberg, Jackson Liscombe, and Jennifer Venditti. 2003 · 2003
Earlier work this paper cites.
A model of textual affect sensing using real-world knowledge
Hugo Liu, Henry Lieberman, and Ted Selker. 2003 · 2003
Earlier work this paper cites.
Janyce Wiebe, Theresa Wilson, Rebecca Bruce, Matthew Bell, and Melanie Martin. 2004 · 2004
Earlier work this paper cites.
Affect in* Text and Speech
Ebba Cecilia Ovesdotter Alm. 2008 · 2008
Earlier work this paper cites.
Cheap and fast–but is it good? evaluating non-expert annotations for natural language tasks
Rion Snow, Brendan O’connor, Dan Jurafsky, and Andrew Y Ng. 2008 · 2008
Earlier work this paper cites.
Sentiment analysis and subjectivity
Bing Liu et al. 2010 · 2010
Earlier work this paper cites.
How reliable are annotations via crowdsourcing: a study about inter-annotator agreement for multi-label image annotation
Stefanie Nowak and Stefan Rüger. 2010 · 2010
Earlier work this paper cites.
Subjective natural language problems: Motivations, applications, characterizations, and implications
Cecilia Ovesdotter Alm. 2011 · 2011
Earlier work this paper cites.
Statistical modality tagging from rule-based annotations and crowdsourcing
Vinodkumar Prabhakaran, Michael Bloodgood, Mona Diab, Bonnie Dorr, Lori Levin, Christine Piatko, Owen Rambow, and Benjamin Van Durme. 2012 · 2012
Earlier work this paper cites.
Detecting hate speech on the world wide web
William Warner and Julia Hirschberg. 2012 · 2012
Earlier work this paper cites.
Crowd truth: Harnessing disagreement in crowdsourcing a relation extraction gold standard
Lora Aroyo and Chris Welty. 2013 · 2013
Cited alongside, same era.
Modelling annotator bias with multi-task gaussian processes: An application to machine translation quality estimation
Trevor Cohn and Lucia Specia. 2013 · 2013
Cited alongside, same era.
Learning part-of-speech taggers with inter-annotator agreement loss
Barbara Plank, Dirk Hovy, and Anders Søgaard. 2014 · 2014
Cited alongside, same era.
Are you a racist or am i seeing things? annotator influence on hate speech detection on twitter
Zeerak Waseem. 2016 · 2016
Cited alongside, same era.
Hateful symbols or hateful people? predictive features for hate speech detection on twitter
Zeerak Waseem and Dirk Hovy. 2016 · 2016
Cited alongside, same era.
Every rating matters: Joint learning of subjective labels and individual annotators for speech emotion classification
Huang-Cheng Chou and Chi-Chun Lee. 2019 · 2019
Later among the works it cites.
Annotating social media data from vulnerable populations: Evaluating disagreement between domain experts and graduate student annotators
Desmond Patton, Philipp Blandfort, William Frey, Michael Gaskell, and Svebor Karaman. 2019 · 2019
Later among the works it cites.
GoEmotions: A dataset of fine-grained emotions
Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020 · 2020
Later among the works it cites.
The gab hate corpus: A collection of 27k posts annotated for hate speech
Brendan Kennedy, Mohammad Atari, Aida Mostafazadeh Davani, Leigh Yeh, Ali Omrani, Yehsong Kim, Kris Coombs Jr., Shreya Havaldar, Gwenyth Portillo-Wightman, Elaine Gonzalez, Joe Hoover, Aida Azatian, Gabriel Cardenas, Alyzeh Hussain, Austin Lara, Adam Omary, Christina Park, Xin Wang, Clarisa Wijaya, Yong Zhang, Beth Meyerowitz, and Morteza Dehghani. 2020 · 2020
Later among the works it cites.
Detecting stance in media on global warming
Yiwei Luo, Dallas Card, and Dan Jurafsky. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017 · 2017
Cited alongside, same era.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman. 2018 · 2018
Cited alongside, same era.
Addressing age-related bias in sentiment analysis
Mark Díaz, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle. 2018 · 2018
Cited alongside, same era.
Large scale crowdsourcing and characterization of twitter abusive behavior
Antigoni Maria Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. 2018 · 2018
Cited alongside, same era.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2018 · 2018
Cited alongside, same era.
Who said what: Modeling individual labelers improves classification
Melody Guan, Varun Gulshan, Andrew Dai, and Geoffrey Hinton. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Investigating annotator bias with a graph-based approach
Maximilian Wich, Hala Al Kuwatly, and Georg Groh. 2020 · 2020
Later among the works it cites.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Aida Mostafazadeh Davani, Mark Diaz, and Vinodkumar Prabhakaran. 2021 · 2021
Closest in time.
Beyond black & white: Leveraging annotator disagreement via soft-label multi-task learning
Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, and Massimo Poesio. 2021 · 2021
Closest in time.
Toxic comment classification challenge
Jigsaw. 2018 · 2021
Closest in time.
Unintended bias in toxicity classification
Jigsaw. 2019 · 2021
Closest in time.