Fetching the paper…
Reading the bibliography…
Human variation in labeling is often considered noise.
Maximum likelihood estimation of observer error-rates using the em algorithm
Alexander Philip Dawid and Allan M Skene. 1979 · 1979
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
Generalized entropy regularization or: There’s nothing special about label smoothing
Clara Meister, Elizabeth Salesky, and Ryan Cotterell. 2020 · 2005
Earlier work this paper cites.
The reliability of anaphoric annotation, reconsidered: Taking ambiguity into account
Massimo Poesio and Ron Artstein. 2005 · 2005
Earlier work this paper cites.
Inter-coder agreement for computational linguistics
Ron Artstein and Massimo Poesio. 2008 · 2008
Earlier work this paper cites.
Analyzing disagreements
Beata Beigman Klebanov, Eyal Beigman, and Danie Diermaier. 2008 · 2008
Earlier work this paper cites.
Reliability measurement without limits
Dennis Reidsma and Jean Carletta. 2008 · 2008
Earlier work this paper cites.
Exploiting ‘subjective’ annotations
Dennis Reidsma and Rieks op den Akker. 2008 · 2008
Earlier work this paper cites.
Get another label? improving data quality and data mining using multiple, noisy labelers
Victor S. Sheng, Foster Provost, and Panagiotis G. Ipeirotis. 2008 · 2008
Earlier work this paper cites.
Cheap and fast–but is it good? evaluating non-expert annotations for natural language tasks
Rion Snow, Brendan O’connor, Dan Jurafsky, and Andrew Y Ng. 2008 · 2008
Earlier work this paper cites.
Learning with annotation noise
Eyal Beigman and Beata Beigman-Klebanov. 2009 · 2009
Earlier work this paper cites.
Chris J Kennedy, Geoff Bacon, Alexander Sahn, and Claudia von Vacano. 2020 · 2009
Earlier work this paper cites.
Word sense annotation of polysemous words by multiple annotators
Rebecca J Passonneau, Ansaf Salleb-Aouissi, Vikas Bhardwaj, and Nancy Ide. 2010 · 2010
Earlier work this paper cites.
Hard problems of tagset conversion
Daniel Zeman. 2010 · 2010
Earlier work this paper cites.
Subjective natural language problems: Motivations, applications, characterizations, and implications
Cecilia Ovesdotter Alm. 2011 · 2011
Earlier work this paper cites.
Part-of-speech tagging from 97% to 100%: is it time for some linguistics?
Christopher D Manning. 2011 · 2011
Earlier work this paper cites.
Did it happen? the pragmatic complexity of veridicality assessment
Marie-Catherine de Marneffe, Christopher D. Manning, and Christopher Potts. 2012 · 2012
Earlier work this paper cites.
Natural Language Annotation for Machine Learning: A guide to corpus-building for applications
James Pustejovsky and Amber Stubbs. 2012 · 2012
Earlier work this paper cites.
Discourse structure and computation: Past, present and future
Bonnie Webber and Aravind Joshi. 2012 · 2012
Earlier work this paper cites.
Modelling annotator bias with multi-task Gaussian processes: An application to machine translation quality estimation
Trevor Cohn and Lucia Specia. 2013 · 2013
Earlier work this paper cites.
Empirical analysis of aggregation methods for collective annotation
Ciyang Qing, Ulle Endriss, Raquel Fernández, and Justin Kruger. 2014 · 2014
Earlier work this paper cites.
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty. 2015 · 2015
Earlier work this paper cites.
How far are we from fully automatic high quality grammatical error correction?
Christopher Bryant and Hwee Tou Ng. 2015 · 2015
Earlier work this paper cites.
Noise or additional information? leveraging crowdsource annotation item agreement for natural language tasks
Emily Jamison and Iryna Gurevych. 2015 · 2015
Earlier work this paper cites.
Obtaining Well Calibrated Probabilities Using Bayesian Binning
Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015 · 2015
Earlier work this paper cites.
Anchoring and agreement in syntactic annotations
Yevgeni Berzak, Yan Huang, Andrei Barbu, Anna Korhonen, and Boris Katz. 2016 · 2016
Earlier work this paper cites.
Broad Twitter corpus: A diverse named entity recognition resource
Leon Derczynski, Kalina Bontcheva, and Ian Roberts. 2016 · 2016
Earlier work this paper cites.
Embracing error to enable rapid crowdsourcing
Ranjay A Krishna, Kenji Hata, Stephanie Chen, Joshua Kravitz, David A Shamma, Li Fei-Fei, and Michael S Bernstein. 2016 · 2016
Earlier work this paper cites.
Understanding discourse on work and job-related well-being in public social media
Tong Liu, Christopher Homan, Cecilia Ovesdotter Alm, Megan Lytle, Ann Marie White, and Henry Kautz. 2016 · 2016
Earlier work this paper cites.
Supersense tagging with inter-annotator disagreement
Héctor Martínez Alonso, Anders Johannsen, and Barbara Plank. 2016 · 2016
Earlier work this paper cites.
Why is that relevant? collecting annotator rationales for relevance judgments
Tyler McDonnell, Matthew Lease, Mucahid Kutlu, and Tamer Elsayed. 2016 · 2016
Cited alongside, same era.
The good, the bad, and the disagreement: Complex ground truth in rhetorical structure analysis
Debopam Das, Manfred Stede, and Maite Taboada. 2017 · 2017
Cited alongside, same era.
Training deep neural-networks using a noise adaptation layer
Jacob Goldberger and Ehud Ben-Reuven. 2016 · 2017
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017 · 2017
Cited alongside, same era.
Soft label memorization-generalization for natural language inference
John P Lalor, Hao Wu, and Hong Yu. 2017 · 2017
Cited alongside, same era.
A computational exploration of pejorative language in social media
Liviu P. Dinu, Ioan-Bogdan Iordache, Ana Sabina Uban, and Marcos Zampieri. 2021 · 2021
Later among the works it cites.
Did they answer? subjective acts and intents in conversational discourse
Elisa Ferracane, Greg Durrett, Junyi Jessy Li, and Katrin Erk. 2021 · 2021
Later among the works it cites.
Beyond black & white: Leveraging annotator disagreement via soft-label multi-task learning
Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, and Massimo Poesio. 2021 · 2021
Later among the works it cites.
The disagreement deconvolution: Bringing machine learning performance metrics in line with reality
Mitchell L Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S Bernstein. 2021 · 2021
Later among the works it cites.
How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emily M Bender and Batya Friedman. 2018 · 2018
Cited alongside, same era.
Crowd disagreement about medical images is informative
Veronika Cheplygina and Josien PW Pluim. 2018 · 2018
Cited alongside, same era.
Crowdsourcing semantic label propagation in relation classification
Anca Dumitrache, Lora Aroyo, and Chris Welty. 2018 · 2018
Cited alongside, same era.
Content analysis: An introduction to its methodology
Klaus Krippendorff. 2018 · 2018
Cited alongside, same era.
A case for a range of acceptable annotations
Jennimaria Palomaki, Olivia Rhinehart, and Michael Tseng. 2018 · 2018
Cited alongside, same era.
Deep learning from crowds
Filipe Rodrigues and Francisco Pereira. 2018 · 2018
Cited alongside, same era.
The commitmentbank: Investigating projection in naturally occurring discourse
Marie-Catherine de Marneffe, Mandy Simons, and Judith Tonhauser. 2019 · 2019
Cited alongside, same era.
EaSe: A diagnostic tool for VQA based on answer diversity
Shailza Jolly, Sandro Pezzelle, and Moin Nabi. 2021 · 2021
Later among the works it cites.
Reconsidering annotator disagreement about racist language: Noise or signal?
Savannah Larimore, Ian Kennedy, Breon Haskett, and Alina Arseniev-Koehler. 2021 · 2021
Later among the works it cites.
Agreeing to disagree: Annotating offensive language datasets with annotators’ disagreement
Elisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini, and Sara Tonelli. 2021 · 2021
Later among the works it cites.
On releasing annotator-level labels and information in datasets
Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. 2021 · 2021
Later among the works it cites.
Survey equivalence: A procedure for measuring classifier accuracy against human labels
Paul Resnick, Yuqing Kong, Grant Schoenebeck, and Tim Weninger. 2021 · 2021
Later among the works it cites.
Targeting the benchmark: On methodology in current natural language processing research
David Schlangen. 2021 · 2021
Later among the works it cites.
SemEval-2021 task 12: Learning with disagreements
Alexandra Uma, Tommaso Fornaciari, Anca Dumitrache, Tristan Miller, Jon Chamberlain, Barbara Plank, Edwin Simpson, and Massimo Poesio. 2021a · 2021
Later among the works it cites.
Investigating annotator bias in abusive language datasets
Maximilian Wich, Christian Widmer, Gerhard Hagerer, and Georg Groh. 2021 · 2021
Later among the works it cites.
Cartography active learning
Mike Zhang and Barbara Plank. 2021 · 2021
Later among the works it cites.
Learning with different amounts of annotation: From zero to many labels
Shujian Zhang, Chengyue Gong, and Eunsol Choi. 2021 · 2021
Later among the works it cites.
Identifying inherent disagreement in natural language inference
Xinliang Frederick Zhang and Marie-Catherine de Marneffe. 2021 · 2021
Later among the works it cites.
The ARMIS dataset of misogyny in arabic tweets
Dina Almanea and Massimo Poesio. 2022 · 2022
Closest in time.
Stop measuring calibration when humans disagree
Joris Baan, Wilker Aziz, Barbara Plank, and Raquel Fernandez. 2022 · 2022
Closest in time.
CrossRE: A Cross-Domain Dataset for Relation Extraction
Elisa Bassignana and Barbara Plank. 2022 · 2022
Closest in time.
All that glitters… interannotator agreement in natural language processing
Lars Borin. 2022 · 2022
Closest in time.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022 · 2022
Closest in time.
Jury learning: Integrating dissenting voices into machine learning models
Mitchell L Gordon, Michelle S Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S Bernstein. 2022 · 2022
Closest in time.
Challenges and strategies in cross-cultural NLP
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders Søgaard. 2022 · 2022
Closest in time.
Investigating reasons for disagreement in natural language inference
Nan-Jiang Jiang and Marie-Catherine de Marneffe. 2022 · 2022
Closest in time.
Annotation error detection: Analyzing the past and present for a more coherent future
Jan-Christoph Klie, Bonnie Webber, and Iryna Gurevych. 2022 · 2022
Closest in time.
Establishing annotation quality in multi-label annotations
Marian Marchal, Merel Scholman, Frances Yung, and Vera Demberg. 2022 · 2022
Closest in time.
Statistical methods for annotation analysis
Silviu Paun, Ron Artstein, and Massimo Poesio. 2022 · 2022
Closest in time.
Gcdt: A chinese rst treebank for multigenre and multilingual discourse parsing
Siyao Peng, Yang Janet Liu, and Amir Zeldes. 2022 · 2022
Closest in time.
Two contrasting data annotation paradigms for subjective NLP tasks
Paul Rottger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022 · 2022
Closest in time.
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022 · 2022
Closest in time.