Fetching the paper…
Reading the bibliography…
We commonly use agreement measures to assess the utility of judgements made by human annotators in Natural Language Processing (NLP) tasks.
Assessing agreement on classification tasks: the kappa statistic
Jean Carletta. 1996 · 1996
Earlier work this paper cites.
Intra-rater and inter-rater reliability of a weight-bearing lunge measure of ankle dorsiflexion
Kim Bennell, Richard Talbot, Henry Wajswelner, Wassana Techovanich, David Kelly, and AJ Hall. 1998 · 1998
Earlier work this paper cites.
An annotation scheme for discourse-level argumentation in research articles
Simone Teufel, Jean Carletta, and Marc Moens. 1999 · 1999
Earlier work this paper cites.
A review and analysis of research on the test–retest reliability of professional judgment
Robert H. Ashton. 2000 · 2000
Earlier work this paper cites.
Limb apraxia, pantomine, and lexical gesture in aphasic speakers: Preliminary findings
Miranda Rose and Jacinta Douglas. 2003 · 2003
Earlier work this paper cites.
CIU and main event analyses of the structured discourse of older and younger adults
Gilson Capilouto, Heather Harris Wright, and Stacy A. Wagovich. 2005 · 2005
Earlier work this paper cites.
A corpus for studying addressing behavior in multi-party dialogues
Natasa Jovanovic, Rieks op den Akker, and Anton Nijholt. 2005 · 2005
Earlier work this paper cites.
Annotation guidelines for Czech-English word alignment
Ivana Kruijff-Korbayová, Klára Chvátalová, and Oana Postolache. 2006 · 2006
Earlier work this paper cites.
(meta-) evaluation of machine translation
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2007 · 2007
Earlier work this paper cites.
Inconsistency of expert judgment-based estimates of software development effort
Stein Grimstad and Magne Jørgensen. 2007 · 2007
Earlier work this paper cites.
Further meta-evaluation of machine translation
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2008 · 2008
Earlier work this paper cites.
Children’s oral reading corpus (CHOREC): Description and assessment of annotator agreement
Leen Cleuren, Jacques Duchateau, Pol Ghesquière, and Hugo Van hamme. 2008 · 2008
Earlier work this paper cites.
An examination of judge reliability at a major U.S. wine competition
Robert T. Hodgson. 2008 · 2008
Earlier work this paper cites.
How well does active learning actually
Jason Baldridge and Alexis Palmer. 2009 · 2009
Earlier work this paper cites.
Findings of the 2009 Workshop on Statistical Machine Translation
Chris Callison-Burch, Philipp Koehn, Christof Monz, and Josh Schroeder. 2009 · 2009
Earlier work this paper cites.
Preferred reporting items for systematic reviews and meta-analyses: The prisma statement
David Moher, Alessandro Liberati, Jennifer Tetzlaff, Douglas G. Altman, and the PRISMA Group. 2009 · 2009
Earlier work this paper cites.
Multilingual extension of PDTB-style annotation: The case of TED multilingual discourse bank
Deniz Zeyrek, Amália Mendes, and Murathan Kurfalı. 2018 · 2009
Earlier work this paper cites.
Findings of the 2010 joint workshop on statistical machine translation and metrics for machine translation
Chris Callison-Burch, Philipp Koehn, Christof Monz, Kay Peterson, Mark Przybocki, and Omar Zaidan. 2010 · 2010
Earlier work this paper cites.
Is my judge a good one?
Olivier Hamon. 2010 · 2010
Earlier work this paper cites.
Improving the post-editing experience using translation recommendation: A user study
Yifan He, Yanjun Ma, Johann Roturier, Andy Way, and Josef van Genabith. 2010 · 2010
Earlier work this paper cites.
Annotating underquantification
Aurelie Herbelot and Ann Copestake. 2010 · 2010
Earlier work this paper cites.
Enriching word alignment with linguistic tags
Xuansong Li, Niyu Ge, Stephen Grimes, Stephanie M. Strassel, and Kazuaki Maeda. 2010 · 2010
Earlier work this paper cites.
Discrete vs. continuous rating scales for language evaluation in NLP
Anja Belz and Eric Kow. 2011 · 2011
Earlier work this paper cites.
Getting expert quality from the crowd for machine translation evaluation
Luisa Bentivogli, Marcello Federico, Giovanni Moretti, and Michael Paul. 2011 · 2011
Earlier work this paper cites.
Quiz-based evaluation of machine translation
Jan Berka, Martin Černý, and Ondřej Bojar. 2011 · 2011
Earlier work this paper cites.
Findings of the 2011 workshop on statistical machine translation
Chris Callison-Burch, Philipp Koehn, Christof Monz, and Omar Zaidan. 2011 · 2011
Earlier work this paper cites.
Performance of automated scoring for children’s oral reading
Ryan Downey, David Rubin, Jian Cheng, and Jared Bernstein. 2011 · 2011
Earlier work this paper cites.
Automatic emotion classification for interpersonal communication
Frederik Vaassen and Walter Daelemans. 2011 · 2011
Earlier work this paper cites.
Findings of the 2012 workshop on statistical machine translation
Chris Callison-Burch, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia. 2012 · 2012
Earlier work this paper cites.
Annotation schemes to encode domain knowledge in medical narratives
Wilson McCoy, Cecilia Ovesdotter Alm, Cara Calvelli, Rui Li, Jeff B. Pelz, Pengcheng Shi, and Anne Haake. 2012 · 2012
Cited alongside, same era.
Source-language dictionaries help non-expert users to enlarge target-language dictionaries for machine translation
Víctor M. Sánchez-Cartagena, Miquel Esplà-Gomis, and Juan Antonio Pérez-Ortiz. 2012 · 2012
Cited alongside, same era.
Findings of the 2013 Workshop on Statistical Machine Translation
Ondřej Bojar, Christian Buck, Chris Callison-Burch, Christian Federmann, Barry Haddow, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia. 2013 · 2013
Cited alongside, same era.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Cited alongside, same era.
Learning whom to trust with MACE
Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, and Eduard Hovy. 2013 · 2013
Cited alongside, same era.
Implicit and explicit racial attitudes changed during Black Lives Matter
Jeremy Sawyer and Anup Gampa. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Later among the works it cites.
A crowdsourced corpus of multiple judgments and disagreement on anaphoric interpretation
Massimo Poesio, Jon Chamberlain, Silviu Paun, Juntao Yu, Alexandra Uma, and Udo Kruschwitz. 2019 · 2019
Later among the works it cites.
Analysis of automatic annotation suggestions for hard discourse-level tasks in expert domains
Claudia Schulz, Christian M. Meyer, Jan Kiesewetter, Michael Sailer, Elisabeth Bauer, Martin R. Fischer, Frank Fischer, and Iryna Gurevych. 2019 · 2019
Later among the works it cites.
Correct me if you can: Learning from error corrections and markings
Julia Kreutzer, Nathaniel Berger, and Stefan Riezler. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Statistical machine translation for automobile marketing texts
Samuel Läubli, Mark Fishel, Manuela Weibel, and Martin Volk. 2013 · 2013
Cited alongside, same era.
Findings of the 2014 workshop on statistical machine translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014 · 2014
Cited alongside, same era.
A human judgement corpus and a metric for Arabic MT evaluation
Houda Bouamor, Hanan Alshikhabobakr, Behrang Mohit, and Kemal Oflazer. 2014 · 2014
Cited alongside, same era.
Situation entity annotation
Annemarie Friedrich and Alexis Palmer. 2014 · 2014
Cited alongside, same era.
Decision style in a clinical reasoning corpus
Limor Hochberg, Cecilia Ovesdotter Alm, Esa M. Rantanen, Caroline M. DeLong, and Anne Haake. 2014a · 2014
Cited alongside, same era.
Query-focused opinion summarization for user-generated content
Lu Wang, Hema Raghavan, Claire Cardie, and Vittorio Castelli. 2014 · 2014
Cited alongside, same era.
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty. 2015 · 2015
Cited alongside, same era.
Shallow discourse annotation for Chinese TED talks
Wanqiu Long, Xinyi Cai, James Reid, Bonnie Webber, and Deyi Xiong. 2020 · 2020
Later among the works it cites.
Views of sexual assault following #MeToo: The role of gender and individual differences
Hanna Szekeres, Eric Shuman, and Tamar Saguy. 2020 · 2020
Later among the works it cites.
Annotating verbal MWEs in Irish for the PARSEME shared task 1.2
Abigail Walsh, Teresa Lynn, and Jennifer Foster. 2020 · 2020
Later among the works it cites.
On exposure bias, hallucination and domain shift in neural machine translation
Chaojun Wang and Rico Sennrich. 2020 · 2020
Later among the works it cites.
Findings of the 2021 conference on machine translation (WMT21)
Farhad Akhbardeh, Arkady Arkhangorodsky, Magdalena Biesialska, Ondřej Bojar, Rajen Chatterjee, Vishrav Chaudhary, Marta R. Costa-jussa, Cristina España-Bonet, Angela Fan, Christian Federmann, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Leonie Harter, Kenneth Heafield, Christopher Homan, Matthias Huck, Kwabena Amponsah-Kaakyire, Jungo Kasai, Daniel Khashabi, Kevin Knight, Tom Kocmi, Philipp Koehn, Nicholas Lourie, Christof Monz, Makoto Morishita, Masaaki Nagata, Ajay Nagesh, Toshiaki Nakazawa, Matteo Negri, Santanu Pal, Allahsera Auguste Tapo, Marco Turchi, Valentin Vydrin, and Marcos Zampieri. 2021 · 2021
Later among the works it cites.
Toward a perspectivist turn in ground truthing for predictive computing
Valerio Basile, Federico Cabitza, Andrea Campagner, and Michael Fell. 2021a · 2021
Later among the works it cites.
Sociolinguistically Driven Approaches for Just Natural Language Processing
Su Lin Blodgett. 2021 · 2021
Later among the works it cites.
ConvAbuse: Data, analysis, and benchmarks for nuanced abuse detection in conversational AI
Amanda Cercas Curry, Gavin Abercrombie, and Verena Rieser. 2021 · 2021
Later among the works it cites.
SemEval-2021 task 11: NLPContributionGraph - structuring scholarly NLP contributions for a research knowledge graph
Jennifer D’Souza, Sören Auer, and Ted Pedersen. 2021 · 2021
Later among the works it cites.
SuperSim: a test set for word similarity and relatedness in Swedish
Simon Hengchen and Nina Tahmasebi. 2021 · 2021
Later among the works it cites.
Whit’s the richt pairt o speech: PoS tagging for Scots
Harm Lameris and Sara Stymne. 2021 · 2021
Later among the works it cites.
Agreeing to disagree: Annotating offensive language datasets with annotators’ disagreement
Elisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini, and Sara Tonelli. 2021 · 2021
Later among the works it cites.
Derek Chauvin found guilty of George Floyd’s murder
Chris McGreal. 2021 · 2021
Later among the works it cites.
Beyond fair pay: Ethical implications of NLP crowdsourcing
Boaz Shmueli, Jan Fell, Soumya Ray, and Lun-Wei Ku. 2021 · 2021
Later among the works it cites.
Preregistering NLP research
Emiel van Miltenburg, Chris van der Lee, and Emiel Krahmer. 2021 · 2021
Later among the works it cites.
Exploring the importance of source text in automatic post-editing for context-aware machine translation
Chaojun Wang, Christian Hardmeier, and Rico Sennrich. 2021 · 2021
Later among the works it cites.
Theory-grounded measurement of U.S. social stereotypes in English language models
Yang Cao, Anna Sotnikova, Hal Daumé III, Rachel Rudinger, and Linda Zou. 2022 · 2022
Later among the works it cites.
StereoKG: Data-driven knowledge graph construction for cultural knowledge and stereotypes
Awantee Deshpande, Dana Ruiter, Marius Mosbach, and Dietrich Klakow. 2022 · 2022
Later among the works it cites.
The ‘problem’ of human label variation: On ground truth in data, modeling and evaluation
Barbara Plank. 2022 · 2022
Later among the works it cites.
Two contrasting data annotation paradigms for subjective NLP tasks
Paul Rottger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022 · 2022
Later among the works it cites.
Exploiting social media content for self-supervised style transfer
Dana Ruiter, Thomas Kleinbauer, Cristina España-Bonet, Josef van Genabith, and Dietrich Klakow. 2022 · 2022
Later among the works it cites.
Temporal and second language influence on intra-annotator agreement and stability in hate speech labelling
Gavin Abercrombie, Dirk Hovy, and Vinodkumar Prabhakaran. 2023 · 2023
Closest in time.
VariErr NLI: Separating annotation error from human label variation
Leon Weber-Genzel, Siyao Peng, Marie-Catherine De Marneffe, and Barbara Plank. 2024 · 2024
Closest in time.