Fetching the paper…
Reading the bibliography…
Human label variation arises when annotators assign different labels to the same item for valid reasons, while annotation errors occur when labels are assigned for invalid reasons.
Multi-stage document ranking with BERT
Rodrigo Frassetto Nogueira, Wei Yang, Kyunghyun Cho, and Jimmy Lin. 2019 · 1910
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Detecting Errors in Part-of-Speech Annotation
Markus Dickinson and W. Detmar Meurers. 2003 · 2003
Earlier work this paper cites.
Identifying Incorrect Labels in the CoNLL-2003 Corpus
Frederick Reiss, Hong Xu, Bryan Cutler, Karthik Muthuraman, and Zachary Eichenberger. 2020 · 2003
Earlier work this paper cites.
Lucas Beyer, Olivier J Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aäron van den Oord. 2020 · 2006
Earlier work this paper cites.
Part-of-speech tagging from 97% to 100%: Is it time for some linguistics?
Christopher D Manning. 2011 · 2011
Earlier work this paper cites.
Did it happen? The pragmatic complexity of veridicality assessment
Marie-Catherine de Marneffe, Christopher D. Manning, and Christopher Potts. 2012 · 2012
Earlier work this paper cites.
Crowd truth: Harnessing disagreement in crowdsourcing a relation extraction gold standard
Lora Aroyo and Chris Welty. 2013 · 2013
Earlier work this paper cites.
Linguistically debatable or just plain wrong?
Barbara Plank, Dirk Hovy, and Anders Søgaard. 2014 · 2014
Earlier work this paper cites.
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty. 2015 · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Spotting Spurious Data with Neural Networks
Hadi Amiri, Timothy Miller, and Guergana Savova. 2018 · 2016
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
Unsupervised Label Noise Modeling and Loss Correction
Eric Arazo, Diego Ortego, Paul Albert, Noel O’Connor, and Kevin Mcguinness. 2019 · 2019
Earlier work this paper cites.
O2U-Net: A Simple Noisy Label Detection Approach for Deep Neural Networks
Jinchi Huang, Lie Qu, Rongfei Jia, and Binqiang Zhao. 2019 · 2019
Earlier work this paper cites.
Outlier detection for improved data quality and diversity in dialog systems
Stefan Larson, Anish Mahendran, Andrew Lee, Jonathan K. Kummerfeld, Parker Hill, Michael A. Laurenzano, Johann Hauswald, Lingjia Tang, and Jason Mars. 2019 · 2019
Earlier work this paper cites.
Inherent disagreements in human textual inferences
Ellie Pavlick and Tom Kwiatkowski. 2019 · 2019
Cited alongside, same era.
CrossWeigh: Training Named Entity Tagger from Imperfect Annotations
Zihan Wang, Jingbo Shang, Liyuan Liu, Lihao Lu, Jiacheng Liu, and Jiawei Han. 2019 · 2019
Cited alongside, same era.
TACRED Revisited: A Thorough Evaluation of the TACRED Relation Extraction Task
Christoph Alt, Aleksandra Gabryszak, and Leonhard Hennig. 2020 · 2020
Cited alongside, same era.
Abductive commonsense reasoning
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen tau Yih, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
Investigating reasons for disagreement in natural language inference
Nan-Jiang Jiang and Marie-Catherine de Marneffe. 2022 · 2022
Later among the works it cites.
The “problem” of human label variation: On ground truth in data, modeling and evaluation
Barbara Plank. 2022 · 2022
Later among the works it cites.
Two contrasting data annotation paradigms for subjective NLP tasks
Paul Rottger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022 · 2022
Later among the works it cites.
When does dough become a bagel? Analyzing the remaining mistakes on ImageNet
Vijay Vasudevan, Benjamin Caine, Raphael Gontijo Lopes, Sara Fridovich-Keil, and Rebecca Roelofs. 2022 · 2022
Later among the works it cites.
Toward a perspectivist turn in ground truthing for predictive computing
Federico Cabitza, Andrea Campagner, and Valerio Basile. 2023 · 2023
Later among the works it cites.
Ecologically valid explanations for label variation in NLI
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What can we learn from collective human opinions on natural language inference data?
Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020 · 2020
Cited alongside, same era.
Would you describe a leopard as yellow? Evaluating crowd-annotations with justified and informative disagreement
Pia Sommerauer, Antske Fokkens, and Piek Vossen. 2020 · 2020
Cited alongside, same era.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
What will it take to fix benchmarking in natural language understanding?
Samuel R. Bowman and George Dahl. 2021 · 2021
Cited alongside, same era.
A data-centric approach for training deep neural networks with less data
Mohammad Motamedi, Nikolay Sakharnykh, and Tim Kaldewey. 2021 · 2021
Cited alongside, same era.
Pervasive label errors in test sets destabilize machine learning benchmarks
Curtis G Northcutt, Anish Athalye, and Jonas Mueller. 2021 · 2021
Cited alongside, same era.
Learning from Disagreement: A Survey
Alexandra N. Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021 · 2021
Cited alongside, same era.
Nan-Jiang Jiang, Chenhao Tan, and Marie-Catherine de Marneffe. 2023 · 2023
Later among the works it cites.
Annotation error detection: Analyzing the past and present for a more coherent future
Jan-Christoph Klie, Bonnie Webber, and Iryna Gurevych. 2023 · 2023
Later among the works it cites.
DataPerf: Benchmarks for Data-Centric AI Development
Mark Mazumder, Colby Banbury, Xiaozhe Yao, Bojan Karlaš, William Gaviria Rojas, Sudnya Diamos, Greg Diamos, Lynn He, Alicia Parrish, Hannah Rose Kirk, Jessica Quaye, Charvi Rastogi, Douwe Kiela, David Jurado, David Kanter, Rafael Mosquera, Will Cukierski, Juan Ciro, Lora Aroyo, Bilge Acun, Lingjiao Chen, Mehul Raje, Max Bartolo, Evan Sabri Eyuboglu, Amirata Ghorbani, Emmett Goodman, Addison Howard, Oana Inel, Tariq Kane, Christine R. Kirkpatrick, D. Sculley, Tzu-Sheng Kuo, Jonas W Mueller, Tristan Thrush, Joaquin Vanschoren, Margaret Warren, Adina Williams, Serena Yeung, Newsha Ardalani, Praveen Paritosh, Ce Zhang, James Y Zou, Carole-Jean Wu, Cody Coleman, Andrew Ng, Peter Mattson, and Vijay Janapa Reddi. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
CleanCoNLL: A nearly noise-free named entity recognition dataset
Susanna Rücker and Alan Akbik. 2023 · 2023
Later among the works it cites.
Metadata archaeology: Unearthing data subsets by leveraging training dynamics
Shoaib Ahmed Siddiqui, Nitarshan Rajkumar, Tegan Maharaj, David Krueger, and Sara Hooker. 2023 · 2023
Later among the works it cites.
ActiveAED: A human in the loop improves annotation error detection
Leon Weber and Barbara Plank. 2023 · 2023
Later among the works it cites.
Efficiently programming large language models using sglang
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2023 · 2023
Later among the works it cites.
Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs
Simone Balloccu, Patrícia Schmidtová, Mateusz Lango, and Ondrej Dusek. 2024 · 2024
Closest in time.
Label Smarter, Not Harder: CleverLabel for Faster Annotation of Ambiguous Image Classification with Higher Quality
Lars Schmarje, Vasco Grossmann, Tim Michels, Jakob Nazarenus, Monty Santarossa, Claudius Zelenka, and Reinhard Koch. 2024 · 2024
Closest in time.