Fetching the paper…
Reading the bibliography…
Annotated data is an essential ingredient in natural language processing for training and evaluating machine learning models.
RoBERTa: A robustly optimized BERT pretraining approach
Liu, Yinhan, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
A New Measure of Rank Correlation
Kendall, M. G. 1938 · 1938
Earlier work this paper cites.
Individual Comparisons by Ranking Methods
Wilcoxon, Frank. 1945 · 1945
Earlier work this paper cites.
Statistical Theories of Mental Test Scores
Lord, F.M., M.R. Novick, and Allan Birnbaum. 1968 · 1968
Earlier work this paper cites.
So Lernt Man Leben [How to Learn to Live]
Leitner, Sebastian. 1974 · 1974
Earlier work this paper cites.
Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm
Dawid, A. P. and A. M. Skene. 1979 · 1979
Earlier work this paper cites.
The ATIS spoken language systems pilot corpus
Hemphill, Charles T., John J. Godfrey, and George R. Doddington. 1990 · 1990
Earlier work this paper cites.
MUC-5 evaluation metrics
Chinchor, Nancy and Beth Sundheim. 1993 · 1993
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
Platt, John C. 1999 · 1999
Earlier work this paper cites.
LOF: Identifying density-based local outliers
Breunig, Markus M., Hans-Peter Kriegel, Raymond T. Ng, and Jörg Sander. 2000 · 2000
Earlier work this paper cites.
The Detection of Inconsistency in Manually Tagged Text
van Halteren, Hans. 2000 · 2000
Earlier work this paper cites.
Rank aggregation methods for the Web
Dwork, Cynthia, Ravi Kumar, Moni Naor, and D. Sivakumar. 2001 · 2001
Earlier work this paper cites.
Obtaining calibrated probability estimates from decision trees and naive Bayesian classifiers
Zadrozny, Bianca and Charles Elkan. 2001 · 2001
Earlier work this paper cites.
(Semi-)Automatic Detection of Errors in PoS-Tagged Corpora
Kvĕtoň, Pavel and Karel Oliva. 2002 · 2002
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Zadrozny, Bianca and Charles Elkan. 2002 · 2002
Earlier work this paper cites.
Identifying Incorrect Labels in the CoNLL-2003 Corpus
Reiss, Frederick, Hong Xu, Bryan Cutler, Karthik Muthuraman, and Zachary Eichenberger. 2020 · 2003
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Tjong Kim Sang, Erik F. and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
A Companion to Digital Humanities
Schreibman, Susan, Ray Siemens, and John Unsworth, editors. 2004 · 2004
Earlier work this paper cites.
Example-Based Robust Outlier Detection in High Dimensional Datasets
Cui Zhu, H. Kitagawa, and C. Faloutsos. 2005 · 2005
Earlier work this paper cites.
Detecting errors in discontinuous structural annotation
Dickinson, Markus and W. Detmar Meurers. 2005 · 2005
Earlier work this paper cites.
The relationship between Precision-Recall and ROC curves
Davis, Jesse and Mark Goadrich. 2006 · 2006
Earlier work this paper cites.
From detecting errors to automatically correcting them
Dickinson, Markus. 2006 · 2006
Earlier work this paper cites.
Active annotation
Vlachos, Andreas. 2006 · 2006
Earlier work this paper cites.
Coherence and Coreference Revisited
Kehler, A., L. Kertz, H. Rohde, and J. L. Elman. 2007 · 2007
Earlier work this paper cites.
Learning from noisy labels with deep neural networks: A survey
Song, Hwanjun, Minseok Kim, Dongmin Park, Yooju Shin, and Jae-Gil Lee. 2020 · 2007
Earlier work this paper cites.
Corpora in Language Acquisition Research: History, Methods, Perspectives , volume 6 of Trends in Language Acquisition Research
Behrens, Heike, editor. 2008 · 2008
Earlier work this paper cites.
On Detecting Errors in Dependency Treebanks
Boyd, Adriane, Markus Dickinson, and W. Detmar Meurers. 2008 · 2008
Earlier work this paper cites.
Detecting errors in semantic annotation
Dickinson, Markus and Chong Min Lee. 2008 · 2008
Earlier work this paper cites.
Introduction to Information Retrieval
Manning, Christopher D., Prabhakar Raghavan, and Hinrich Schütze. 2008 · 2008
Earlier work this paper cites.
Correcting a POS-Tagged corpus using three complementary methods
Loftsson, Hrafn. 2009 · 2009
Earlier work this paper cites.
Numbers Rule: The Vexing Mathematics of Democracy, from Plato to the Present
Szpiro, George. 2010 · 2010
Earlier work this paper cites.
Error detection for treebank validation
Ambati, Bharat Ram, Rahul Agarwal, Mridul Gupta, Samar Husain, and Dipti Misra Sharma. 2011 · 2011
Earlier work this paper cites.
Reducing the need for double annotation
Dligach, Dmitriy and Martha Palmer. 2011 · 2011
Cited alongside, same era.
Part-of-speech tagging for Twitter: Annotation, features, and experiments
Gimpel, Kevin, Nathan Schneider, Brendan O’Connor, Dipanjan Das, Daniel Mills, Jacob Eisenstein, Michael Heilman, Dani Yogatama, Jeffrey Flanigan, and Noah A. Smith. 2011 · 2011
Cited alongside, same era.
Part-of-Speech Tagging from 97% to 100%: Is It Time for Some Linguistics?
Manning, Christopher D. 2011 · 2011
Cited alongside, same era.
Assignment Problems: Revised Reprint
Burkard, Rainer, Mauro Dell’Amico, and Silvano Martello. 2012 · 2012
Cited alongside, same era.
Approximating theoretical linguistics classification in real data: The case of german "nach" particle verbs
Haselbach, Boris, Kerstin Eckart, Wolfgang Seeker, Kurt Eberle, and Ulrich Heid. 2012 · 2012
Cited alongside, same era.
Learning whom to trust with MACE
Errator: A tool to help detect annotation errors in the Universal Dependencies project
Wisniewski, Guillaume. 2018 · 2018
Later among the works it cites.
FLAIR: An Easy-to-Use Framework for State-of-the-Art NLP
Akbik, Alan, Tanja Bergmann, Duncan Blythe, Kashif Rasul, Stefan Schweter, and Roland Vollgraf. 2019 · 2019
Later among the works it cites.
Sentiment Analysis Is Not Solved! Assessing and Probing Sentiment Classification
Barnes, Jeremy, Lilja Øvrelid, and Erik Velldal. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Outlier Detection for Improved Data Quality and Diversity in Dialog Systems
Larson, Stefan, Anish Mahendran, Andrew Lee, Jonathan K. Kummerfeld, Parker Hill, Michael A. Laurenzano, Johann Hauswald, Lingjia Tang, and Jason Mars. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hovy, Dirk, Taylor Berg-Kirkpatrick, Ashish Vaswani, and Eduard Hovy. 2013 · 2013
Cited alongside, same era.
Natural Language Annotation for Machine Learning
Pustejovsky, J. and Amber Stubbs. 2013 · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, Richard, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
GloVe: Global vectors for word representation
Pennington, Jeffrey, Richard Socher, and Christopher D. Manning. 2014 · 2014
Cited alongside, same era.
POS error detection in automatically annotated corpora
Rehbein, Ines. 2014 · 2014
Cited alongside, same era.
Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation
Aroyo, Lora and Chris Welty. 2015 · 2015
Cited alongside, same era.
Detection of Annotation Errors in Corpora: Detection of Annotation Errors in Corpora
Dickinson, Markus. 2015 · 2015
Cited alongside, same era.
Turning Silver into Gold: Error-Focused Corpus Reannotation with Active Learning
Ménard, Pierre André and Antoine Mougeot. 2019 · 2019
Later among the works it cites.
Inherent Disagreements in Human Textual Inferences
Pavlick, Ellie and Tom Kwiatkowski. 2019 · 2019
Later among the works it cites.
To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks
Peters, Matthew E., Sebastian Ruder, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, Nils and Iryna Gurevych. 2019 · 2019
Later among the works it cites.
DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter
Sanh, Victor, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
CrossWeigh: Training Named Entity Tagger from Imperfect Annotations
Wang, Zihan, Jingbo Shang, Liyuan Liu, Lihao Lu, Jiacheng Liu, and Jiawei Han. 2019 · 2019
Later among the works it cites.
A Study of Incorrect Paraphrases in Crowdsourced User Utterances
Yaghoub-Zadeh-Fard, Mohammad-Ali, Boualem Benatallah, Moshe Chai Barukh, and Shayan Zamanirad. 2019 · 2019
Later among the works it cites.
TACRED Revisited: A Thorough Evaluation of the TACRED Relation Extraction Task
Alt, Christoph, Aleksandra Gabryszak, and Leonhard Hennig. 2020 · 2020
Later among the works it cites.
Not a cute stroke: Analysis of Rule- and Neural Network-based Information Extraction Systems for Brain Radiology Reports
Grivas, Andreas, Beatrice Alex, Claire Grover, Richard Tobin, and William Whiteley. 2020 · 2020
Later among the works it cites.
Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks
Gururangan, Suchin, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Later among the works it cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, Urvashi, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2020
Later among the works it cites.
Multivariate confidence calibration for object detection
Küppers, Fabian, Jan Kronenberger, Amirhossein Shantia, and Anselm Haselhoff. 2020 · 2020
Later among the works it cites.
Inconsistencies in Crowdsourced Slot-Filling Annotations: A Typology and Identification Methods
Larson, Stefan, Adrian Cheung, Anish Mahendran, Kevin Leach, and Jonathan K. Kummerfeld. 2020 · 2020
Later among the works it cites.
Universal Dependencies v2: An evergrowing multilingual treebank collection
Nivre, Joakim, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020 · 2020
Later among the works it cites.
Dataset Cartography: Mapping and Diagnosing Datasets with Training Dynamics
Swayamdipta, Swabha, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020 · 2020
Later among the works it cites.
We Need to Consider Disagreement in Evaluation
Basile, Valerio, Michael Fell, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio, and Alexandra Uma. 2021 · 2021
Later among the works it cites.
Beyond Black & White: Leveraging Annotator Disagreement via Soft-Label Multi-Task Learning
Fornaciari, Tommaso, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, and Massimo Poesio. 2021 · 2021
Later among the works it cites.
Analysing the Noise Model Error for Realistic Noisy Label Data
Hedderich, Michael A., Dawei Zhu, and Dietrich Klakow. 2021 · 2021
Later among the works it cites.
Confident Learning: Estimating Uncertainty in Dataset Labels
Northcutt, Curtis, Lu Jiang, and Isaac Chuang. 2021 · 2021
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
Northcutt, Curtis G., Anish Athalye, and Jonas Mueller. 2021 · 2021
Later among the works it cites.
Annotation inconsistency and entity bias in MultiWOZ
Qian, Kun, Ahmad Beirami, Zhouhan Lin, Ankita De, Alborz Geramifard, Zhou Yu, and Chinnadhurai Sankar. 2021 · 2021
Later among the works it cites.
Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?
Rodriguez, Pedro, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, and Jordan Boyd-Graber. 2021 · 2021
Later among the works it cites.
How Certain is Your Transformer?
Shelmanov, Artem, Evgenii Tsymbalov, Dmitri Puzyrev, Kirill Fedyanin, Alexander Panchenko, and Maxim Panov. 2021 · 2021
Later among the works it cites.
Re-TACRED: Addressing Shortcomings of the TACRED Dataset
Stoica, George, Emmanouil Antonios Platanios, and Barnabas Poczos. 2021 · 2021
Later among the works it cites.
A general-purpose crowdsourcing computational quality control toolkit for python
Ustalov, Dmitry, Nikita Pavlichenko, Vladimir Losev, Iulian Giliazev, and Evgeny Tulin. 2021 · 2021
Later among the works it cites.
Crowdsourcing Learning as Domain Adaptation: A Case Study on Named Entity Recognition
Zhang, Xin, Guangwei Xu, Yueheng Sun, Meishan Zhang, and Pengjun Xie. 2021 · 2021
Later among the works it cites.
Meta label correction for noisy label learning
Zheng, Guoqing, Ahmed Hassan Awadallah, and Susan Dumais. 2021 · 2021
Later among the works it cites.