Fetching the paper…
Reading the bibliography…
The use of machine learning (ML)-based language models (LMs) to monitor content online is on the rise.
A coefficient of agreement for nominal scales
J. Cohen · 1960
Earlier work this paper cites.
Management research based on the paradigm of the design sciences: the quest for field-tested and grounded technological rules
J. E. v. Aken · 2004
Earlier work this paper cites.
Design science in information systems research
A. R. Hevner, S. T. March, J. Park, and S. Ram · 2004
Earlier work this paper cites.
Annotating expressions of opinions and emotions in language
J. Wiebe, T. Wilson, and C. Cardie · 2005
Earlier work this paper cites.
Answering the call for a standard reliability measure for coding data
A. F. Hayes and K. Krippendorff · 2007
Earlier work this paper cites.
Multi-label classification: An overview
G. Tsoumakas and I. Katakis · 2007
Earlier work this paper cites.
Cheap and fast–but is it good? evaluating non-expert annotations for natural language tasks
R. Snow, B. O’connor, D. Jurafsky, and A. Y. Ng · 2008
Earlier work this paper cites.
Design science research evaluation
K. Peffers, M. Rothenberger, T. Tuunanen, and R. Vaezi · 2012
Earlier work this paper cites.
Positioning and presenting design science research for maximum impact
S. Gregor and A. R. Hevner · 2013
Earlier work this paper cites.
Inter-annotator agreement
R. Artstein · 2017
Cited alongside, same era.
Using a visual abstract as a lens for communicating and promoting design science research in software engineering
M.-A. Storey, E. Engstrom, M. Höst, P. Runeson, and E. Bjarnason · 2017
Cited alongside, same era.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
E. M. Bender and B. Friedman · 2018
Cited alongside, same era.
M. Geva, Y. Goldberg, and J. Berant · 2019
Cited alongside, same era.
The risk of racial bias in hate speech detection
M. Sap, D. Card, S. Gabriel, Y. Choi, and N. A. Smith · 2019
Cited alongside, same era.
Hatexplain: A benchmark dataset for explainable hate speech detection
B. Mathew, P. Saha, S. M. Yimam, C. Biemann, P. Goyal, and A. Mukherjee · 2020
Later among the works it cites.
Differential tweetment: Mitigating racial dialect bias in harmful tweet detection
A. Ball-Burack, M. S. A. Lee, J. Cobbe, and J. Singh · 2021
Closest in time.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
A. M. Davani, M. Díaz, and V. Prabhakaran · 2021
Closest in time.
“garbage in, garbage out” revisited: What do machine learning application papers report about human-labeled training data?
R. S. Geiger, D. Cope, J. Ip, M. Lotosh, A. Shah, J. Weng, and R. Tang · 2021
Closest in time.
The disagreement deconvolution: Bringing machine learning performance metrics in line with reality
M. L. Gordon, K. Zhou, K. Patel, T. Hashimoto, and M. S. Bernstein · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Sap, S. Gabriel, L. Qin, D. Jurafsky, N. A. Smith, and Y. Choi · 2019
Cited alongside, same era.
Detoxify
L. Hanu and Unitary team · 2020
Cited alongside, same era.
Using text classification to improve annotation quality by improving annotator consistency
E. Ishita, S. Fukuda, Y. Tomiura, and D. W. Oard · 2020
Cited alongside, same era.
Closest in time.
An expert annotated dataset for the detection of online misogyny
E. Guest, B. Vidgen, A. Mittos, N. Sastry, G. Tyson, and H. Margetts · 2021
Closest in time.
What is the ground truth? reliability of multi-annotator data for audio tagging
I. Martin-Morato and A. Mesaros · 2021
Closest in time.
Pervasive label errors in test sets destabilize machine learning benchmarks
C. G. Northcutt, A. Athalye, and J. Mueller · 2021
Closest in time.