Fetching the paper…
Reading the bibliography…
Generic `toxicity' classifiers continue to be used for evaluating the potential for harm in natural language generation, despite mounting evidence of their shortcomings.
Angry White Men: American Masculinity at the End of an Era , first edition
Michael Kimmel. 2013 · 2013
Earlier work this paper cites.
Down Girl: The Logic of Misogyny
Kate Manne. 2017 · 2017
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2019 · 2019
Earlier work this paper cites.
Exploring misogyny across the manosphere in reddit
Tracie Farrell, Miriam Fernandez, Jakub Novotny, and Harith Alani. 2019 · 2019
Earlier work this paper cites.
Alphas, betas, and incels: Theorizing the masculinities of the manosphere
Debbie Ging. 2019 · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019 · 2019
Earlier work this paper cites.
The pushshift reddit dataset
Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020 · 2020
Earlier work this paper cites.
On the use of Jargon and Word Embeddings to Explore Subculture within the Reddit’s Manosphere
Tracie Farrell, Oscar Araque, Miriam Fernandez, and Harith Alani. 2020 · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Fortifying toxic speech detectors against veiled toxicity
Xiaochuang Han and Yulia Tsvetkov. 2020 · 2020
Earlier work this paper cites.
Detoxify: Toxic comment classification with pytorch lightning and huggingface transformers
Laura Hanu and Unitary team. 2020 · 2020
Cited alongside, same era.
A Critical Audit of Accuracy and Demographic Biases within Toxicity Detection Tools
Jiachen Jiang. 2020 · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Cited alongside, same era.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Cited alongside, same era.
Reading between the demographic lines: Resolving sources of bias in toxicity classifiers
Elizabeth Reichert, Helen Qiu, and Jasmine Bayrooti. 2020 · 2020
Cited alongside, same era.
Examining incel subculture on reddit
Brenna Helm, Ryan Scrivens, Thomas J. Holt, Steve Chermak, and Richard Frank. 2022 · 2022
Later among the works it cites.
Platforming Hate: An Analysis of Incel Ideology on Reddit’s r/Braincels
Jeffrey J. Millar. 2022 · 2022
Later among the works it cites.
Characteristics of harmful text: Towards rigorous benchmarking of language models
Maribeth Rauh, John Mellor, Jonathan Uesato, Po-Sen Huang, Johannes Welbl, Laura Weidinger, Sumanth Dathathri, Amelia Glaese, Geoffrey Irving, and Iason Gabriel. 2022 · 2022
Later among the works it cites.
Identifying Different Layers of Online Misogyny
Wienke Strathern and Juergen Pfeffer. 2022 · 2022
Later among the works it cites.
Operationalising ‘toxicity’ in the manosphere: Automation, platform governance and community health
Verity Trott, Jennifer Beckett, and Venessa Paech. 2022 · 2022
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Cited alongside, same era.
NoisyHate: Benchmarking Content Moderation Machine Learning Models with Human-Written Perturbations Online
Yiran Ye, Thai Le, and Dongwon Lee. 2023 · 2020
Cited alongside, same era.
HateCheck: Functional tests for hate speech detection models
Paul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert. 2021 · 2021
Cited alongside, same era.
Fighting hate speech, silencing drag queens? artificial intelligence in content moderation and risks to LGBTQ voices online
Dias Oliva Thiago, Antonialli Dennys Marcelo, and Alessandra Gomes. 2021 · 2021
Cited alongside, same era.
Incels on Reddit: A study in social norms and decentralised moderation
Rosalie Gillett and Nicolas Suzor. 2022 · 2022
Cited alongside, same era.
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, and Edward Raff. 2023 · 2023
Closest in time.
Holistic Evaluation of Language Models
Rishi Bommasani, Percy Liang, and Tony Lee. 2023 · 2023
Closest in time.
A benchmark for toxic comment classification on civil comments dataset
Corentin Duchêne, Henri Jamet, Pierre Guillaume, and Réda Dehak. 2023 · 2023
Closest in time.
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
Julia Mendelsohn, Ronan Le Bras, Yejin Choi, and Maarten Sap. 2023 · 2023
Closest in time.
On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research
Luiza Pozzobon, Beyza Ermis, Patrick Lewis, and Sara Hooker. 2023 · 2023
Closest in time.