Fetching the paper…
Reading the bibliography…
Counterfactually Augmented Data (CAD) aims to improve out-of-domain generalizability, an indicator of model robustness.
Scikit-learn: Machine learning in python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. 2011 · 2011
Earlier work this paper cites.
Classification and Its Consequences for Online Harassment: Design Insights from HeartMob
Lindsay Blackwell, Jill Dimond, Sarita Schoenebeck, and Cliff Lampe. 2017 · 2017
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Earlier work this paper cites.
Custodians of the Internet
Tarleton Gillespie. 2018 · 2018
Earlier work this paper cites.
SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter
Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso, and Manuela Sanguinetti. 2019 · 2019
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019 · 2019
Earlier work this paper cites.
The platform governance triangle: Conceptualising the informal regulation of online content
Robert Gorwa. 2019 · 2019
Earlier work this paper cites.
A just and comprehensive strategy for using NLP to address online abuse
David Jurgens, Libby Hemphill, and Eshwar Chandrasekharan. 2019 · 2019
Earlier work this paper cites.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. 2019 · 2019
Earlier work this paper cites.
Unintended Bias in Misogyny Detection
Debora Nozza, Claudia Volpetti, and Elisabetta Fersini. 2019 · 2019
Earlier work this paper cites.
Behind the screen
Sarah T Roberts. 2019 · 2019
Cited alongside, same era.
Challenges and frontiers in abusive content detection
Bertie Vidgen, Alex Harris, Dong Nguyen, Rebekah Tromble, Scott Hale, and Helen Margetts. 2019 · 2019
Cited alongside, same era.
Language (Technology) is Power: A Critical Survey of “Bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Cited alongside, same era.
Evaluating Models’ Local Decision Boundaries via Contrast Sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020 · 2020
Cited alongside, same era.
Content moderation, ai, and the question of scale
Tarleton Gillespie. 2020 · 2020
Cited alongside, same era.
Disproportionate removals and differing content moderation experiences for conservative, transgender, and black social media users: Marginalization and moderation gray areas
Oliver L Haimson, Daniel Delmonaco, Peipei Nie, and Andrea Wegner. 2021 · 2021
Later among the works it cites.
Causal direction of data collection matters: Implications of causal and anticausal learning for NLP
Zhijing Jin, Julius von Kügelgen, Jingwei Ni, Tejas Vaidhya, Ayush Kaushal, Mrinmaya Sachan, and Bernhard Schoelkopf. 2021 · 2021
Later among the works it cites.
Text as causal mediators: Research design for causal estimates of differential treatment of social groups via language aspects
Katherine Keith, Douglas Rice, and Brendan O’Connor. 2021 · 2021
Later among the works it cites.
Removing spurious features can hurt accuracy and affect groups disproportionately
Fereshte Khani and Percy Liang. 2021 · 2021
Later among the works it cites.
Detecting abusive language on online platforms: A critical analysis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Algorithmic content moderation: Technical and political challenges in the automation of platform governance
Robert Gorwa, Reuben Binns, and Christian Katzenbach. 2020 · 2020
Cited alongside, same era.
Contextualizing hate speech classifiers with post-hoc explanation
Brendan Kennedy, Xisen Jin, Aida Mostafazadeh Davani, Morteza Dehghani, and Xiang Ren. 2020 · 2020
Cited alongside, same era.
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Cited alongside, same era.
Aaa: Fair evaluation for abuse detection systems wanted
Agostina Calabrese, Michele Bevilacqua, Björn Ross, Rocco Tripodi, and Roberto Navigli. 2021 · 2021
Cited alongside, same era.
Causal inference in natural language processing: Estimation, prediction, interpretation and beyond
Amir Feder, Katherine A Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E Roberts, et al. 2021 · 2021
Cited alongside, same era.
We ‘said her name’and got zucked”: Black Women Calling-out the Carceral Logics of Digital Platforms
Kishonna L Gray and Krysten Stein. 2021 · 2021
Cited alongside, same era.
Preslav Nakov, Vibha Nayak, Kyle Dent, Ameya Bhatawdekar, Sheikh Muhammad Sarwar, Momchil Hardalov, Yoan Dinkov, Dimitrina Zlatkova, Guillaume Bouchard, and Isabelle Augenstein. 2021 · 2021
Later among the works it cites.
Overview of exist 2021: sexism identification in social networks
Francisco Rodriguez-Sanchez, Jorge Carrillo de Albornoz, Laura Plaza, Julio Gonzalo, Paolo Rosso, Miriam Comet, and Trinidad Donoso. 2021 · 2021
Later among the works it cites.
HateCheck: Functional tests for hate speech detection models
Paul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert. 2021 · 2021
Later among the works it cites.
Call me sexist, but…: Revisiting sexism detection using psychological scales and adversarial samples
Mattia Samory, Indira Sen, Julian Kohne, Fabian Flöck, and Claudia Wagner. 2021 · 2021
Later among the works it cites.
How does counterfactually augmented data impact models for social computing constructs?
Indira Sen, Mattia Samory, Fabian Flöck, Claudia Wagner, and Isabelle Augenstein. 2021 · 2021
Later among the works it cites.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2021 · 2021
Later among the works it cites.
Fact Checking with Insufficient Evidence
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2022 · 2022
Closest in time.