Fetching the paper…
Reading the bibliography…
As NLP models are increasingly deployed in socially situated settings such as online abusive content detection, it is crucial to ensure that these models are robust.
Global aggregations of local explanations for black box models
Ilse van der Linden, Hinda Haned, and Evangelos Kanoulas. 2019 · 1907
Earlier work this paper cites.
Mining and summarizing customer reviews
Minqing Hu and Bing Liu. 2004 · 2004
Earlier work this paper cites.
Biographies, Bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification
John Blitzer, Mark Dredze, and Fernando Pereira. 2007 · 2007
Earlier work this paper cites.
Targeting the benchmark: On methodology in current natural language processing research
David Schlangen. 2020 · 2007
Earlier work this paper cites.
Transpeople, transprejudice and pathologization: A seven-country factor analytic study
Sam Winter, Pornthip Chalungsooth, Yik Koon Teh, Nongnuch Rojanalert, Kulthida Maneerat, Ying Wuen Wong, Anne Beaumont, Loretta Man Wah Ho, Francis “Chuck” Gomez, and Raymond Aquino Macapagal. 2009 · 2009
Earlier work this paper cites.
Customizing triggers with concealed data poisoning
Eric Wallace, Tony Z Zhao, Shi Feng, and Sameer Singh. 2020 · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. 2011 · 2011
Earlier work this paper cites.
Dynasent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela. 2020 · 2012
Earlier work this paper cites.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2020 · 2012
Earlier work this paper cites.
Counterfactuals
David Lewis. 2013 · 2013
Earlier work this paper cites.
Direct and indirect treatment effects: causal chains and mediation analysis with instrumental variables
Markus Frölich and Martin Huber. 2014 · 2014
Earlier work this paper cites.
Interpretation and identification of causal mediation
Judea Pearl. 2014 · 2014
Earlier work this paper cites.
The politics of measurement and action
Kathleen H. Pine and Max Liboiron. 2015 · 2015
Earlier work this paper cites.
Whipping girl: A transsexual woman on sexism and the scapegoating of femininity
Julia Serano. 2016 · 2016
Earlier work this paper cites.
Analyzing the targets of hate in online social media
Leandro Silva, Mainack Mondal, Denzil Correa, Fabrício Benevenuto, and Ingmar Weber. 2016 · 2016
Earlier work this paper cites.
Beyond the belmont principles: Ethical challenges, practices, and beliefs in the online data research community
Jessica Vitak, Katie Shilton, and Zahra Ashktorab. 2016 · 2016
Earlier work this paper cites.
Are you a racist or am I seeing things? annotator influence on hate speech detection on Twitter
Zeerak Waseem. 2016 · 2016
Earlier work this paper cites.
Classification and its consequences for online harassment: Design insights from heartmob
Lindsay Blackwell, Jill Dimond, Sarita Schoenebeck, and Cliff Lampe. 2017 · 2017
Earlier work this paper cites.
A survey on hate speech detection using natural language processing
Anna Schmidt and Michael Wiegand. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Internet research ethics for the social age: New challenges, cases, and contexts
Michael Zimmer and Katharina Kinder-Kurlanda. 2017 · 2017
Cited alongside, same era.
Adversarial example generation with syntactically controlled paraphrase networks
Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Causal and counterfactual inference
Judea Pearl. 2018 · 2018
Cited alongside, same era.
Turning words into consumer preferences: How sentiment analysis is framed in research and the news media
Cornelius Puschmann and Alison Powell. 2018 · 2018
Cited alongside, same era.
Data poisoning attack against unsupervised node embedding methods
Mingjie Sun, Jian Tang, Huichen Li, Bo Li, Chaowei Xiao, Yao Chen, and Dawn Song. 2018 · 2018
Cited alongside, same era.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. 2020 · 2020
Later among the works it cites.
TextAttack: A framework for adversarial attacks, data augmentation, and adversarial training in NLP
John Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020 · 2020
Later among the works it cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Later among the works it cites.
Comparative evaluation of label-agnostic selection bias in multilingual hate speech datasets
Nedjma Ousidhoum, Yangqiu Song, and Dit-Yan Yeung. 2020 · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter
Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso, and Manuela Sanguinetti. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019 · 2019
Cited alongside, same era.
Counterfactual fairness in text classification through robustness
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H. Chi, and Alex Beutel. 2019 · 2019
Cited alongside, same era.
Facebook while black: Users call it getting ‘zucked,’say talking about racism is censored as hate speech
Jessica Guynn. 2019 · 2019
Cited alongside, same era.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Cited alongside, same era.
A just and comprehensive strategy for using NLP to address online abuse
David Jurgens, Libby Hemphill, and Eshwar Chandrasekharan. 2019 · 2019
Cited alongside, same era.
Learning what makes a difference from counterfactual examples and gradient supervision
Damien Teney, Ehsan Abbasnedjad, and Anton van den Hengel. 2020 · 2020
Later among the works it cites.
Directions in abusive language training data, a systematic review: Garbage in, garbage out
Bertie Vidgen and Leon Derczynski. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Measurement and fairness
Abigail Z. Jacobs and Hanna Wallach. 2021 · 2021
Closest in time.
Perspective is reducing toxicity in the real world
Jigsaw. 2021 · 2021
Closest in time.
An investigation of the (in) effectiveness of counterfactually augmented data
Nitish Joshi and He He. 2021 · 2021
Closest in time.
The use and misuse of counterfactuals in ethical machine learning
Atoosa Kasirzadeh and Andrew Smart. 2021 · 2021
Closest in time.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021 · 2021
Closest in time.
Generate your counterfactuals: Towards controlled counterfactual generation for text
Nishtha Madaan, Inkit Padhi, Naveen Panwar, and Diptikalyan Saha. 2021 · 2021
Closest in time.
Detecting abusive language on online platforms: A critical analysis
Preslav Nakov, Vibha Nayak, Kyle Dent, Ameya Bhatawdekar, Sheikh Muhammad Sarwar, Momchil Hardalov, Yoan Dinkov, Dimitrina Zlatkova, Guillaume Bouchard, and Isabelle Augenstein. 2021 · 2021
Closest in time.
Overview of exist 2021: sexism identification in social networks
Francisco Ródriguez-Sánchez, Jorge Carrillo de Albornoz, Laura Plaza, Julio Gonzalo, Paolo Rosso, Miriam Comet, and Trinidad Donoso. 2021 · 2021
Closest in time.
"call me sexist, but…" : Revisiting sexism detection using psychological scales and adversarial samples
Mattia Samory, Indira Sen, Julian Kohne, Fabian Floeck, and Claudia Wagner. 2021 · 2021
Closest in time.
A Neighbourhood Framework for Resource-Lean Content Flagging
Sheikh Muhammad Sarwar, Dimitrina Zlatkova, Momchil Hardalov, Yoan Dinkov, Isabelle Augenstein, and Preslav Nakov. 2021 · 2021
Closest in time.
Measuring algorithmically infused societies
Claudia Wagner, Markus Strohmaier, Alexandra Olteanu, Emre Kıcıman, Noshir Contractor, and Tina Eliassi-Rad. 2021 · 2021
Closest in time.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2021 · 2021
Closest in time.