Fetching the paper…
Reading the bibliography…
Natural Language Processing (NLP) models risk overfitting to specific terms in the training data, thereby reducing their performance, fairness, and generalizability.
Quantifying the carbon emissions of machine learning
Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres. 2019 · 1910
Earlier work this paper cites.
A mathematical theory of communication
C. E. Shannon. 1948 · 1948
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020 · 1967
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves. 2013 · 2013
Earlier work this paper cites.
What’s in a p-value in NLP?
Anders Søgaard, Anders Johannsen, Barbara Plank, Dirk Hovy, and Hector Martínez Alonso. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? Debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on Twitter
Zeerak Waseem and Dirk Hovy. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
What does attention in neural machine translation pay attention to?
Hamidreza Ghader and Christof Monz. 2017 · 2017
Earlier work this paper cites.
A multi-view sentiment corpus
Debora Nozza, Elisabetta Fersini, and Enza Messina. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Challenges in automated debiasing for toxic language detection
Xuhui Zhou, Maarten Sap, Swabha Swayamdipta, Yejin Choi, and Noah Smith. 2021 · 2017
Earlier work this paper cites.
Universal sentence encoder for English
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil. 2018 · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Earlier work this paper cites.
Adversarial removal of demographic attributes from text data
Yanai Elazar and Yoav Goldberg. 2018 · 2018
Earlier work this paper cites.
Overview of the EVALITA 2018 Task on Automatic Misogyny Identification (AMI)
Elisabetta Fersini, Debora Nozza, and Paolo Rosso. 2018 · 2018
Earlier work this paper cites.
Word embeddings quantify 100 years of gender and ethnic stereotypes
Nikhil Garg, Londa Schiebinger, Dan Jurafsky, and James Zou. 2018 · 2018
Cited alongside, same era.
Reducing gender bias in abusive language detection
Ji Ho Park, Jamin Shin, and Pascale Fung. 2018 · 2018
Cited alongside, same era.
Hierarchical CVAE for fine-grained hate speech classification
Jing Qian, Mai ElSherief, Elizabeth Belding, and William Yang Wang. 2018 · 2018
Cited alongside, same era.
Nuanced metrics for measuring unintended bias with real data for text classification
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2019 · 2019
Cited alongside, same era.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2019a · 2019
Cited alongside, same era.
What does BERT look at? an analysis of BERT’s attention
On identifiability in transformers
Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020 · 2020
Later among the works it cites.
AMI @ EVALITA2020: Automatic Misogyny Identification
Elisabetta Fersini, Debora Nozza, and Paolo Rosso. 2020 · 2020
Later among the works it cites.
Contextualizing hate speech classifiers with post-hoc explanation
Brendan Kennedy, Xisen Jin, Aida Mostafazadeh Davani, Morteza Dehghani, and Xiang Ren. 2020 · 2020
Later among the works it cites.
Jigsaw@ AMI and HaSpeeDe2: Fine-Tuning a Pre-Trained Comment-Domain BERT Model
Alyssa Lees, Jeffrey Sorensen, and Ian Kivlichan. 2020 · 2020
Later among the works it cites.
Comparative evaluation of label-agnostic selection bias in multilingual hate speech datasets
Nedjma Ousidhoum, Yangqiu Song, and Dit-Yan Yeung. 2020 · 2020
Later among the works it cites.
Null it out: Guarding protected attributes by iterative nullspace projection
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019b · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
FERMI at SemEval-2019 task 5: Using sentence embeddings to identify hate speech against immigrants and women in Twitter
Vijayasaradhi Indurthi, Bakhtiyar Syed, Manish Shrivastava, Nikhil Chakravartula, Manish Gupta, and Vasudeva Varma. 2019 · 2019
Cited alongside, same era.
Towards hierarchical importance attribution: Explaining compositional semantics for neural sequence models
Xisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue, and Xiang Ren. 2019 · 2019
Cited alongside, same era.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Cited alongside, same era.
Unintended bias in misogyny detection
Debora Nozza, Claudia Volpetti, and Elisabetta Fersini. 2019 · 2019
Cited alongside, same era.
Multilingual and multi-aspect hate speech analysis
Nedjma Ousidhoum, Zizheng Lin, Hongming Zhang, Yangqiu Song, and Dit-Yan Yeung. 2019 · 2019
Cited alongside, same era.
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
Data augmentation for discrimination prevention and bias disambiguation
Shubham Sharma, Yunfeng Zhang, Jesús M. Ríos Aliaga, Djallel Bouneffouf, Vinod Muthusamy, and Kush R. Varshney. 2020 · 2020
Later among the works it cites.
Empirical analysis of multi-task learning for reducing identity bias in toxic comment detection
Ameya Vaidya, Feng Mai, and Yue Ning. 2020 · 2020
Later among the works it cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Demographics should not be the reason of toxicity: Mitigating discrimination in text classifications with instance weighting
Guanhua Zhang, Bing Bai, Junqi Zhang, Kun Bai, Conghui Zhu, and Tiejun Zhao. 2020 · 2020
Later among the works it cites.
Stereotype and skew: Quantifying gender bias in pre-trained and fine-tuned language models
Daniel de Vassimon Manela, David Errington, Thomas Fisher, Boris van Breugel, and Pasquale Minervini. 2021 · 2021
Later among the works it cites.
Debiasing pre-trained contextualised embeddings
Masahiro Kaneko and Danushka Bollegala. 2021 · 2021
Later among the works it cites.
Exposing the limits of zero-shot cross-lingual hate speech detection
Debora Nozza. 2021 · 2021
Later among the works it cites.
HONEST: Measuring hurtful sentence completion in language models
Debora Nozza, Federico Bianchi, and Dirk Hovy. 2021 · 2021
Later among the works it cites.
Probing toxic content in large pre-trained language models
Nedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song, and Dit-Yan Yeung. 2021 · 2021
Later among the works it cites.
Effective attention sheds light on interpretability
Kaiser Sun and Ana Marasović. 2021 · 2021
Later among the works it cites.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2021 · 2021
Later among the works it cites.