Fetching the paper…
Reading the bibliography…
Text classifiers have promising applications in high-stake tasks such as resume screening and content moderation.
An empirical study on learning fairness metrics for compas data with human supervision
Hanchen Wang, Nina Grgic-Hlaca, Preethi Lahoti, Krishna P. Gummadi, and Adrian Weller · 1910
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani, and Máté Lengyel · 2011
Earlier work this paper cites.
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan · 2017
Earlier work this paper cites.
Deep bayesian active learning with image data
Yarin Gal, Riashat Islam, and Zoubin Ghahramani · 2017
Earlier work this paper cites.
Counterfactual fairness
Matt J Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva · 2017
Earlier work this paper cites.
The limits of abstract evaluation metrics: The case of hate speech detection
Alexandra Olteanu, Kartik Talamadupula, and Kush R Varshney · 2017
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman · 2018
Earlier work this paper cites.
Online learning with an unknown fairness metric
Stephen Gillen, Christopher Jung, Michael Kearns, and Aaron Roth · 2018
Earlier work this paper cites.
Harini Kannan, Alexey Kurakin, and Ian Goodfellow · 2018
Earlier work this paper cites.
Crowdsourcing subjective tasks: the case study of understanding toxicity in online discussions
Lora Aroyo, Lucas Dixon, Nithum Thain, Olivia Redfield, and Rachel Rosen · 2019
Earlier work this paper cites.
Fairness and Machine Learning: Limitations and Opportunities
Solon Barocas, Moritz Hardt, and Arvind Narayanan · 2019
Earlier work this paper cites.
End-to-end resume parsing and finding candidates for a job description using bert
Vedant Bhatia, Prateek Rawat, Ajit Kumar, and Rajiv Ratn Shah · 2019
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Lee Kention, and Kristina Toutanova · 2019
Earlier work this paper cites.
Counterfactual fairness in text classification through robustness
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H Chi, and Alex Beutel · 2019
Earlier work this paper cites.
Operationalizing individual fairness with pairwise fair representations
Preethi Lahoti, Krishna P Gummadi, and Gerhard Weikum · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
How to fine-tune bert for text classification?
Chi Sun, Xipeng Qiu, Yige Xu, and Xuanjing Huang · 2019
Cited alongside, same era.
Training individually fair ml models with sensitive subspace robustness
Mikhail Yurochkin, Amanda Bower, and Yuekai Sun · 2019
Cited alongside, same era.
The elephant in the interpretability room: Why use attention as explanation when we have saliency methods?
Jasmijn Bastings and Katja Filippova · 2020
Cited alongside, same era.
Metric-free individual fairness in online learning
Yahav Bechavod, Christopher Jung, and Steven Z Wu · 2020
Cited alongside, same era.
Soliciting stakeholders’ fairness notions in child maltreatment predictive systems
Hao-Fei Cheng, Logan Stapleton, Ruiqi Wang, Paige Bullock, Alexandra Chouldechova, Zhiwei Steven Steven Wu, and Haiyi Zhu · 2021
Later among the works it cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Later among the works it cites.
An algorithmic framework for fairness elicitation
Christopher Jung, Michael Kearns, Seth Neel, Aaron Roth, Logan Stapleton, and Zhiwei Steven Wu · 2021
Later among the works it cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2021
Later among the works it cites.
Lewis: Levenshtein editing for unsupervised text style transfer
Machel Reid and Victor Zhong · 2021
Later among the works it cites.
The fabrics of machine moderation: Studying the technical, normative, and organizational structure of perspective api
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language (technology) is power: A critical survey of “bias” in nlp
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, et al · 2020
Cited alongside, same era.
Fine-tuning bert for low-resource natural language understanding via active learning
Daniel Grießhaber, Johannes Maucher, and Ngoc Thang Vu · 2020
Cited alongside, same era.
Metric learning for individual fairness
Christina Ilvento · 2020
Cited alongside, same era.
Stable style transformer: Delete and generate approach with encoder-decoder for text style transfer
Joosung Lee · 2020
Cited alongside, same era.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Cited alongside, same era.
Bernhard Rieder and Yarden Skop · 2021
Later among the works it cites.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Later among the works it cites.
Individual fairness revisited: transferring techniques from adversarial robustness
Samuel Yeom and Matt Fredrikson · 2021
Later among the works it cites.
Sensei: Sensitive set invariance for enforcing individual fairness
Mikhail Yurochkin and Yuekai Sun · 2021
Later among the works it cites.
Clean or annotate: How to spend a limited data collection budget
Derek Chen, Zhou Yu, and Samuel R Bowman · 2022
Closest in time.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran · 2022
Closest in time.
Jury learning: Integrating dissenting voices into machine learning models
Mitchell L Gordon, Michelle S Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S Bernstein · 2022
Closest in time.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar · 2022
Closest in time.
Deep learning for text style transfer: A survey
Di Jin, Zhijing Jin, Zhiting Hu, Olga Vechtomova, and Rada Mihalcea · 2022
Closest in time.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Closest in time.
Latent space smoothing for individually fair representations
Momchil Peychev, Anian Ruoss, Mislav Balunović, Maximilian Baader, and Martin Vechev · 2022
Closest in time.
Perturbation augmentation for fairer NLP
Rebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith, Douwe Kiela, and Adina Williams · 2022
Closest in time.
Tailor: Generating and perturbing text with semantic controls
Alexis Ross, Tongshuang Wu, Hao Peng, Matthew E Peters, and Matt Gardner · 2022
Closest in time.
“I’m sorry to hear that”: Finding new biases in language models with a holistic descriptor dataset
Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams · 2022
Closest in time.