Fetching the paper…
Reading the bibliography…
Large language models are becoming the go-to solution for the ever-growing number of tasks.
Biases in large language models: Origins, inventory, and discussion
Roberto Navigli, Simone Conia, and Björn Ross · 1955
Earlier work this paper cites.
Econometric Theory
A.S. Goldberger, W.A. Shenhart, and S.S. Wilks · 1964
Earlier work this paper cites.
Direct and indirect effects
Judea Pearl · 2001
Earlier work this paper cites.
Gender Bias in Coreference Resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme · 2002
Earlier work this paper cites.
Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2003
Earlier work this paper cites.
Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M. Shieber · 2004
Earlier work this paper cites.
The Winograd Schema Challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam Tauman Kalai · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Evaluating gender bias in machine translation
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
Counterfactual Data Augmentation for Mitigating Gender Stereotypes in Languages with Rich Morphology
Ran Zmigrod, S. J. Mielke, Hanna M. Wallach, and Ryan Cotterell · 2019
Cited alongside, same era.
Language (technology) is Power: A Critical Survey of ”bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna M. Wallach · 2020
Cited alongside, same era.
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman · 2020
Cited alongside, same era.
Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach · 2021
Cited alongside, same era.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov · 2021
Cited alongside, same era.
Transformer Feed-Forward Layers Are Key-Value Memories
Debiasing Pre-trained Language Models via Efficient Fine-tuning
Michael Gira, Ruisu Zhang, and Kangwook Lee · 2022
Later among the works it cites.
Auto-Debias: Debiasing Masked Language Models with Automated Biased Prompts
Yue Guo, Yi Yang, and Ahmed Abbasi · 2022
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Later among the works it cites.
Don’t forget about pronouns: Removing gender bias in language models without losing factual gender information
Tomasz Limisiewicz and David Mareček · 2022
Later among the works it cites.
Locating and Editing Factual Associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Later among the works it cites.
Fast model editing at scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Cited alongside, same era.
Measuring massive multitask language understanding, 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Sustainable modular debiasing of language models
Anne Lauscher, Tobias Lueken, and Goran Glavaš · 2021
Cited alongside, same era.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2021
Cited alongside, same era.
Gender Bias in Machine Translation
Beatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri, and Marco Turchi · 2021
Cited alongside, same era.
A Survey on Gender Bias in Natural Language Processing
Karolina Stanczak and Isabelle Augenstein · 2021
Cited alongside, same era.
Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models
Pieter Delobelle, Ewoenam Tokpo, Toon Calders, and Bettina Berendt · 2022
Cited alongside, same era.
Later among the works it cites.
Linear Adversarial Concept Erasure
Shauli Ravfogel, Michael Twiton, Yoav Goldberg, and Ryan Cotterell · 2022
Later among the works it cites.
LEACE: Perfect linear concept erasure in closed form
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman · 2023
Closest in time.
Mass-Editing Memory in a Transformer
Kevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov, and David Bau · 2023
Closest in time.
A Trip Towards Fairness: Bias and De-biasing in Large Language Models
Leonardo Ranaldi, Elena Sofia Ruzzetti, Davide Venditti, Dario Onorati, and Fabio Massimo Zanzotto · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Undesirable biases in nlp: Averting a crisis of measurement, 2023
Oskar van der Wal, Dominik Bachmann, Alina Leidinger, Leendert van Maanen, Willem Zuidema, and Katrin Schulz · 2023
Closest in time.
An Empirical Analysis of Parameter-efficient Methods for Debiasing Pre-trained Language Models
Zhongbin Xie and Thomas Lukasiewicz · 2023
Closest in time.