Fetching the paper…
Reading the bibliography…
Fairness in Language Models (LMs) remains a longstanding challenge, given the inherent biases in training data that can be perpetuated by models and affect the downstream tasks.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Equalizing gender biases in neural machine translation with word embeddings techniques
Font, J. E.; and Costa-Jussa, M. R. 2019 · 1901
Earlier work this paper cites.
Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology
Zmigrod, R.; Mielke, S. J.; Wallach, H.; and Cotterell, R. 2019b · 1906
Earlier work this paper cites.
The Woman Worked as a Babysitter: On Biases in Language Generation
Sheng, E.; Chang, K.-W.; Natarajan, P.; and Peng, N. 2019 · 1909
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Scao, T. L.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. M. 2020 · 1910
Earlier work this paper cites.
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Nangia, N.; Vania, C.; Bhalerao, R.; and Bowman, S. R. 2020 · 1967
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020 · 2001
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Nadeem, M.; Bethke, A.; and Reddy, S. 2020 · 2004
Earlier work this paper cites.
Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection
Ravfogel, S.; Elazar, Y.; Gonen, H.; Twiton, M.; and Goldberg, Y. 2020a · 2004
Earlier work this paper cites.
Towards Debiasing Sentence Representations
Liang, P. P.; Li, I. M.; Zheng, E.; Lim, Y. C.; Salakhutdinov, R.; and Morency, L.-P. 2020b · 2007
Earlier work this paper cites.
Measuring and Reducing Gendered Correlations in Pre-trained Models
Webster, K.; Wang, X.; Tenney, I.; Beutel, A.; Pitler, E.; Pavlick, E.; Chen, J.; Chi, E.; and Petrov, S. 2021 · 2010
Earlier work this paper cites.
VADER: A Parsimonious Rule-Based Model for Sentiment Analysis of Social Media Text
Hutto, C.; and Gilbert, E. 2014 · 2014
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation
Pennington, J.; Socher, R.; and Manning, C. 2014 · 2014
Earlier work this paper cites.
Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
Bolukbasi, T.; Chang, K.-W.; Zou, J.; Saligrama, V.; and Kalai, A. 2016 · 2016
Cited alongside, same era.
Equality of opportunity in supervised learning
Hardt, M.; Price, E.; and Srebro, N. 2016 · 2016
Cited alongside, same era.
Pointer Sentinel Mixture Models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2016 · 2016
Cited alongside, same era.
Semantics derived automatically from language corpora contain human-like biases
Caliskan, A.; Bryson, J. J.; and Narayanan, A. 2017 · 2017
Cited alongside, same era.
Learning Gender-Neutral Word Embeddings
Zhao, J.; Zhou, Y.; Li, Z.; Wang, W.; and Chang, K.-W. 2018 · 2018
Cited alongside, same era.
Looking for a Handsome Carpenter! Debiasing GPT-3 Job Advertisements
Borchers, C.; Gala, D.; Gilburt, B.; Oravkin, E.; Bounsi, W.; Asano, Y. M.; and Kirk, H. 2022 · 2022
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, E.; Wang, X.; Dehghani, M.; Brahma, S.; et al. 2022 · 2022
Later among the works it cites.
Auto-Debias: Debiasing Masked Language Models with Automated Biased Prompts
Guo, Y.; Yang, Y.; and Abbasi, A. 2022 · 2022
Later among the works it cites.
Mix and Match: Learning-free Controllable Text Generation using Energy Language Models
Mireshghallah, F.; Goyal, K.; and Berg-Kirkpatrick, T. 2022 · 2022
Later among the works it cites.
Upstream Mitigation Is Not
Steed, R.; Panda, S.; Kobren, A.; and Wick, M. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
May, C.; Wang, A.; Bordia, S.; Bowman, S. R.; and Rudinger, R. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Cited alongside, same era.
How can we know what language models know?
Jiang, Z.; Xu, F. F.; Araki, J.; and Neubig, G. 2020 · 2020
Cited alongside, same era.
Investigating Gender Bias in Language Models Using Causal Mediation Analysis
Vig, J.; Gehrmann, S.; Belinkov, Y.; Qian, S.; Nevo, D.; Singer, Y.; and Shieber, S. 2020 · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M.; Gebru, T.; McMillan-Major, A.; and Shmitchell, S. 2021 · 2021
Cited alongside, same era.
Towards Understanding and Mitigating Social Biases in Language Models
Liang, P. P.; Wu, C.; Morency, L.-P.; and Salakhutdinov, R. 2021 · 2021
Cited alongside, same era.
Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus-Based Bias in NLP
Schick, T.; Udupa, S.; and Schütze, H. 2021 · 2021
Cited alongside, same era.
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Biderman, S.; Schoelkopf, H.; Anthony, Q.; Bradley, H.; O’Brien, K.; Hallahan, E.; Khan, M. A.; Purohit, S.; Prashanth, U. S.; Raff, E.; Skowron, A.; Sutawika, L.; and van der Wal, O. 2023 · 2023
Closest in time.
Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models
Ferrara, E. 2023 · 2023
Closest in time.
Detoxifying Text with MaRCo: Controllable Revision with Experts and Anti-Experts
Hallinan, S.; Liu, A.; Choi, Y.; and Sap, M. 2023 · 2023
Closest in time.
Large Language Models are Zero-Shot Reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2023 · 2023
Closest in time.
When Do Pre-Training Biases Propagate to Downstream Tasks? A Case Study in Text Summarization
Ladhak, F.; Durmus, E.; Suzgun, M.; Zhang, T.; Jurafsky, D.; McKeown, K.; and Hashimoto, T. 2023 · 2023
Closest in time.
In-Depth Look at Word Filling Societal Bias Measures
Pikuliak, M.; Beňová, I.; and Bachratý, V. 2023 · 2023
Closest in time.
Prompting GPT-3 To Be Reliable
Si, C.; Gan, Z.; Yang, Z.; Wang, S.; Wang, J.; Boyd-Graber, J.; and Wang, L. 2023 · 2023
Closest in time.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Ferrer, C. C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goyal, N.; Hartshorn, A.; Hosseini, S.; Hou, R.; Inan, H.; Kardas, M.; Kerkez, V.; Khabsa, M.; Kloumann, I.; Korenev, A.; Koura, P. S.; Lachaux, M.-A.; Lavril, T.; Lee, J.; Liskovich, D.; Lu, Y.; Mao, Y.; Martinet, X.; Mihaylov, T.; Mishra, P.; Molybog, I.; Nie, Y.; Poulton, A.; Reizenstein, J.; Rungta, R.; Saladi, K.; Schelten, A.; Silva, R.; Smith, E. M.; Subramanian, R.; Tan, X. E.; Tang, B.; Taylor, R.; Williams, A.; Kuan, J. X.; Xu, P.; Yan, Z.; Zarov, I.; Zhang, Y.; Fan, A.; Kambadur, M.; Narang, S.; Rodriguez, A.; Stojnic, R.; Edunov, S.; and Scialom, T. 2023 · 2023
Closest in time.