Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are prone to inheriting and amplifying societal biases embedded within their training data, potentially reinforcing harmful stereotypes related to gender, occupation, and other sensitive categories.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T.; Chang, K.-W.; Zou, J. Y.; Saligrama, V.; and Kalai, A. T. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Caliskan, A.; Bryson, J. J.; and Narayanan, A. 2017 · 2017
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L.; Biderman, S.; Black, S.; Golding, L.; Hoppe, T.; Foster, C.; Phang, J.; He, H.; Thite, A.; Nabeshima, N.; et al. 2020 · 2020
Earlier work this paper cites.
Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Blodgett, S. L.; Lopez, G.; Olteanu, A.; Sim, R.; and Wallach, H. 2021 · 2021
Earlier work this paper cites.
Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models
Kirk, H. R.; Jun, Y.; Volpin, F.; Iqbal, H.; Benussi, E.; Dreyer, F.; Shtedritski, A.; and Asano, Y. 2021 · 2021
Earlier work this paper cites.
The Falcon Series of Open Language Models
Almazrouei, E.; Alobeidli, H.; Alshamsi, A.; Cappelli, A.; Cojocaru, R.; Debbah, M.; Étienne Goffinet; Hesslow, D.; Launay, J.; Malartic, Q.; Mazzotta, D.; Noune, B.; Pannier, B.; and Penedo, G. 2023 · 2023
Earlier work this paper cites.
Distilling Large Language Models using Skill-Occupation Graph Context for HR-Related Tasks
Pezeshkpour, P.; Iso, H.; Lake, T.; Bhutani, N.; and Hruschka, E. 2023 · 2023
Earlier work this paper cites.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Wang, B.; Chen, W.; Pei, H.; Xie, C.; Kang, M.; Zhang, C.; Xu, C.; Xiong, Z.; Dutta, R.; Schaeffer, R.; et al. 2023 · 2023
Cited alongside, same era.
Large language models are not robust multiple choice selectors
Zheng, C.; Zhou, H.; Meng, F.; Zhou, J.; and Huang, M. 2023 · 2023
Cited alongside, same era.
GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Black, S.; Gao, L.; Wang, P.; Leahy, C.; Biderman, S.; Golding, L.; Hallahan, E.; Anthony, Q.; Hilton, B.; Phang, P.; Frieder, S.; McDonell, K.; Mishkin, P.; Purohit, S.; Thite, A.; Reynolds, L.; and Carlini, N. 2022b · 2024
Cited alongside, same era.
Auditing the Use of Language Models to Guide Hiring Decisions
Gaebler, J. D.; Goel, S.; Huq, A.; and Tambe, P. 2024 · 2024
Cited alongside, same era.
Self-Debiasing Large Language Models: Zero-Shot Recognition and Reduction of Stereotypes
OffsetBias: Leveraging Debiased Data for Tuning Evaluators
Park, J.; Jwa, S.; Ren, M.; Kim, D.; and Choi, S. 2024 · 2024
Closest in time.
Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence
The White House. 2023 · 2024
Closest in time.
Current Population Survey: Table 11. Employed persons by detailed occupation, sex, race, and Hispanic or Latino ethnicity
U.S. Bureau of Labor Statistics. 2024 · 2024
Closest in time.
Undesirable biases in NLP: Addressing challenges of measurement
Van der Wal, O.; Bachmann, D.; Leidinger, A.; van Maanen, L.; Zuidema, W.; and Schulz, K. 2024 · 2024
Closest in time.
Wan, Y.; and Chang, K.-W. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gallegos, I. O.; Rossi, R. A.; Barrow, J.; Tanjim, M. M.; Yu, T.; Deilamsalehy, H.; Zhang, R.; Kim, S.; and Dernoncourt, F. 2024 · 2024
Cited alongside, same era.
COBIAS: Contextual Reliability in Bias Assessment
Govil, P.; Jain, H.; Bonagiri, V. K.; Chadha, A.; Kumaraguru, P.; Gaur, M.; and Dey, S. 2024 · 2024
Cited alongside, same era.
GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Black, S.; Biderman, S.; Hallahan, E.; Anthony, Q.; Gao, L.; Golding, L.; He, H.; Leahy, C.; McDonell, K.; Phang, J.; Pieler, M.; Prashanth, U. S.; Purohit, S.; Reynolds, L.; Tow, J.; Wang, B.; and Weinbach, S. 2022a
Cited in the paper.
BBQ: A hand-built bias benchmark for question answering
Parrish, A.; Chen, A.; Nangia, N.; Padmakumar, V.; Phang, J.; Thompson, J.; Htut, P. M.; and Bowman, S. 2022 · 2086
Closest in time.