Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are known to perpetuate stereotypes and exhibit biases.
Gender schema theory: A cognitive account of sex typing
Sandra Lipsitz Bem. 1981 · 1981
Earlier work this paper cites.
The dip test of unimodality
J. A. Hartigan and P. M. Hartigan. 1985 · 1985
Earlier work this paper cites.
Evidence that gendered wording in job advertisements exists and sustains gender inequality
Danielle Gaucher, Justin Friesen, and Aaron C Kay. 2011 · 2011
Earlier work this paper cites.
Language use in African American communities
Sonja Lanehart, Ayesha M Malik, and SL Lanehart. 2015 · 2015
Earlier work this paper cites.
Skinny-dip: Clustering in a sea of noise
Samuel Maurus and Claudia Plant. 2016 · 2016
Earlier work this paper cites.
Preventing fairness gerrymandering: Auditing and learning for subgroup fairness
Michael Kearns, Seth Neel, Aaron Roth, and Zhiwei Steven Wu. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Syntactic and cognitive issues in investigating gendered coreference
Lauren Ackerman. 2019 · 2019
Earlier work this paper cites.
Toward gender-inclusive coreference resolution
Yang Trista Cao and Hal Daumé III. 2020 · 2020
Earlier work this paper cites.
Detecting gender stereotypes: Lexicon vs. supervised learning methods
Jenna Cryan, Shiliang Tang, Xinyi Zhang, Miriam Metzger, Haitao Zheng, and Ben Y. Zhao. 2020 · 2020
Earlier work this paper cites.
Investigating African-American Vernacular English in transformer-based text generation
Sophie Groenwold, Lily Ou, Aesha Parekh, Samhita Honnavalli, Sharon Levy, Diba Mirza, and William Yang Wang. 2020 · 2020
Earlier work this paper cites.
Gender bias in text: Origin, taxonomy, and implications
Jad Doughman, Wael Khreich, Maya El Gharib, Maha Wiss, and Zahraa Berjawi. 2021 · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, and 1 others. 2021 · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
DExperts: Decoding-time controlled text generation with experts and anti-experts
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
NeuroLogic decoding: (un)supervised neural text generation with predicate logic constraints
Ximing Lu, Peter West, Rowan Zellers, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Theories of “gender” in NLP bias research
Hannah Devinney, Jenny Björklund, and Henrik Björklund. 2022 · 2022
Cited alongside, same era.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, and 1 others. 2023 · 2023
Cited alongside, same era.
Marked personas: Using natural language prompts to measure stereotypes in language models
Myra Cheng, Esin Durmus, and Dan Jurafsky. 2023 · 2023
Cited alongside, same era.
Identifying and adapting transformer-components responsible for gender bias in an English language model
Abhijith Chintam, Rahel Beloch, Willem Zuidema, Michael Hanna, and Oskar van der Wal. 2023 · 2023
Cited alongside, same era.
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda. 2024 · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, and 1 others. 2024 · 2024
Later among the works it cites.
BiasAlert: A plug-and-play tool for social bias detection in LLMs
Zhiting Fan, Ruizhe Chen, Ruiling Xu, and Zuozhu Liu. 2024 · 2024
Later among the works it cites.
Linguistic bias in ChatGPT: Language models reinforce dialect discrimination
Eve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi, Xavier Yin, and Dan Klein. 2024 · 2024
Later among the works it cites.
Granite 3.0 language models
IBM Granite Team. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The capacity for moral self-correction in large language models
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, and 1 others. 2023 · 2023
Cited alongside, same era.
Llama guard: LLM-based input-output safeguard for human-AI conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and 1 others. 2023 · 2023
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Benjamin Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, and 44 others. 2023 · 2023
Cited alongside, same era.
Using ChatGPT to generate gendered language
Shweta Soundararajan, Manuela Nayantara Jeyaraj, and Sarah Jane Delany. 2023 · 2023
Cited alongside, same era.
Linear representations of sentiment in large language models
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, and 1 others. 2023 · 2023
Cited alongside, same era.
Activation addition: Steering language models without optimization
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J Vazquez, Ulisse Mini, and Monte MacDiarmid. 2023 · 2023
Cited alongside, same era.
Language models represent space and time
Wes Gurnee and Max Tegmark. 2024 · 2024
Later among the works it cites.
Evaluating gender bias in large language models via chain-of-thought prompting
Masahiro Kaneko, Danushka Bollegala, Naoaki Okazaki, and Timothy Baldwin. 2024 · 2024
Later among the works it cites.
RewardBench: Evaluating reward models for language modeling
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, and 1 others. 2024 · 2024
Later among the works it cites.
The geometry of truth: Emergent linear structure in large language model representations of true/false datasets
Samuel Marks and Max Tegmark. 2024 · 2024
Later among the works it cites.
Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, and 1 others. 2024 · 2024
Later among the works it cites.
A trip towards fairness: Bias and de-biasing in large language models
Leonardo Ranaldi, Elena Sofia Ruzzetti, Davide Venditti, Dario Onorati, and Fabio Massimo Zanzotto. 2024 · 2024
Later among the works it cites.
Steering llama 2 via contrastive activation addition
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner. 2024 · 2024
Later among the works it cites.
Programming refusal with conditional activation steering
Bruce W. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Erik Miehling, Pierre Dognin, Manish Nagireddy, and Amit Dhurandhar. 2025 · 2025
Closest in time.
Rejected dialects: Biases against african american language in reward models
Joel Mire, Zubin Trivadi Aysola, Daniel Chechelnitsky, Nicholas Deas, Chrysoula Zerva, and Maarten Sap. 2025 · 2025
Closest in time.