Fetching the paper…
Reading the bibliography…
Large language models (LLMs), despite their remarkable capabilities, are susceptible to generating biased and discriminatory responses.
Limitations of the application of fourfold table analysis to hospital data
Joseph Berkson · 1946
Earlier work this paper cites.
Sample selection bias as a specification error
James J Heckman · 1979
Earlier work this paper cites.
Varieties of selection bias
James Heckman · 1990
Earlier work this paper cites.
Models for sample selection bias
Christopher Winship and Robert D Mare · 1992
Earlier work this paper cites.
Causation, Prediction, and Search
Peter Spirtes, Clark Glymour, and Richard Scheines · 1993
Earlier work this paper cites.
Causal inference in the presence of latent variables and selection bias
Peter Spirtes, Christopher Meek, and Thomas Richardson · 1995
Earlier work this paper cites.
Interventions and causal inference
Frederick Eberhardt and Richard Scheines · 2007
Earlier work this paper cites.
On the completeness of orientation rules for causal discovery in the presence of latent confounders and selection bias
Jiji Zhang · 2008
Earlier work this paper cites.
Building classifiers with independency constraints
Toon Calders, Faisal Kamiran, and Mykola Pechenizkiy · 2009
Earlier work this paper cites.
Classifying without discriminating
Faisal Kamiran and Toon Calders · 2009
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Controlling selection bias in causal inference
Elias Bareinboim and Judea Pearl · 2012
Earlier work this paper cites.
Discrimination in online ad delivery
Latanya Sweeney · 2013
Earlier work this paper cites.
Recovering causal effects from selection bias
Elias Bareinboim and Jin Tian · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
A statistical framework for fair predictive algorithms
Kristian Lum and James Johndrow · 2016
Earlier work this paper cites.
On the identifiability and estimation of functional causal models in the presence of outcome-dependent selection
Kun Zhang, Jiji Zhang, Biwei Huang, Bernhard Schölkopf, and Clark Glymour · 2016
Earlier work this paper cites.
Avoiding discrimination through causal reasoning
Niki Kilbertus, Mateo Rojas-Carulla, Giambattista Parascandolo, Moritz Hardt, Dominik Janzing, and Bernhard Schölkopf · 2017
Earlier work this paper cites.
Counterfactual fairness
Matt Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva · 2017
Earlier work this paper cites.
Elements of Causal Inference: Foundations and Learning Algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Fair inference on outcomes
Razieh Nabi and Ilya Shpitser · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2018
Earlier work this paper cites.
Identifying and reducing gender bias in word-level language models
Shikha Bordia and Samuel R Bowman · 2019
Earlier work this paper cites.
Path-specific counterfactual fairness
Silvia Chiappa · 2019
Earlier work this paper cites.
Identification of causal effects in the presence of selection bias
Juan D Correa, Jin Tian, and Elias Bareinboim · 2019
Earlier work this paper cites.
Hila Gonen and Yoav Goldberg · 2019
Cited alongside, same era.
Learning optimal fair policies
Razieh Nabi, Daniel Malinsky, and Ilya Shpitser · 2019
Cited alongside, same era.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng · 2019
Cited alongside, same era.
General transportability of soft interventions: Completeness results
Juan Correa and Elias Bareinboim · 2020
Cited alongside, same era.
Causal modeling for fairness in dynamical systems
Elliot Creager, David Madras, Toniann Pitassi, and Richard Zemel · 2020
Cited alongside, same era.
On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning
Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang · 2022
Later among the works it cites.
Prompting gpt-3 to be reliable
Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, and Lijuan Wang · 2022
Later among the works it cites.
On the fairness of causal algorithmic recourse
Julius von Kügelgen, Amir-Hossein Karimi, Umang Bhatt, Isabel Valera, Adrian Weller, and Bernhard Schölkopf · 2022
Later among the works it cites.
Yiwei Wang, Muhao Chen, Wenxuan Zhou, Yujun Cai, Yuxuan Liang, Dayiheng Liu, Baosong Yang, Juncheng Liu, and Bryan Hooi · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Miguel A Hernán and James M Robins · 2020
Cited alongside, same era.
Causal discovery from heterogeneous/nonstationary data
Biwei Huang, Kun Zhang, Jiji Zhang, Joseph Ramsey, Ruben Sanchez-Romero, Clark Glymour, and Bernhard Schölkopf · 2020
Cited alongside, same era.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2020
Cited alongside, same era.
How do fair decisions fare in long-term qualification?
Xueru Zhang, Ruibo Tu, Yang Liu, Mingyan Liu, Hedvig Kjellstrom, Kun Zhang, and Cheng Zhang · 2020
Cited alongside, same era.
Large language models associate muslims with violence
Abubakar Abid, Maheen Farooqi, and James Zou · 2021
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Cited alongside, same era.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner · 2021
Cited alongside, same era.
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
The capacity for moral self-correction in large language models
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al · 2023
Later among the works it cites.
Can large language models infer causation from correlation?
Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff, Mrinmaya Sachan, Rada Mihalcea, Mona Diab, and Bernhard Schölkopf · 2023
Later among the works it cites.
Identifying selection bias from observational data
David Kaltenpoth and Jilles Vreeken · 2023
Later among the works it cites.
Causal reasoning and large language models: Opening a new frontier for causality
Emre Kıcıman, Robert Ness, Amit Sharma, and Chenhao Tan · 2023
Later among the works it cites.
Trustworthy llms: A survey and guideline for evaluating large language models’ alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li · 2023
Later among the works it cites.
In-contextual bias suppression for large language models
Daisuke Oba, Masahiro Kaneko, and Danushka Bollegala · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Chatgpt: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope
Partha Pratim Ray · 2023
Later among the works it cites.
The political biases of chatgpt
David Rozado · 2023
Later among the works it cites.
Evaluating and mitigating discrimination in language model decisions
Alex Tamkin, Amanda Askell, Liane Lovitt, Esin Durmus, Nicholas Joseph, Shauna Kravec, Karina Nguyen, Jared Kaplan, and Deep Ganguli · 2023
Later among the works it cites.
Large language model unlearning
Yuanshun Yao, Xiaojun Xu, and Yang Liu · 2023
Later among the works it cites.
Understanding causality with large language models: Feasibility and opportunities
Cheng Zhang, Stefan Bauer, Paul Bennett, Jiangfeng Gao, Wenbo Gong, Agrin Hilmkil, Joel Jennings, Chao Ma, Tom Minka, Nick Pawlowski, et al · 2023
Later among the works it cites.
Take a step back: Evoking reasoning via abstraction in large language models
Huaixiu Steven Zheng, Swaroop Mishra, Xinyun Chen, Heng-Tze Cheng, Ed H Chi, Quoc V Le, and Denny Zhou · 2023
Later among the works it cites.
Causal-debias: Unifying debiasing in pretrained language models and fine-tuning via causal invariant learning
Fan Zhou, Yuzhou Mao, Liu Yu, Yi Yang, and Ting Zhong · 2023
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku
AI Anthropic · 2024
Closest in time.
Thinking fair and slow: On the efficacy of structured prompts for debiasing language models
Shaz Furniturewala, Surgan Jandial, Abhinav Java, Pragyan Banerjee, Simra Shahid, Sumit Bhatia, and Kokil Jaidka · 2024
Closest in time.
Procedural fairness through decoupling objectionable data generating components
Zeyu Tang, Jialu Wang, Yang Liu, Peter Spirtes, and Kun Zhang · 2024
Closest in time.
Causal prompting: Debiasing large language model prompting based on front-door adjustment
Congzhi Zhang, Linhai Zhang, Jialong Wu, Deyu Zhou, and Yulan He · 2024
Closest in time.
Detecting and identifying selection structure in sequential data
Yujia Zheng, Zeyu Tang, Yiwen Qiu, Bernhard Schölkopf, and Kun Zhang · 2024
Closest in time.
Fairness through difference awareness: Measuring desired group discrimination in LLMs
Angelina Wang, Michelle Phan, Daniel E Ho, and Sanmi Koyejo · 2025
Closest in time.