Fetching the paper…
Reading the bibliography…
The growing integration of Large Language Models (LLMs) into critical societal domains has raised concerns about embedded biases that can perpetuate stereotypes and undermine fairness.
‘‘Language models are few-shot learners’’
Tom Brown et al · 1901
Earlier work this paper cites.
‘‘Language models are few-shot learners’’
Tom Brown et al · 1901
Earlier work this paper cites.
‘‘CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models’’
Nikita Nangia, Clara Vania, Rasika Bhalerao and Samuel Bowman · 1967
Earlier work this paper cites.
‘‘CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models’’
Nikita Nangia, Clara Vania, Rasika Bhalerao and Samuel Bowman · 1967
Earlier work this paper cites.
‘‘The measurement of observer agreement for categorical data’’
J Landis and Gary Koch · 1977
Earlier work this paper cites.
‘‘The measurement of observer agreement for categorical data’’
J Landis and Gary Koch · 1977
Earlier work this paper cites.
‘‘Gender trouble’’
Judith Butler · 2002
Earlier work this paper cites.
‘‘Gender trouble’’
Judith Butler · 2002
Earlier work this paper cites.
‘‘Mixed methods assessment of the influence of demographics on medical advice of ChatGPT’’
Katerina Andreadis et al · 2009
Earlier work this paper cites.
‘‘Stigma: Notes on the management of spoiled identity’’
Erving Goffman · 2009
Earlier work this paper cites.
‘‘Mixed methods assessment of the influence of demographics on medical advice of ChatGPT’’
Katerina Andreadis et al · 2009
Earlier work this paper cites.
‘‘Stigma: Notes on the management of spoiled identity’’
Erving Goffman · 2009
Earlier work this paper cites.
‘‘Demarginalizing the intersection of race and sex: A black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics’’
Kimberlé Crenshaw · 2013
Earlier work this paper cites.
‘‘Demarginalizing the intersection of race and sex: A black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics’’
Kimberlé Crenshaw · 2013
Earlier work this paper cites.
‘‘Racial formation in the United States’’
Michael Omi and Howard Winant · 2014
Earlier work this paper cites.
‘‘Racial formation in the United States’’
Michael Omi and Howard Winant · 2014
Earlier work this paper cites.
‘‘Semantics derived automatically from language corpora contain human-like biases’’
Aylin Caliskan, Joanna Bryson and Arvind Narayanan · 2017
Earlier work this paper cites.
‘‘Semantics derived automatically from language corpora contain human-like biases’’
Aylin Caliskan, Joanna Bryson and Arvind Narayanan · 2017
Earlier work this paper cites.
‘‘Identifying and reducing gender bias in word-level language models’’
Shikha Bordia and Samuel Bowman · 2019
Earlier work this paper cites.
‘‘Measuring Bias in Contextualized Word Representations’’
Keita Kurita et al · 2019
Earlier work this paper cites.
‘‘On Measuring Social Biases in Sentence Encoders’’
Chandler May et al · 2019
Earlier work this paper cites.
‘‘Identifying and reducing gender bias in word-level language models’’
Shikha Bordia and Samuel Bowman · 2019
Earlier work this paper cites.
‘‘Measuring Bias in Contextualized Word Representations’’
Keita Kurita et al · 2019
Earlier work this paper cites.
‘‘On Measuring Social Biases in Sentence Encoders’’
Chandler May et al · 2019
Earlier work this paper cites.
‘‘The State and Fate of Linguistic Diversity and Inclusion in the NLP World’’
Pratik Joshi et al · 2020
Earlier work this paper cites.
‘‘The State and Fate of Linguistic Diversity and Inclusion in the NLP World’’
Pratik Joshi et al · 2020
Earlier work this paper cites.
‘‘Persistent anti-muslim bias in large language models’’
Abubakar Abid, Maheen Farooqi and James Zou · 2021
Earlier work this paper cites.
‘‘Bold: Dataset and metrics for measuring biases in open-ended language generation’’
Jwala Dhamala et al · 2021
Earlier work this paper cites.
‘‘Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases’’
Wei Guo and Aylin Caliskan · 2021
Earlier work this paper cites.
‘‘Five sources of bias in natural language processing’’
Dirk Hovy and Shrimai Prabhumoye · 2021
Earlier work this paper cites.
‘‘A survey on bias and fairness in machine learning’’
Ninareh Mehrabi et al · 2021
Earlier work this paper cites.
‘‘StereoSet: Measuring stereotypical bias in pretrained language models’’
Moin Nadeem, Anna Bethke and Siva Reddy · 2021
Earlier work this paper cites.
‘‘HONEST: Measuring Hurtful Sentence Completion in Language Models’’
Debora Nozza, Federico Bianchi and Dirk Hovy · 2021
Earlier work this paper cites.
‘‘Persistent anti-muslim bias in large language models’’
Abubakar Abid, Maheen Farooqi and James Zou · 2021
Earlier work this paper cites.
‘‘Bold: Dataset and metrics for measuring biases in open-ended language generation’’
Jwala Dhamala et al · 2021
Earlier work this paper cites.
‘‘Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases’’
Wei Guo and Aylin Caliskan · 2021
Earlier work this paper cites.
‘‘Five sources of bias in natural language processing’’
Dirk Hovy and Shrimai Prabhumoye · 2021
Earlier work this paper cites.
‘‘A survey on bias and fairness in machine learning’’
Ninareh Mehrabi et al · 2021
Earlier work this paper cites.
‘‘StereoSet: Measuring stereotypical bias in pretrained language models’’
Moin Nadeem, Anna Bethke and Siva Reddy · 2021
Earlier work this paper cites.
‘‘HONEST: Measuring Hurtful Sentence Completion in Language Models’’
Debora Nozza, Federico Bianchi and Dirk Hovy · 2021
Earlier work this paper cites.
‘‘Evaluating the feasibility of ChatGPT in healthcare: an analysis of multiple clinical and research scenarios’’
Marco Cascella, Jonathan Montomoli, Valentina Bellini and Elena Bignami · 2023
Earlier work this paper cites.
‘‘Should ChatGPT be biased? Challenges and risks of bias in large language models’’
Emilio Ferrara · 2023
Earlier work this paper cites.
‘‘Llama guard: Llm-based input-output safeguard for human-ai conversations’’
Hakan Inan et al · 2023
Cited alongside, same era.
‘‘Gender bias and stereotypes in large language models’’
Hadas Kotek, Rikker Dockum and David Sun · 2023
Cited alongside, same era.
‘‘Holistic evaluation of language models’’
Percy Liang et al · 2023
Cited alongside, same era.
‘‘Biases in large language models: origins, inventory, and discussion’’
Roberto Navigli, Simone Conia and Björn Ross · 2023
Cited alongside, same era.
‘‘Neural machine translation for low-resource languages: A survey’’
Surangika Ranathunga et al · 2023
Cited alongside, same era.
‘‘Evaluating interfaced llm bias’’
Kai-Ching Yeh, Jou-An Chi, Da-Chen Lian and Shu-Kai Hsieh · 2023
‘‘Large Language Models are not Fair Evaluators’’
Peiyi Wang et al · 2024
Later among the works it cites.
‘‘Self-preference bias in llm-as-a-judge’’
Koki Wataoka, Tsubasa Takahashi and Ryokan Ri · 2024
Later among the works it cites.
‘‘Disparities in seizure outcomes revealed by large language models’’
Kevin Xie et al · 2024
Later among the works it cites.
‘‘Jailbreak attacks and defenses against large language models: A survey’’
Sibo Yi et al · 2024
Later among the works it cites.
‘‘Ultramedical: Building specialized generalists in biomedicine’’
Kaiyan Zhang et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
‘‘Low-Resource Languages Jailbreak GPT-4’’
Zheng Yong, Cristina Menghini and Stephen Bach · 2023
Cited alongside, same era.
‘‘Judging llm-as-a-judge with mt-bench and chatbot arena’’
Lianmin Zheng et al · 2023
Cited alongside, same era.
‘‘Evaluating the feasibility of ChatGPT in healthcare: an analysis of multiple clinical and research scenarios’’
Marco Cascella, Jonathan Montomoli, Valentina Bellini and Elena Bignami · 2023
Cited alongside, same era.
‘‘Should ChatGPT be biased? Challenges and risks of bias in large language models’’
Emilio Ferrara · 2023
Cited alongside, same era.
‘‘Llama guard: Llm-based input-output safeguard for human-ai conversations’’
Hakan Inan et al · 2023
Cited alongside, same era.
‘‘Gender bias and stereotypes in large language models’’
Hadas Kotek, Rikker Dockum and David Sun · 2023
Cited alongside, same era.
Marah Abdin et al · 2024
Later among the works it cites.
‘‘Understanding intrinsic socioeconomic biases in large language models’’
Mina Arzaghi, Florian Carichon and Golnoosh Farnadi · 2024
Later among the works it cites.
‘‘Measuring implicit bias in explicitly unbiased large language models’’
Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky and Thomas Griffiths · 2024
Later among the works it cites.
‘‘Are large language models really bias-free? jailbreak prompts for assessing adversarial robustness to bias elicitation’’
Riccardo Cantini, Giada Cosenza, Alessio Orsino and Domenico Talia · 2024
Later among the works it cites.
‘‘A survey on evaluation of large language models’’
Yupeng Chang et al · 2024
Later among the works it cites.
‘‘(A)I am not a lawyer, but…: engaging legal experts towards responsible LLM policies for legal advice’’
Inyoung Cheong et al · 2024
Later among the works it cites.
‘‘Med42-v2: A Suite of Clinical LLMs’’
Clément Christophe et al · 2024
Later among the works it cites.
‘‘Deepseek-v3 technical report’’
DeepSeek-AI et al · 2024
Later among the works it cites.
‘‘Bells: A framework towards future proof benchmarks for the evaluation of llm safeguards’’
Diego Dorn, Alexandre Variengien, Charbel-RaphaÃĢl Segerie and Vincent Corruble · 2024
Later among the works it cites.
‘‘Bias and fairness in large language models: A survey’’
Isabel Gallegos et al · 2024
Later among the works it cites.
‘‘Gemma 2: Improving open language models at a practical size’’
Gemma Team et al · 2024
Later among the works it cites.
‘‘The llama 3 herd of models’’
Aaron Grattafiori et al · 2024
Later among the works it cites.
‘‘ChatGPT in education: A blessing or a curse? A qualitative study exploring early adopters’ utilization and perceptions’’
Reza Hadi Mogavi et al · 2024
Later among the works it cites.
‘‘GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models’’
Haibo Jin et al · 2024
Later among the works it cites.
‘‘Investigating Subtler Biases in LLMs: Ageism, Beauty, Institutional, and Nationality Bias in Generative Models’’
Mahammed Kamruzzaman, Md Shovon and Gene Kim · 2024
Later among the works it cites.
‘‘Prometheus: Inducing Fine-Grained Evaluation Capability in Language Models’’
Seungone Kim et al · 2024
Later among the works it cites.
‘‘Generative Judge for Evaluating Alignment’’
Junlong Li et al · 2024
Later among the works it cites.
‘‘AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models’’
Xiaogeng Liu, Nan Xu, Muhao Chen and Chaowei Xiao · 2024
Later among the works it cites.
‘‘Social Bias Probing: Fairness Benchmarking for Language Models’’
Marta Manerba, Karolina Stanczak, Riccardo Guidotti and Isabelle Augenstein · 2024
Later among the works it cites.
‘‘Tree of attacks: Jailbreaking black-box llms automatically’’
Anay Mehrotra et al · 2024
Later among the works it cites.
‘‘A survey of small language models’’
Chien Nguyen et al · 2024
Later among the works it cites.
‘‘What’s in a name? Auditing large language models for race and gender bias’’
Alejandro Salinas, Amit Haim and Julian Nyarko · 2024
Later among the works it cites.
‘‘ALERT: A Comprehensive Benchmark for Assessing Large Language Models’ Safety through Red Teaming’’
Simone Tedeschi et al · 2024
Later among the works it cites.
‘‘On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective’’
Jindong Wang et al · 2024
Later among the works it cites.
‘‘Large Language Models are not Fair Evaluators’’
Peiyi Wang et al · 2024
Later among the works it cites.
‘‘Self-preference bias in llm-as-a-judge’’
Koki Wataoka, Tsubasa Takahashi and Ryokan Ri · 2024
Later among the works it cites.
‘‘Disparities in seizure outcomes revealed by large language models’’
Kevin Xie et al · 2024
Later among the works it cites.
‘‘Jailbreak attacks and defenses against large language models: A survey’’
Sibo Yi et al · 2024
Later among the works it cites.
‘‘Ultramedical: Building specialized generalists in biomedicine’’
Kaiyan Zhang et al · 2024
Later among the works it cites.
‘‘Jailbreaking black box large language models in twenty queries’’
Patrick Chao et al · 2025
Closest in time.
‘‘Evaluating and addressing demographic disparities in medical large language models: a systematic review’’
Mahmud Omar et al · 2025
Closest in time.
‘‘LLMs Reproduce Stereotypes of Sexual and Gender Minorities’’
Ruby Ostrow and Adam Lopez · 2025
Closest in time.
‘‘JudgeLM: Fine-tuned Large Language Models are Scalable Judges’’
Lianghui Zhu, Xinggang Wang and Xinlong Wang · 2025
Closest in time.
‘‘Jailbreaking black box large language models in twenty queries’’
Patrick Chao et al · 2025
Closest in time.
‘‘Evaluating and addressing demographic disparities in medical large language models: a systematic review’’
Mahmud Omar et al · 2025
Closest in time.
‘‘LLMs Reproduce Stereotypes of Sexual and Gender Minorities’’
Ruby Ostrow and Adam Lopez · 2025
Closest in time.
‘‘JudgeLM: Fine-tuned Large Language Models are Scalable Judges’’
Lianghui Zhu, Xinggang Wang and Xinlong Wang · 2025
Closest in time.