Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) reproduce social biases, yet prevailing evaluations score models in isolation, obscuring how biases persist across families and releases.
Divergence measures based on the shannon entropy
Jianhua Lin · 1991
Earlier work this paper cites.
Discrimination in online ad delivery
Latanya Sweeney · 2013
Earlier work this paper cites.
Distance measures in author profiling
Mirco Kocher and Jacques Savoy · 2017
Earlier work this paper cites.
Pivoted document length normalization
Amit Singhal, Chris Buckley, and Manclar Mitra · 2017
Earlier work this paper cites.
Older adults perceptions of technology and barriers to interacting with tablet computers: a focus group study
Eleftheria Vaportzis, Maria Giatsi Clausen, and Alan J Gow · 2017
Earlier work this paper cites.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
TweetEval: Unified benchmark and comparative evaluation for tweet classification
Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa Anke, and Leonardo Neves · 2020
Earlier work this paper cites.
UNQOVERing stereotyping biases via underspecified questions
Tao Li, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Vivek Srikumar · 2020
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman · 2020
Earlier work this paper cites.
Similarity analysis of contextual word representation models
John Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass · 2020
Earlier work this paper cites.
Measuring biases of word embeddings: What similarity measures and descriptive statistics to use?
Hossein Azarpanah and Mohsen Farhadloo · 2021
Earlier work this paper cites.
Modeldiff: Testing-based dnn similarity comparison for model reuse detection
Yuanchun Li, Ziqi Zhang, Bingyan Liu, Ziyue Yang, and Yunxin Liu · 2021
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2021
Earlier work this paper cites.
Are you stealing my model? sample correlation for fingerprinting deep neural networks
Jiyang Guan, Jian Liang, and Ran He · 2022
Earlier work this paper cites.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman · 2022
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Fairness and bias in artificial intelligence: A brief survey of sources, impacts, and mitigation strategies
Emilio Ferrara · 2023
Cited alongside, same era.
Large language model (llm) bias index–llmbi
Abiodun Finbarrs Oketunji, Muhammad Anas, and Deepthi Saina · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Evaluating interfaced llm bias
Kai-Ching Yeh, Jou-An Chi, Da-Chen Lian, and Shu-Kai Hsieh · 2023
Lin Shi, Chiyu Ma, Wenhua Liang, Weicheng Ma, and Soroush Vosoughi · 2024
Closest in time.
Large language models are inconsistent and biased evaluators
Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara · 2024
Closest in time.
Beats: Bias evaluation and assessment test suite for large language models
Alok Abhishek, Lisa Erickson, and Tushar Bandopadhyay · 2025
Closest in time.
Certifying counterfactual bias in llms
Isha Chaudhary, Qian Hu, Manoj Kumar, Morteza Ziyadi, Rahul Gupta, and Gagandeep Singh · 2025
Closest in time.
Polyrating: A cost-effective and bias-aware rating system for llm evaluation
Jasper Dekoninck, Maximilian Baader, and Martin Vechev · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gptbias: A comprehensive framework for evaluating bias in large language models
Jiaxu Zhao, Meng Fang, Shirui Pan, Wenpeng Yin, and Mykola Pechenizkiy · 2023
Cited alongside, same era.
Llm diagnostic toolkit: Evaluating llms for ethical issues
Mehdi Bahrami, Ryosuke Sonoda, and Ramya Srinivasan · 2024
Cited alongside, same era.
Bias and unfairness in information retrieval systems: New challenges in the llm era
Sunhao Dai, Chen Xu, Shicheng Xu, Liang Pang, Zhenhua Dong, and Jun Xu · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Cited alongside, same era.
Bias and fairness in large language models: A survey
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed · 2024
Cited alongside, same era.
Calm: A multi-task benchmark for comprehensive assessment of language model bias
Vipul Gupta, Pranav Narayanan Venkit, Hugo Laurençon, Shomir Wilson, and Rebecca J Passonneau · 2024
Cited alongside, same era.
Proflingo: A fingerprinting-based copyright protection scheme for large language models
Heng Jin, Chaoyu Zhang, Shanghao Shi, Wenjing Lou, and Y Thomas Hou · 2024
Cited alongside, same era.
Closest in time.
Fairmt-bench: Benchmarking fairness for multi-turn dialogue in conversational llms
Zhiting Fan, Ruizhe Chen, Tianxiang Hu, and Zuozhu Liu · 2025
Closest in time.
Gemini api models
Google AI Developers · 2025
Closest in time.
Similarity of neural architectures using adversarial attack transferability
Jaehui Hwang, Dongyoon Han, Byeongho Heo, Song Park, Sanghyuk Chun, and Jong-Seok Lee · 2025
Closest in time.
Similarity of neural network models: A survey of functional and representational measures
Max Klabunde, Tobias Schumacher, Markus Strohmaier, and Florian Lemmerich · 2025
Closest in time.
Investigating bias in LLM-based bias detection: Disparities between LLMs and human perception
Luyang Lin, Lingzhi Wang, Jinsong Guo, and Kam-Fai Wong · 2025
Closest in time.
Gpt-4o mini: advancing cost-efficient intelligence
OpenAI · 2025
Closest in time.
Introducing GPT-5
OpenAI · 2025
Closest in time.
Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, et al · 2025
Closest in time.
Ceb: Compositional evaluation benchmark for fairness in large language models
Song Wang, Peng Wang, Tong Zhou, Yushun Dong, Zhen Tan, and Jundong Li · 2025
Closest in time.
Justice or prejudice? quantifying biases in llm-as-a-judge
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, et al · 2025
Closest in time.