Fetching the paper…
Reading the bibliography…
Bias in Large Language Models (LLMs) significantly undermines their reliability and fairness.
The socioeconomic gradient and chronic illness and associated risk factors in australia
John D. Glover, Diana M. Hetzel, and Sarah K. Tennant · 2004
Earlier work this paper cites.
Sparse autoencoder
Andrew Ng · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
k-sparse autoencoders, 2014
Alireza Makhzani and Brendan Frey · 2014
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan · 2017
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie J. Cai, James Wexler, Fernanda B. Viégas, and Rory Sayres · 2018
Earlier work this paper cites.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger · 2019
Earlier work this paper cites.
Measuring bias in contextualized word representations
Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel R Bowman, and Rachel Rudinger · 2019
Earlier work this paper cites.
Understanding the origins of bias in word embeddings
Marc-Etienne Brunet, Colleen Alkalay-Houlihan, Ashton Anderson, and Richard Zemel · 2019
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng · 2019
Earlier work this paper cites.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2020
Earlier work this paper cites.
On measuring and mitigating biased inferences of word embeddings
Sunipa Dev, Tao Li, Jeff M Phillips, and Vivek Srikumar · 2020
Earlier work this paper cites.
Reducing sentiment bias in language models via counterfactual evaluation
Po-Sen Huang, Huan Zhang, Ray Jiang, Robert Stanforth, Johannes Welbl, Jack Rae, Vishal Maini, Dani Yogatama, and Pushmeet Kohli · 2020
Earlier work this paper cites.
Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models
Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Frederic Dreyer, Aleksandar Shtedritski, and Yuki Asano · 2021
Earlier work this paper cites.
Measuring biases of word embeddings: What similarity measures and descriptive statistics to use?
Hossein Azarpanah and Mohsen Farhadloo · 2021
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2021
Earlier work this paper cites.
Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases
Wei Guo and Aylin Caliskan · 2021
Earlier work this paper cites.
Quantifying social biases in NLP: A generalization and empirical comparison of extrinsic fairness metrics
Paula Czarnowska, Yogarshi Vyas, and Kashif Shah · 2021
Earlier work this paper cites.
RedditBias: A real-world resource for bias evaluation and debiasing of conversational language models
Soumya Barikeri, Anne Lauscher, Ivan Vulić, and Goran Glavaš · 2021
Earlier work this paper cites.
On measures of biases and harms in NLP
Sunipa Dev, Emily Sheng, Jieyu Zhao, Aubrie Amstutz, Jiao Sun, Yu Hou, Mattie Sanseverino, Jiin Kim, Akihiro Nishi, Nanyun Peng, and Kai-Wei Chang · 2022
Cited alongside, same era.
Unmasking the mask–evaluating social biases in masked language models
Masahiro Kaneko and Danushka Bollegala · 2022
Cited alongside, same era.
Gender bias and stereotypes in large language models
Hadas Kotek, Rikker Dockum, and David Sun · 2023
Cited alongside, same era.
A survey on fairness in large language models
Yingji Li, Mengnan Du, Rui Song, Xin Wang, and Ying Wang · 2023
Cited alongside, same era.
Large language models propagate race-based medicine
Jesutofunmi A Omiye, Jenna C Lester, Simon Spichak, Veronica Rotemberg, and Roxana Daneshjou · 2023
Cited alongside, same era.
Evaluating large language models: A comprehensive survey, 2023
Extracting unlearned information from llms with activation steering, 2024
Atakan Seyitoğlu, Aleksei Kuvshinov, Leo Schwinn, and Stephan Günnemann · 2024
Later among the works it cites.
Efficient training of sparse autoencoders for large language models via layer groups
Davide Ghilardi, Federico Belotti, and Marco Molinari · 2024
Later among the works it cites.
Efficient dictionary learning with switch sparse autoencoders
Anish Mudide, Joshua Engels, Eric J Michaud, Max Tegmark, and Christian Schroeder de Witt · 2024
Later among the works it cites.
Jumping ahead: Improving reconstruction fidelity with jumprelu sparse autoencoders
Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat, Arthur Conmy, Vikrant Varma, János Kramár, and Neel Nanda · 2024
Later among the works it cites.
Interpreting preference models with sparse autoencoders
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Supryadi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, and Deyi Xiong · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models, 2023
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Neuronpedia: Interactive reference and tooling for analyzing neural networks, 2023
Johnny Lin · 2023
Cited alongside, same era.
Llms are biased teachers: Evaluating llm bias in personalized education
Iain Weissburg, Sathvika Anand, Sharon Levy, and Haewon Jeong · 2024
Cited alongside, same era.
Bias and volatility: A statistical framework for evaluating large language model's stereotypes and the associated generation inconsistency
Yiran Liu, Ke Yang, Zehan Qi, Xiao Liu, Yang Yu, and ChengXiang Zhai · 2024
Cited alongside, same era.
Climb: A benchmark of clinical bias in large language models, 2024
Yubo Zhang, Shudi Hou, Mingyu Derek Ma, Wei Wang, Muhao Chen, and Jieyu Zhao · 2024
Cited alongside, same era.
Can sparse autoencoders be used to decompose and interpret steering vectors?
Harry Mayne, Yushi Yang, and Adam Mahdi · 2024
Cited alongside, same era.
Luke R. Smith and Jonas Brinkmann · 2024
Later among the works it cites.
Effectiveness of sparse autoencoder for understanding and removing gender bias in LLMs
Praveen Hegde · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Later among the works it cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Later among the works it cites.
Search-based automatic repair for fairness and accuracy in decision-making software
Max Hort, Jie M. Zhang, Federica Sarro, and Mark Harman · 2024
Later among the works it cites.
A survey on fairness in large language models, 2024
Yingji Li, Mengnan Du, Rui Song, Xin Wang, and Ying Wang · 2024
Later among the works it cites.
Explore spurious correlations at the concept level in language models for text classification
Yuhang Zhou, Paiheng Xu, Xiaoyu Liu, Bang An, Wei Ai, and Furong Huang · 2024
Later among the works it cites.
Towards detecting unanticipated bias in large language models, 2024
Anna Kruspe · 2024
Later among the works it cites.
Evaluation and mitigation of cognitive biases in medical language models
Samuel Schmidgall, Carl Harris, Ime Essien, Daniel Olshvang, Tawsifur Rahman, Ji Woong Kim, Rojin Ziaei, Jason Eshraghian, Peter Abadir, and Rama Chellappa · 2024
Later among the works it cites.
Lg-cav: Train any concept activation vector with language guidance
Qihan Huang, Jie Song, Mengqi Xue, Haofei Zhang, Bingde Hu, Huiqiong Wang, Hao Jiang, Xingen Wang, and Mingli Song · 2024
Later among the works it cites.
Sociodemographic biases in medical decision making by large language models
Mahmud Omar, Shelly Soffer, Reem Agbareia, Nicola Luigi Bragazzi, Donald U Apakama, Carol R Horowitz, Alexander W Charney, Robert Freeman, Benjamin Kummer, Benjamin S Glicksberg, et al · 2025
Closest in time.
Controlling large language models through concept activation vectors
Hanyu Zhang, Xiting Wang, Chengao Li, Xiang Ao, and Qing He · 2025
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupre la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2025
Closest in time.
Writing style matters: An examination of bias and fairness in information retrieval systems
Hongliu Cao · 2025
Closest in time.
Justice or prejudice? quantifying biases in LLM-as-a-judge
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, Nitesh V Chawla, and Xiangliang Zhang · 2025
Closest in time.
Bias detection and fairness in large language models for financial services
Rahul Vats, Shekhar Agrawal, and Srinivasa Chippada · 2025
Closest in time.
Explaining explainability: Recommendations for effective use of concept activation vectors
Angus Nicolson, Lisa Schut, Alison Noble, and Yarin Gal · 2025
Closest in time.
Controlling large language models through concept activation vectors, 2025
Hanyu Zhang, Xiting Wang, Chengao Li, Xiang Ao, and Qing He · 2025
Closest in time.
Edu-values: Towards evaluating the chinese education values of large language models, 2025
Peiyi Zhang, Yazhou Zhang, Bo Wang, Lu Rong, Prayag Tiwari, and Jing Qin · 2025
Closest in time.
Socioeconomic status and mental health — Wikipedia, the free encyclopedia, 2024
Wikipedia contributors · 2025
Closest in time.