Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are increasingly being deployed in high-stakes applications like hiring, yet their potential for unfair decision-making remains understudied in generative and retrieval settings.
Intrinsic bias metrics do not correlate with application bias
Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya, and Adam Lopez. 2021 · 1940
Earlier work this paper cites.
Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel
JP Kincaid. 1975 · 1975
Earlier work this paper cites.
Controlling the false discovery rate: a practical and powerful approach to multiple testing
Yoav Benjamini and Yosef Hochberg. 1995 · 1995
Earlier work this paper cites.
Multiple significance tests: the bonferroni method
J Martin Bland and Douglas G Altman. 1995 · 1995
Earlier work this paper cites.
Measuring and reducing gendered correlations in pre-trained models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2021 · 2010
Earlier work this paper cites.
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. 2012 · 2012
Earlier work this paper cites.
The problem with bias: Allocative versus representational harms in machine learning
Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017 · 2017
Earlier work this paper cites.
Counterfactual fairness
Matt Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva. 2017 · 2017
Earlier work this paper cites.
Insight - Amazon scraps secret AI recruiting tool that showed bias against women
Jeffrey Dastin. 2018 · 2018
Earlier work this paper cites.
Fairness of extractive text summarization
Anurag Shandilya, Kripabandhu Ghosh, and Saptarshi Ghosh. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
Understanding undesirable word embedding associations
Kawin Ethayarajh, David Duvenaud, and Graeme Hirst. 2019 · 2019
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019 · 2019
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Earlier work this paper cites.
Mitigating bias in algorithmic hiring: evaluating claims and practices
Manish Raghavan, Solon Barocas, Jon Kleinberg, and Karen Levy. 2020 · 2020
Earlier work this paper cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Earlier work this paper cites.
What does it mean to ’solve’ the problem of discrimination in hiring? social, technical and legal perspectives from the uk on automated hiring systems
Javier Sánchez-Monedero, Lina Dencik, and Lilian Edwards. 2020 · 2020
Cited alongside, same era.
Large language models associate muslims with violence
Abubakar Abid, Maheen Farooqi, and James Zou. 2021 · 2021
Cited alongside, same era.
Resume dataset
Snehaan Bhawal. 2021 · 2021
Cited alongside, same era.
Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021 · 2021
Cited alongside, same era.
Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models
Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Frederic Dreyer, Aleksandar Shtedritski, and Yuki Asano. 2021 · 2021
Cited alongside, same era.
Command-r
Cohere. 2024 · 2024
Later among the works it cites.
Stop! in the name of flaws: Disentangling personal names and sociodemographic attributes in NLP
Vagrant Gautam, Arjun Subramonian, Anne Lauscher, and Os Keyes. 2024 · 2024
Later among the works it cites.
Identifying and improving disability bias in gpt-based resume screening
Kate Glazko, Yusuf Mohammed, Ben Kosa, Venkatesh Potluri, and Jennifer Mankoff. 2024 · 2024
Later among the works it cites.
What’s in a name? auditing large language models for race and gender bias
Amit Haim, Alejandro Salinas, and Julian Nyarko. 2024 · 2024
Later among the works it cites.
6 ways to automate Recruit CRM with Zapier
Hannah Herman. 2024 · 2024
Later among the works it cites.
Using AI to streamline the recruiting process
Humanly. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021 · 2021
Cited alongside, same era.
A framework for understanding sources of harm throughout the machine learning life cycle
Harini Suresh and John Guttag. 2021 · 2021
Cited alongside, same era.
On the intrinsic and extrinsic fairness evaluation metrics for contextualized language representations
Yang Trista Cao, Yada Pruksachatkun, Kai-Wei Chang, Rahul Gupta, Varun Kumar, Jwala Dhamala, and Aram Galstyan. 2022 · 2022
Cited alongside, same era.
Nichelle and nancy: The influence of demographic attributes and tokenization length on first name biases
Haozhe An and Rachel Rudinger. 2023 · 2023
Cited alongside, same era.
Marked personas: Using natural language prompts to measure stereotypes in language models
Myra Cheng, Esin Durmus, and Dan Jurafsky. 2023 · 2023
Cited alongside, same era.
What’s in my big data?
Yanai Elazar, Akshita Bhagia, Ian Helgi Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Evan Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, Hanna Hajishirzi, Noah A. Smith, and Jesse Dodge. 2023 · 2023
Cited alongside, same era.
"i wouldn’t say offensive but…": Disability-centered perspectives on large language models
Vinitha Gadiraju, Shaun Kane, Sunipa Dev, Alex Taylor, Ding Wang, Emily Denton, and Robin Brewer. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
Later among the works it cites.
Meta llama 3
Meta AI. 2024 · 2024
Later among the works it cites.
“you gotta be a doctor, lin” : An investigation of name-based bias of large language models in employment recommendations
Huy Nghiem, John Prindle, Jieyu Zhao, and Hal Daumé Iii. 2024 · 2024
Later among the works it cites.
Bias in news summarization: Measures, pitfalls and corpora
Julius Steen and Katja Markert. 2024 · 2024
Later among the works it cites.
Do large language models rank fairly? an empirical study on the fairness of LLMs as rankers
Yuan Wang, Xuyang Wu, Hsin-Tai Wu, Zhiqiang Tao, and Yi Fang. 2024 · 2024
Later among the works it cites.
Gender, race, and intersectional bias in resume screening via language model retrieval
Kyra Wilson and Aylin Caliskan. 2024 · 2024
Later among the works it cites.
A study of implicit ranking unfairness in large language models
Chen Xu, Wenjie Wang, Yuxin Li, Liang Pang, Jun Xu, and Tat-Seng Chua. 2024 · 2024
Later among the works it cites.
Openai’s gpt is a recruiter’s dream tool. tests show there’s racial bias
Leon Yin, Davey Alba, and Leonardo Nicoletti. 2024 · 2024
Later among the works it cites.
Fair abstractive summarization of diverse perspectives
Yusen Zhang, Nan Zhang, Yixin Liu, Alexander Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao, Dragomir Radev, Kathleen McKeown, and Rui Zhang. 2024 · 2024
Later among the works it cites.
How ai is changing recruitment
Boston Consulting Group. 2025 · 2025
Closest in time.
Improving fairness of large language models in multi-document summarization
Haoyuan Li, Rui Zhang, and Snigdha Chaturvedi. 2025 · 2025
Closest in time.