Fetching the paper…
Reading the bibliography…
Current datasets for unwanted social bias auditing are limited to studying protected demographic features such as race and gender.
The Curious Case of Neural Text Degeneration
Holtzman, A.; Buys, J.; Forbes, M.; and Choi, Y. 2019 · 1904
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Nadeem, M.; Bethke, A.; and Reddy, S. 2020 · 2004
Earlier work this paper cites.
Understanding searches better than ever before
Nayak, P. 2019 · 2019
Earlier work this paper cites.
Language (Technology) is Power: A Critical Survey of “Bias” in NLP
Blodgett, S. L.; Barocas, S.; Daumé III, H.; and Wallach, H. 2020 · 2020
Earlier work this paper cites.
Social Chemistry 101: Learning to Reason about Social and Moral Norms
Forbes, M.; Hwang, J. D.; Shwartz, V.; Sap, M.; and Choi, Y. 2020 · 2020
Earlier work this paper cites.
Homeless Patients Associate Clinician Bias With Suboptimal Care for Mental Illness, Addictions, and Chronic Pain
Gilmer, C.; and Buccieri, K. 2020 · 2020
Earlier work this paper cites.
UNQOVERing Stereotyping Biases via Underspecified Questions
Li, T.; Khashabi, D.; Khot, T.; Sabharwal, A.; and Srikumar, V. 2020a · 2020
Earlier work this paper cites.
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Nangia, N.; Vania, C.; Bhalerao, R.; and Bowman, S. R. 2020 · 2020
Earlier work this paper cites.
Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets
Blodgett, S. L.; Lopez, G.; Olteanu, A.; Sim, R.; and Wallach, H. 2021 · 2021
Earlier work this paper cites.
BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
Dhamala, J.; Sun, T.; Kumar, V.; Krishna, S.; Pruksachatkun, Y.; Chang, K.; and Gupta, R. 2021 · 2021
Earlier work this paper cites.
Using Machine Learning to Reduce Toxicity Online
Perspective API. 2021 · 2021
Earlier work this paper cites.
AI and the Everything in the Whole Wide World Benchmark
Raji, I. D.; Denton, E.; Bender, E. M.; Hanna, A.; and Paullada, A. 2021 · 2021
Earlier work this paper cites.
On Measuring Social Biases in Prompt-Based Multi-Task Learning
Akyürek, A. F.; Paik, S.; Kocyigit, M. Y.; Akbiyik, S.; Runyun, S. L.; and Wijaya, D. 2022 · 2022
Cited alongside, same era.
Your Fairness May Vary: Pretrained Language Model Fairness in Toxic Text Classification
Baldini, I.; Wei, D.; Ramamurthy, K. N.; Yurochkin, M.; and Singh, M. 2022 · 2022
Cited alongside, same era.
Trustworthy Social Bias Measurement
Bommasani, R.; and Liang, P. 2022 · 2022
Cited alongside, same era.
The Dangers of Underclaiming: Reasons for Caution When Reporting How NLP Systems Fail
Bowman, S. 2022 · 2022
Cited alongside, same era.
Large Language Models are Zero-Shot Reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022 · 2022
Cited alongside, same era.
Holistic Evaluation of Language Models
Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; Newman, B.; Yuan, B.; Yan, B.; Zhang, C.; Cosgrove, C.; Manning, C. D.; Ré, C.; Acosta-Navas, D.; Hudson, D. A.; Zelikman, E.; Durmus, E.; Ladhak, F.; Rong, F.; Ren, H.; Yao, H.; Wang, J.; Santhanam, K.; Orr, L.; Zheng, L.; Yuksekgonul, M.; Suzgun, M.; Kim, N.; Guha, N.; Chatterji, N.; Khattab, O.; Henderson, P.; Huang, Q.; Chi, R.; Xie, S. M.; Santurkar, S.; Ganguli, S.; Hashimoto, T.; Icard, T.; Zhang, T.; Chaudhary, V.; Wang, W.; Li, X.; Mai, Y.; Zhang, Y.; and Koreeda, Y. 2022 · 2022
Social bias, discrimination and inequity in healthcare: mechanisms, implications and recommendations
Webster, C. S.; Taylor, S.; Thomas, C.; and Weller, J. M. 2022 · 2022
Later among the works it cites.
AI Spring? Four Takeaways from Major Releases in Foundation Models
Bommasani, R. 2023 · 2023
Closest in time.
AI is acting ‘pro-anorexia’ and tech companies aren’t stopping it
Fowler, G. A. 2023 · 2023
Closest in time.
Bias Against 93 Stigmatized Groups in Masked Language Models and Downstream Sentiment Classification Tasks
Mei, K.; Fereidooni, S.; and Caliskan, A. 2023 · 2023
Closest in time.
SocialStigmaQA: A Benchmark to Uncover Stigma Amplification in Generative Language Models
Nagireddy, M.; Chiazor, L.; Singh, M.; and Baldini, I. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ChatGPT: Optimizing Language Models for Dialogue
OpenAI. 2022 · 2022
Cited alongside, same era.
BBQ: A hand-built bias benchmark for question answering
Parrish, A.; Chen, A.; Nangia, N.; Padmakumar, V.; Phang, J.; Thompson, J.; Htut, P. M.; and Bowman, S. 2022a · 2022
Cited alongside, same era.
The Tail Wagging the Dog: Dataset Construction Biases of Social Bias Benchmarks
Selvam, N. R.; Dev, S.; Khashabi, D.; Khot, T.; and Chang, K.-W. 2022 · 2022
Cited alongside, same era.
“I’m sorry to hear that”: Finding New Biases in Language Models with a Holistic Descriptor Dataset
Smith, E. M.; Hall, M.; Kambadur, M.; Presani, E.; and Williams, A. 2022 · 2022
Cited alongside, same era.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Srivastava, A.; et al. 2022 · 2022
Cited alongside, same era.
Measure and Improve Robustness in NLP Models: A Survey
Wang, X.; Wang, H.; and Yang, D. 2022 · 2022
Cited alongside, same era.
Bias in Clinical Risk Prediction Models: Challenges in Application to Observational Health Data
Park, Y.; Singh, M.; Sylla, I.; Xiao, E.; Hu, J.; and Das, A. 2021 · 2023
Closest in time.
The Tail Wagging the Dog: Dataset Construction Biases of Social Bias Benchmarks
Selvam, N.; Dev, S.; Khashabi, D.; Khot, T.; and Chang, K.-W. 2023 · 2023
Closest in time.
On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning
Shaikh, O.; Zhang, H.; Held, W.; Bernstein, M.; and Yang, D. 2023 · 2023
Closest in time.
Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
Sun, Z.; Shen, Y.; Zhou, Q.; Zhang, H.; Chen, Z.; Cox, D.; Yang, Y.; and Gan, C. 2023 · 2023
Closest in time.
UL2: Unifying Language Learning Paradigms
Tay, Y.; Dehghani, M.; Tran, V. Q.; Garcia, X.; Wei, J.; Wang, X.; Chung, H. W.; Bahri, D.; Schuster, T.; Zheng, H. S.; Zhou, D.; Houlsby, N.; and Metzler, D. 2023 · 2023
Closest in time.
Turpin, M.; Michael, J.; Perez, E.; and Bowman, S. R. 2023 · 2023
Closest in time.
BBQ: A hand-built bias benchmark for question answering
Parrish, A.; Chen, A.; Nangia, N.; Padmakumar, V.; Phang, J.; Thompson, J.; Htut, P. M.; and Bowman, S. 2022b · 2086
Closest in time.