Fetching the paper…
Reading the bibliography…
Recently, there has been an increase in efforts to understand how large language models (LLMs) propagate and amplify social biases.
Gender bias in coreference resolution: Evaluation and debiasing methods
J. Zhao, T. Wang, M. Yatskar, V. Ordonez, and K.-W. Chang · 2003
Earlier work this paper cites.
Examining gender and race bias in two hundred sentiment analysis systems
S. Kiritchenko and S. Mohammad · 2005
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
S. Bowman, G. Angeli, C. Potts, and C. D. Manning · 2015
Earlier work this paper cites.
Ex machina: Personal attacks seen at scale
E. Wulczyn, N. Thain, and L. Dixon · 2017
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
L. Dixon, J. Li, J. Sorensen, N. Thain, and L. Vasserman · 2018
Earlier work this paper cites.
Semeval-2018 task 1: Affect in tweets
S. Mohammad, F. Bravo-Marquez, M. Salameh, and S. Kiritchenko · 2018
Earlier work this paper cites.
Reducing gender bias in abusive language detection
J. H. Park, J. Shin, and P. Fung · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
R. Rudinger, J. Naradowsky, B. Leonard, and B. Van Durme · 2018
Earlier work this paper cites.
Randomness in neural network training: Characterizing the impact of tooling
D. Zhuang, X. Zhang, S. Song, and S. Hooker · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Measuring bias in contextualized word representations
K. Kurita, N. Vyas, A. Pareek, A. W. Black, and Y. Tsvetkov · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
C. May, A. Wang, S. Bordia, S. R. Bowman, and R. Rudinger · 2019
Earlier work this paper cites.
Perturbation sensitivity analysis to detect unintended model biases
V. Prabhakaran, B. Hutchinson, and M. Mitchell · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Cited alongside, same era.
The woman worked as a babysitter: On biases in language generation
E. Sheng, K.-W. Chang, P. Natarajan, and N. Peng · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Bad seeds: Evaluating lexical methods for bias measurement
M. Antoniak and D. Mimno · 2021
Later among the works it cites.
Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
S. L. Blodgett, G. Lopez, A. Olteanu, R. Sim, and H. Wallach · 2021
Later among the works it cites.
Robustness gym: Unifying the NLP evaluation landscape
K. Goel, N. F. Rajani, J. Vig, Z. Taschdjian, M. Bansal, and C. Ré · 2021
Later among the works it cites.
Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models
H. R. Kirk, Y. Jun, F. Volpin, H. Iqbal, E. Benussi, F. Dreyer, A. Shtedritski, and Y. Asano · 2021
Later among the works it cites.
Automatic construction of evaluation suites for natural language generation datasets
S. Mille, K. Dhole, S. Mahamood, L. Perez-Beltrachini, V. P. Gangal, M. Kale, E. van Miltenburg, and S. Gehrmann · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Dev, T. Li, J. M. Phillips, and V. Srikumar · 2020
Cited alongside, same era.
Reducing sentiment bias in language models via counterfactual evaluation
P.-S. Huang, H. Zhang, R. Jiang, R. Stanforth, J. Welbl, J. Rae, V. Maini, D. Yogatama, and P. Kohli · 2020
Cited alongside, same era.
UNQOVERing stereotyping biases via underspecified questions
T. Li, D. Khashabi, T. Khot, A. Sabharwal, and V. Srikumar · 2020
Cited alongside, same era.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
N. Nangia, C. Vania, R. Bhalerao, and S. R. Bowman · 2020
Cited alongside, same era.
Beyond accuracy: Behavioral testing of NLP models with CheckList
M. T. Ribeiro, T. Wu, C. Guestrin, and S. Singh · 2020
Cited alongside, same era.
Persistent anti-muslim bias in large language models
A. Abid, M. Farooqi, and J. Zou · 2021
Cited alongside, same era.
On the impact of random seeds on the fairness of clinical classifiers
S. Amir, J.-W. van de Meent, and B. Wallace · 2021
Cited alongside, same era.
StereoSet: Measuring stereotypical bias in pretrained language models
M. Nadeem, A. Bethke, and S. Reddy · 2021
Later among the works it cites.
Are my deep learning systems fair? an empirical study of fixed-seed training
S. Qian, V. H. Pham, T. Lutellier, Z. Hu, J. Kim, L. Tan, Y. Yu, J. Chen, and S. Shah · 2021
Later among the works it cites.
HateCheck: Functional tests for hate speech detection models
P. Röttger, B. Vidgen, D. Nguyen, Z. Waseem, H. Margetts, and J. Pierrehumbert · 2021
Later among the works it cites.
Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models
P. Delobelle, E. Tokpo, T. Calders, and B. Berendt · 2022
Closest in time.
Choose your lenses: Flaws in gender bias evaluation
H. Orgad and Y. Belinkov · 2022
Closest in time.
Adaptive testing and debugging of NLP models
M. T. Ribeiro and S. Lundberg · 2022
Closest in time.
The MultiBERTs: BERT Reproductions for Robustness Analysis , 2022
T. Sellam, S. Yadlowsky, I. Tenney, J. Wei, N. Saphra, A. N. D’Amour, T. Linzen, J. Bastings, I. R. Turc, J. Eisenstein, D. Das, and E. Pavlick, editors · 2022
Closest in time.