Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have shown powerful performance and development prospects and are widely deployed in the real world.
1906
Earlier work this paper cites.
1907
Earlier work this paper cites.
D. Khashabi, S. Min, T. Khot, A. Sabharwal, O. Tafjord, P. Clark, H. Hajishirzi, UNIFIEDQA: Crossing format boundaries with a single QA system, in: Proceedings of the Findings of the 2022 Association for Computational Linguistics, EMNLP, 2020, pp. 1896–1907
1907
Earlier work this paper cites.
1909
Earlier work this paper cites.
V. Sanh, L. Debut, J. Chaumond, T. Wolf, Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter , CoRR abs/1910.01108 · 1910
Earlier work this paper cites.
S. Goldfarb-Tarrant, R. Marchant, R. M. Sánchez, M. Pandya, A. Lopez, Intrinsic bias metrics do not correlate with application bias, in: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP, 2021, pp. 1926–1940
1940
Earlier work this paper cites.
N. Nangia, C. Vania, R. Bhalerao, S. R. Bowman, Crows-pairs: A challenge dataset for measuring social biases in masked language models, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP, 2020, pp. 1953–1967
1967
Earlier work this paper cites.
R. G. Miller, The jackknife-a review, Biometrika 61 (1) (1974) 1–15
1974
Earlier work this paper cites.
S. T. Fiske, A. J. Cuddy, P. Glick, J. Xu, A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition, in: Journal of Personality and Social Psychology, 2002, pp. 878–902
2002
Earlier work this paper cites.
doi:10.18653/V1/S18-2005
S. Kiritchenko, S. M. Mohammad, Examining gender and race bias in two hundred sentiment analysis systems , in: Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics, *SEM@NAACL-HLT 2018, New Orleans, Louisiana, USA, June 5-6, 2018, Association for Computational Linguistics, 2018, pp. 43–53 · 2005
Earlier work this paper cites.
2010
Earlier work this paper cites.
T. Li, T. Khot, D. Khashabi, A. Sabharwal, V. Srikumar, Unqovering stereotyping biases via underspecified questions , CoRR abs/2010.02428 · 2010
Earlier work this paper cites.
H. J. Levesque, E. Davis, L. Morgenstern, The winograd schema challenge, in: Proceedings of the 30th Principles of Knowledge Representation and Reasoning, KR, 2012
2012
Earlier work this paper cites.
A. Caliskan, J. J. Bryson, A. Narayanan, Semantics derived automatically from language corpora contain human-like biases, Science 356 (6334) (2017) 183–186
2017
Earlier work this paper cites.
D. M. Cer, M. T. Diab, E. Agirre, I. Lopez-Gazpio, L. Specia, Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation, in: Proceedings of the 11th International Workshop on Semantic Evaluation, SemEval@ACL, 2017, pp. 1–14
2017
Earlier work this paper cites.
doi:10.18653/V1/D17-1082
G. Lai, Q. Xie, H. Liu, Y. Yang, E. H. Hovy, RACE: large-scale reading comprehension dataset from examinations , in: Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017, Association for Computational Linguistics, 2017, pp. 785–794 · 2017
Earlier work this paper cites.
doi:10.18653/V1/D17-1247
M. Sap, M. C. Prasettio, A. Holtzman, H. Rashkin, Y. Choi, Connotation frames of power and agency in modern films , in: Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017, Association for Computational Linguistics, 2017, pp. 2329–2334 · 2017
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, D. Amodei, Deep reinforcement learning from human preferences, in: Proceedings of the 30th Annual Conference on Neural Information Processing Systems, NeurIPS, 2017, pp. 4299–4307
2017
Earlier work this paper cites.
N. Garg, L. Schiebinger, D. Jurafsky, J. Zou, Word embeddings quantify 100 years of gender and ethnic stereotypes, Proc. Natl. Acad. Sci. USA 115 (16) (2018) E3635–E3644
2018
Earlier work this paper cites.
J. Zhao, T. Wang, M. Yatskar, V. Ordonez, K. Chang, Gender bias in coreference resolution: Evaluation and debiasing methods, in: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics, NAACL-HLT, 2018, pp. 15–20
2018
Earlier work this paper cites.
R. Rudinger, J. Naradowsky, B. Leonard, B. V. Durme, Gender bias in coreference resolution, in: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics, NAACL, 2018, pp. 8–14
2018
Earlier work this paper cites.
K. Webster, M. Recasens, V. Axelrod, J. Baldridge, Mind the GAP: A balanced corpus of gendered ambiguous pronouns, Trans. Assoc. Comput. Linguistics 6 (2018) 605–617
2018
Earlier work this paper cites.
doi:10.18653/V1/S18-1001
S. M. Mohammad, F. Bravo-Marquez, M. Salameh, S. Kiritchenko, Semeval-2018 task 1: Affect in tweets , in: Proceedings of The 12th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2018, New Orleans, Louisiana, USA, June 5-6, 2018, Association for Computational Linguistics, 2018, pp. 1–17 · 2018
Earlier work this paper cites.
L. Dixon, J. Li, J. Sorensen, N. Thain, L. Vasserman, Measuring and mitigating unintended bias in text classification, in: Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, AIES, 2018, pp. 67–73
2018
Earlier work this paper cites.
J. H. Park, J. Shin, P. Fung, Reducing gender bias in abusive language detection, in: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP, 2018, pp. 2799–2804
2018
Earlier work this paper cites.
Y. Elazar, Y. Goldberg, Adversarial removal of demographic attributes from text data, in: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP, 2018, pp. 11–21
2018
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, K. Toutanova, BERT: pre-training of deep bidirectional transformers for language understanding, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, NAACL, 2019, pp. 4171–4186
2019
Earlier work this paper cites.
T. Sun, A. Gaut, S. Tang, Y. Huang, M. ElSherief, J. Zhao, D. Mirza, E. M. Belding, K. Chang, W. Y. Wang, Mitigating gender bias in natural language processing: Literature review, in: Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL, 2019, pp. 1630–1640
2019
Earlier work this paper cites.
M. De-Arteaga, A. Romanov, H. M. Wallach, J. T. Chayes, C. Borgs, A. Chouldechova, S. C. Geyik, K. Kenthapadi, A. T. Kalai, Bias in bios: A case study of semantic representation bias in a high-stakes setting, in: Proceedings of the Conference on Fairness, Accountability, and Transparency, FAT, 2019, pp. 120–128
2019
Earlier work this paper cites.
Z. Obermeyer, B. Powers, C. Vogeli, S. Mullainathan, Dissecting racial bias in an algorithm used to manage the health of populations, Science 366 (6464) (2019) 447–453
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al., Language models are unsupervised multitask learners, OpenAI blog 1 (8) (2019) 9
2019
Earlier work this paper cites.
C. May, A. Wang, S. Bordia, S. R. Bowman, R. Rudinger, On measuring social biases in sentence encoders, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, NAACL, 2019, pp. 622–628
2019
Earlier work this paper cites.
Y. C. Tan, L. E. Celis, Assessing social and intersectional biases in contextualized word representations, in: Proceedings of the 32nd Annual Conference on Neural Information Processing Systems, NeurIPS, 2019, pp. 13209–13220
2019
Earlier work this paper cites.
doi:10.18653/V1/D19-1339
E. Sheng, K. Chang, P. Natarajan, N. Peng, The woman worked as a babysitter: On biases in language generation , in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, Association for Computational Linguistics, 2019, pp. 3405–3410 · 2019
Earlier work this paper cites.
R. Jiang, A. Pacchiano, T. Stepleton, H. Jiang, S. Chiappa, Wasserstein fair classification , in: Proceedings of the Thirty-Fifth Conference on Uncertainty in Artificial Intelligence, UAI 2019, Tel Aviv, Israel, July 22-25, 2019, Vol. 115 of Proceedings of Machine Learning Research, AUAI Press, 2019, pp. 862–872. URL http://proceedings.mlr.press/v115/jiang20a.html
2019
Earlier work this paper cites.
R. Zmigrod, S. J. Mielke, H. M. Wallach, R. Cotterell, Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology, in: Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL, 2019, pp. 1651–1661
2019
Earlier work this paper cites.
M. Brunet, C. Alkalay-Houlihan, A. Anderson, R. S. Zemel, Understanding the origins of bias in word embeddings, in: Proceedings of the 36th International Conference on Machine Learning, ICML, 2019, pp. 803–811
2019
Earlier work this paper cites.
S. Dev, J. M. Phillips, Attenuating bias in word vectors, in: Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics, AISTATS, Vol. 89, 2019, pp. 879–887
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, S. Gelly, Parameter-efficient transfer learning for NLP , in: Proceedings of the 36th International Conference on Machine Learning, ICML, Vol. 97 of Proceedings of Machine Learning Research, PMLR, 2019, pp. 2790–2799. URL http://proceedings.mlr.press/v97/houlsby19a.html
2019
Earlier work this paper cites.
S. Garg, V. Perot, N. Limtiaco, A. Taly, E. H. Chi, A. Beutel, Counterfactual fairness in text classification through robustness, in: Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, AIES, 2019, pp. 219–226
2019
Earlier work this paper cites.
M. Sap, D. Card, S. Gabriel, Y. Choi, N. A. Smith, The risk of racial bias in hate speech detection, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, ACL, 2019, pp. 1668–1678
2019
Earlier work this paper cites.
R. Liu, J. Regier, N. Tripuraneni, M. I. Jordan, J. D. McAuliffe, Rao-blackwellized stochastic gradients for discrete distributions , in: Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, Vol. 97 of Proceedings of Machine Learning Research, PMLR, 2019, pp. 4023–4031. URL http://proceedings.mlr.press/v97/liu19c.html
2019
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, D. Amodei, Language models are few-shot learners, in: Proceedings of the 33rd Annual Conference on Neural Information Processing Systems, NeurIPS, 2020
2020
Earlier work this paper cites.
S. L. Blodgett, S. Barocas, H. D. III, H. M. Wallach, Language (technology) is power: A critical survey of "bias" in NLP, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL, 2020, pp. 5454–5476
2020
Earlier work this paper cites.
K. V. Deshpande, S. Pan, J. R. Foulds, Mitigating demographic bias in ai-based resume filtering, in: Proc. 28th UMAP Adjun. - Adjun. Publ. ACM Conf. User Model., Adapt. Pers., ACM, 2020, pp. 268–275
2020
Earlier work this paper cites.
D. Shah, H. A. Schwartz, D. Hovy, Predictive biases in natural language processing models: A conceptual framework and overview, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL, 2020, pp. 5248–5264
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, P. J. Liu, Exploring the limits of transfer learning with a unified text-to-text transformer , J. Mach. Learn. Res. 21 (2020) 140:1–140:67. URL http://jmlr.org/papers/v21/20-074.html
2020
Earlier work this paper cites.
S. Dev, T. Li, J. M. Phillips, V. Srikumar, On measuring and mitigating biased inferences of word embeddings, in: Proceedings of the 34th Association for the Advancement of Artificial Intelligence, AAAI, 2020, pp. 7659–7666
2020
Earlier work this paper cites.
doi:10.18653/V1/2020.FINDINGS-EMNLP.7
P. Huang, H. Zhang, R. Jiang, R. Stanforth, J. Welbl, J. Rae, V. Maini, D. Yogatama, P. Kohli, Reducing sentiment bias in language models via counterfactual evaluation , in: Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, Vol. EMNLP 2020 of Findings of ACL, Association for Computational Linguistics, 2020, pp. 65–83 · 2020
Earlier work this paper cites.
M. Du, F. Yang, N. Zou, X. Hu, Fairness in deep learning: A computational perspective, IEEE Intelligent Systems 36 (4) (2020) 25–34
2020
Cited alongside, same era.
K. Lu, P. Mardziel, F. Wu, P. Amancharla, A. Datta, Gender bias in neural natural language processing, in: Proceedings of the Logic, Language, and Security - Essays Dedicated to Andre Scedrov on the Occasion of His 65th Birthday, Vol. 12300 of Lecture Notes in Computer Science, 2020, pp. 189–202
2020
Cited alongside, same era.
doi:10.18653/V1/2020.EMNLP-MAIN.602
X. Ma, M. Sap, H. Rashkin, Y. Choi, Powertransformer: Unsupervised controllable revision for biased language correction , in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, Association for Computational Linguistics, 2020, pp. 7426–7441 · 2020
Cited alongside, same era.
P. P. Liang, I. M. Li, E. Zheng, Y. C. Lim, R. Salakhutdinov, L. Morency, Towards debiasing sentence representations, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL, 2020, pp. 5502–5515
2020
Cited alongside, same era.
P. Delobelle, B. Berendt, Fairdistillation: mitigating stereotyping in language models, in: Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Springer, 2022, pp. 638–654
2022
Later among the works it cites.
U. Gupta, J. Dhamala, V. Kumar, A. Verma, Y. Pruksachatkun, S. Krishna, R. Gupta, K. Chang, G. V. Steeg, A. Galstyan, Mitigating gender bias in distilled language models via counterfactual role reversal, in: Proceedings of the Findings of the Association for Computational Linguistics, ACL, 2022, pp. 658–678
2022
Later among the works it cites.
Y. Guo, Y. Yang, A. Abbasi, Auto-debias: Debiasing masked language models with automated biased prompts, in: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, ACL, 2022, pp. 1012–1023
2022
Later among the works it cites.
doi:10.18653/V1/2022.FINDINGS-ACL.88
G. Attanasio, D. Nozza, D. Hovy, E. Baralis, Entropy-based attention regularization frees unintended bias mitigation from lists , in: Findings of the Association for Computational Linguistics: ACL, Association for Computational Linguistics, 2022, pp. 1105–1119 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Saunders, B. Byrne, Reducing gender bias in neural machine translation as a domain adaptation problem, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL, 2020, pp. 7724–7736
2020
Cited alongside, same era.
G. Zhang, B. Bai, J. Zhang, K. Bai, C. Zhu, T. Zhao, Demographics should not be the reason of toxicity: Mitigating discrimination in text classifications with instance weighting, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL, 2020, pp. 4134–4145
2020
Cited alongside, same era.
doi:10.18653/V1/2020.EMNLP-MAIN.613
P. A. Utama, N. S. Moosavi, I. Gurevych, Towards debiasing NLU models from unknown biases , in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, Association for Computational Linguistics, 2020, pp. 7597–7610 · 2020
Cited alongside, same era.
P. Lahoti, A. Beutel, J. Chen, K. Lee, F. Prost, N. Thain, X. Wang, E. H. Chi, Fairness without demographics through adversarially reweighted learning, in: Proceedings of the 33rd Annual Conference on Neural Information Processing Systems, NeurIPS, 2020
2020
Cited alongside, same era.
S. Ravfogel, Y. Elazar, H. Gonen, M. Twiton, Y. Goldberg, Null it out: Guarding protected attributes by iterative nullspace projection, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL, 2020, pp. 7237–7256
2020
Cited alongside, same era.
doi:10.18653/V1/2020.EMNLP-MAIN.48
M. Forbes, J. D. Hwang, V. Shwartz, M. Sap, Y. Choi, Social chemistry 101: Learning to reason about social and moral norms , in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, Association for Computational Linguistics, 2020, pp. 653–670 · 2020
Cited alongside, same era.
I. Garrido-Muñoz, A. Montejo-Ráez, F. Martínez-Santiago, L. A. Ureña-López, A survey on bias in deep nlp, Applied Sciences 11 (7) (2021) 3184
2021
Cited alongside, same era.
P. He, X. Liu, J. Gao, W. Chen, Deberta: decoding-enhanced bert with disentangled attention , in: 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021, OpenReview.net, 2021. URL https://openreview.net/forum?id=XPZIaotutsD
2021
Cited alongside, same era.
Later among the works it cites.
J. He, M. Xia, C. Fellbaum, D. Chen, MABEL: attenuating gender bias using textual entailment data, in: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP, 2022, pp. 9681–9702
2022
Later among the works it cites.
doi:10.1145/3534678.3539232
C. Oh, H. Won, J. So, T. Kim, Y. Kim, H. Choi, K. Song, Learning fair representation via distributional contrastive disentanglement , in: KDD ’22: The 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Washington, DC, USA, August 14 - 18, 2022, ACM, 2022, pp. 1295–1305 · 2022
Later among the works it cites.
S. Panda, A. Kobren, M. Wick, Q. Shen, Don’t just clean it, proxy clean it: Mitigating bias by proxy in pre-trained models, in: Proceedings of the Findings of the Association for Computational Linguistics, EMNLP, 2022, pp. 5073–5085
2022
Later among the works it cites.
P. Sattigeri, S. Ghosh, I. Padhi, P. L. Dognin, K. R. Varshney, Fair infinitesimal jackknife: Mitigating the influence of biased training data points without refitting , in: NeurIPS, 2022. URL http://papers.nips.cc/paper_files/paper/2022/hash/e94481b99473c83b2e79d91c64eb37d1-Abstract-Conference.html
2022
Later among the works it cites.
A. Garimella, R. Mihalcea, A. Amarnath, Demographic-aware language model fine-tuning as a bias mitigation technique , in: Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing, AACL/IJCNLP 2022 - Volume 2: Short Papers, Online only, November 20-23, 2022, Association for Computational Linguistics, 2022, pp. 311–319. URL https://aclanthology.org/2022.aacl-short.38
2022
Later among the works it cites.
X. Han, T. Baldwin, T. Cohn, Balancing out bias: Achieving fairness through balanced training, in: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP, 2022, pp. 11335–11350
2022
Later among the works it cites.
S. Ravfogel, M. Twiton, Y. Goldberg, R. D. Cotterell, Linear adversarial concept erasure, in: Proceedings of the 39th International Conference on Machine Learning, ICML, Vol. 162, 2022, pp. 18400–18421
2022
Later among the works it cites.
A. Shen, X. Han, T. Cohn, T. Baldwin, L. Frermann, Does representational fairness imply empirical fairness?, in: Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022, 2022, pp. 81–95
2022
Later among the works it cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, R. Lowe, Training language models to follow instructions with human feedback, in: NeurIPS, 2022
2022
Later among the works it cites.
V. Sanh, A. Webson, C. Raffel, S. H. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, A. Raja, M. Dey, M. S. Bari, C. Xu, U. Thakker, S. S. Sharma, E. Szczechla, T. Kim, G. Chhablani, N. V. Nayak, D. Datta, J. Chang, M. T. Jiang, H. Wang, M. Manica, S. Shen, Z. X. Yong, H. Pandey, R. Bawden, T. Wang, T. Neeraj, J. Rozen, A. Sharma, A. Santilli, T. Févry, J. A. Fries, R. Teehan, T. L. Scao, S. Biderman, L. Gao, T. Wolf, A. M. Rush, Multitask prompted training enables zero-shot task generalization, in: Proceedings of the 10th International Conference on Learning Representations, ICLR, 2022
2022
Later among the works it cites.
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, Q. V. Le, Finetuned language models are zero-shot learners, in: Proceedings of the 10th International Conference on Learning Representations, ICLR, 2022
2022
Later among the works it cites.
P. Delobelle, E. K. Tokpo, T. Calders, B. Berendt, Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models, in: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics, NAACL, 2022, pp. 1693–1706
2022
Later among the works it cites.
Y. T. Cao, Y. Pruksachatkun, K. Chang, R. Gupta, V. Kumar, J. Dhamala, A. Galstyan, On the intrinsic and extrinsic fairness evaluation metrics for contextualized language representations, in: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, ACL, 2022, pp. 561–570
2022
Later among the works it cites.
I. Baldini, D. Wei, K. N. Ramamurthy, M. Singh, M. Yurochkin, Your fairness may vary: Pretrained language model fairness in toxic text classification, in: Proceedings of the Findings of the Association for Computational Linguistics, ACL, 2022, pp. 2245–2262
2022
Later among the works it cites.
X. Lu, S. Welleck, J. Hessel, L. Jiang, L. Qin, P. West, P. Ammanabrolu, Y. Choi, QUARK: controllable text generation with reinforced unlearning , in: NeurIPS, 2022. URL http://papers.nips.cc/paper_files/paper/2022/hash/b125999bde7e80910cbdbd323087df8f-Abstract-Conference.html
2022
Later among the works it cites.
S. Kumar, V. Balachandran, L. Njoo, A. Anastasopoulos, Y. Tsvetkov, Language generation models can cause harm: So what can we do about it? an actionable survey, in: Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, EACL, 2023, pp. 3291–3313
2023
Closest in time.
doi:10.24963/IJCAI.2023/742
U. Gohar, L. Cheng, A survey on intersectional fairness in machine learning: Notions, mitigation, and challenges , in: Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, IJCAI 2023, 19th-25th August 2023, Macao, SAR, China, ijcai.org, 2023, pp. 6619–6627 · 2023
Closest in time.
doi:10.1016/J.INFFUS.2023.101906
D. Jin, L. Wang, H. Zhang, Y. Zheng, W. Ding, F. Xia, S. Pan, A survey on fairness-aware recommender systems , Inf. Fusion 100 (2023) 101906 · 2023
Closest in time.
doi:10.1145/3547333
Y. Wang, W. Ma, M. Zhang, Y. Liu, S. Ma, A survey on the fairness of recommender systems , ACM Transactions on Information Systems 41 (3) (2023) 52:1–52:43 · 2023
Closest in time.
P. He, J. Gao, W. Chen, Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing , in: The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, OpenReview.net, 2023. URL https://openreview.net/pdf?id=sE7-XhLxHA
2023
Closest in time.
Z. Xie, T. Lukasiewicz, An empirical analysis of parameter-efficient methods for debiasing pre-trained language models, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL, 2023, pp. 15730–15745
2023
Closest in time.
doi:10.18653/V1/2023.ACL-SHORT.30
H. Thakur, A. Jain, P. Vaddamanu, P. P. Liang, L. Morency, Language models get a gender makeover: Mitigating gender bias with few-shot data interventions , in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL, Association for Computational Linguistics, 2023, pp. 340–351 · 2023
Closest in time.
C. Amrhein, F. Schottmann, R. Sennrich, S. Läubli, Exploiting biased models to de-bias text: A gender-fair rewriting model, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL, 2023, pp. 4486–4506
2023
Closest in time.
A. Omrani, A. S. Ziabari, C. Yu, P. Golazizian, B. Kennedy, M. Atari, H. Ji, M. Dehghani, Social-group-agnostic bias mitigation via the stereotype content model, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL, 2023, pp. 4123–4139
2023
Closest in time.
Y. Li, M. Du, X. Wang, Y. Wang, Prompt tuning pushes farther, contrastive learning pulls closer: A two-stage approach to mitigate social biases, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL, 2023, pp. 14254–14267
2023
Closest in time.
doi:10.18653/V1/2023.FINDINGS-ACL.369
S. Iskander, K. Radinsky, Y. Belinkov, Shielded representations: Protecting sensitive attributes through iterative gradient-based projection , in: Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023, Association for Computational Linguistics, 2023, pp. 5961–5977 · 2023
Closest in time.
Z. Fatemi, C. Xing, W. Liu, C. Xiong, Improving gender fairness of pre-trained language models without catastrophic forgetting, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL, 2023, pp. 1249–1262
2023
Closest in time.
doi:10.1609/AAAI.V37I9.26279
K. Yang, C. Yu, Y. R. Fung, M. Li, H. Ji, ADEPT: A debiasing prompt framework , in: Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2023, Washington, DC, USA, February 7-14, 2023, AAAI Press, 2023, pp. 10780–10788 · 2023
Closest in time.
doi:10.18653/V1/2023.FINDINGS-ACL.386
L. Hauzenberger, S. Masoudian, D. Kumar, M. Schedl, N. Rekabsaz, Modular and on-demand bias mitigation with attribute-removal subnetworks , in: Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023, Association for Computational Linguistics, 2023, pp. 6192–6214 · 2023
Closest in time.
doi:10.1145/3539618.3591938
L. Yu, Y. Mao, J. Wu, F. Zhou, Mixup-based unified framework to overcome gender bias resurgence , in: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2023, ACM, 2023, pp. 1755–1759 · 2023
Closest in time.
doi:10.1609/AAAI.V37I12.26706
A. Zayed, P. Parthasarathi, G. Mordido, H. Palangi, S. Shabanian, S. Chandar, Deep learning on a healthy data diet: Finding important examples for fairness , in: Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2023, Washington, DC, USA, February 7-14, 2023, AAAI Press, 2023, pp. 14593–14601 · 2023
Closest in time.
H. Orgad, Y. Belinkov, BLIND: Bias removal with no demographics, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL, 2023, pp. 8801–8821
2023
Closest in time.
doi:10.1145/3539597.3570473
S. Park, K. Choi, H. Yu, Y. Ko, Never too late to learn: Regularizing gender bias in coreference resolution , in: Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining, WSDM, ACM, 2023, pp. 15–23 · 2023
Closest in time.
S. Ghanbarzadeh, Y. Huang, H. Palangi, R. C. Moreno, H. Khanpour, Gender-tuning: Empowering fine-tuning for debiasing pre-trained language models, in: Proceedings of the findings of the Association for Computational Linguistics: ACL, 2023, pp. 5448–5458
2023
Closest in time.
R. Wang, P. Cheng, R. Henao, Toward fairness in text generation via mutual information minimization based on importance sampling , in: International Conference on Artificial Intelligence and Statistics, AISTATS, Vol. 206 of Proceedings of Machine Learning Research, PMLR, 2023, pp. 4473–4485. URL https://proceedings.mlr.press/v206/wang23c.html
2023
Closest in time.
M. Cheng, E. Durmus, D. Jurafsky, Marked personas: Using natural language prompts to measure stereotypes in language models, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL, 2023, pp. 1504–1532
2023
Closest in time.
A. Ramezani, Y. Xu, Knowledge of cultural moral norms in large language models, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL, 2023, pp. 428–446
2023
Closest in time.
T. Y. Zhuo, Y. Huang, C. Chen, Z. Xing, Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity, 2023
2023
Closest in time.
E. Fleisig, A. Amstutz, C. Atalla, S. L. Blodgett, H. Daumé III, A. Olteanu, E. Sheng, D. Vann, H. Wallach, Fair-prism: Evaluating fairness-related harms in text generation, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 2023
2023
Closest in time.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, E. P. Xing, Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality (March 2023). URL https://lmsys.org/blog/2023-03-30-vicuna/
2023
Closest in time.
Anthropic, Model card and evaluations for claude models (July 2023). URL https://www-files.anthropic.com/production/images/Model-Card-Claude-2.pdf?dm=1689034733
2023
Closest in time.
S. Santy, J. T. Liang, R. L. Bras, K. Reinecke, M. Sap, Nlpositionality: Characterizing design biases of datasets and models, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL, 2023, pp. 9080–9102
2023
Closest in time.
J. Watson, B. Beekhuizen, S. Stevenson, What social attitudes about gender does BERT encode? leveraging insights from psycholinguistics, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL, 2023, pp. 6790–6809
2023
Closest in time.
A. Parrish, A. Chen, N. Nangia, V. Padmakumar, J. Phang, J. Thompson, P. M. Htut, S. R. Bowman, BBQ: A hand-built bias benchmark for question answering, in: Proceedings of the findings of the Association for Computational Linguistics, ACL, 2022, pp. 2086–2105
2086
Closest in time.