Fetching the paper…
Reading the bibliography…
The benefits and capabilities of pre-trained language models (LLMs) in current and future innovations are vital to any society.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, et al · 1901
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, et al · 1907
Earlier work this paper cites.
RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models. In ACL-IJCNLP . ACL, Online, 1941–1955
Soumya Barikeri, Anne Lauscher, Ivan Vulić, and Goran Glavaš. 2021 · 1955
Earlier work this paper cites.
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models. In EMNLP . ACL, 1953–1967
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel Bowman. 2020 · 1967
Earlier work this paper cites.
Lexical expansion in Maori
Ray Harlow. 1993 · 1993
Earlier work this paper cites.
Identifying and avoiding bias in research
Christopher J Pannucci and Edwin G Wilkins. 2010 · 2010
Earlier work this paper cites.
Fairness through awareness. In the 3rd innovations in theoretical computer science conference . 214–226
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. 2012 · 2012
Earlier work this paper cites.
Scientific research must take gender into account
Londa Schiebinger. 2014 · 2014
Earlier work this paper cites.
Big data’s disparate impact
Solon Barocas and Andrew D Selbst. 2016 · 2016
Earlier work this paper cites.
" Why should i trust you?" Explaining the predictions of any classifier. In ACM SIGKDD . ACM, 1135–1144
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
A difference of perspective? Māori members of parliament and te ao Māori in parliament
Te Hau White. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
The trouble with bias
Kate Crawford. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
English language as thief
Vaughan Rapatahana. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Examining Gender and Race Bias in Two Hundred Sentiment Analysis Systems. In JCLCS . 43–53
Svetlana Kiritchenko and Saif Mohammad. 2018 · 2018
Earlier work this paper cites.
IEEE P7003™ standard for algorithmic bias considerations: work in progress paper. In Int. workshop on software fairness (Gothenburg, Sweden). 38–41
Ansgar Koene, Liz Dowthwaite, and Suchana Seth. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training. Preprint, OpenAI, 1–12
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Gender Bias in Coreference Resolution. In NAACL-HLT . ACL, 8–14
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Mind the GAP: A balanced corpus of gendered ambiguous pronouns
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge. 2018 · 2018
Earlier work this paper cites.
Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods. In NAACL-HLT . ACL, 15–20
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
Australia’s Artificial Intelligence Ethics Framework
Science Australian Government (Department of Industry and Resources). 2019 · 2019
Earlier work this paper cites.
Fairness and Machine Learning: Limitations and Opportunities
Solon Barocas, Moritz Hardt, and Arvind Narayanan. 2019 · 2019
Earlier work this paper cites.
Identifying and Reducing Gender Bias in Word-Level Language Models. In NAACL: SRW . Association for Computational Linguistics, 7–15
Shikha Bordia and Samuel Bowman. 2019 · 2019
Earlier work this paper cites.
Why cultural safety rather than cultural competency is required to achieve health equity: a literature review & recommended definition
Elana Curtis, Rhys Jones, David Tipene-Leach, et al · 2019
Earlier work this paper cites.
Plug and Play Language Models: A Simple Approach to Controlled Text Generation. In ICLR
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HTT . Association for Computational Linguistics, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Recommendation of the Council on Artificial Intelligence, OECD/LEGAL/0449 (Paris)
Organisation for Economic Co-operation and Development (OECD). 2019 · 2019
Earlier work this paper cites.
Counterfactual fairness in text classification through robustness. In AAAI/ACM Conference on AI, Ethics, and Society . 219–226
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H Chi, and Alex Beutel. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP. In ICML . PMLR, 2790–2799
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Measuring Bias in Contextualized Word Representations. In GeBNLP . 166–172
Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019 · 2019
Earlier work this paper cites.
It’s All in the Name: Mitigating Gender Bias with Name-Based Counterfactual Data Substitution. In EMNLP-IJCNLP . ACL, 5267–5275
Rowan Hall Maudslay, Hila Gonen, Ryan Cotterell, and Simone Teufel. 2019 · 2019
Earlier work this paper cites.
On Measuring Social Biases in Sentence Encoders. In NAACL-HLT . ACL, 622–628
Chandler May, Alex Wang, Shikha Bordia, Samuel R Bowman, and Rachel Rudinger. 2019 · 2019
Earlier work this paper cites.
Reducing Gender Bias in Word-Level Language Models with a Gender-Equalizing Loss Function
Yusu Qian, Urwa Muaz, Ben Zhang, and Jae Won Hyun. 2019 · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners. OpenAI blog 1, no. 8, 9
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. 2019 · 2019
Earlier work this paper cites.
The Woman Worked as a Babysitter: On Biases in Language Generation. In EMNLP-IJCNLP . Association for Computational Linguistics, 3407–3412
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2019 · 2019
Earlier work this paper cites.
Māori loanwords: a corpus of New Zealand English tweets. In ACL-SRW . Association for Computational Linguistics, 136–142
David Trye, Andreea S Calude, Felipe Bravo-Marquez, and Te Taka Adrian Gregory Keegan. 2019 · 2019
Earlier work this paper cites.
Counterfactual Data Augmentation for Mitigating Gender Stereotypes in Languages with Rich Morphology. In ACL . 1651–1661
Ran Zmigrod, Sabrina J. Mielke, Hanna Wallach, and Ryan Cotterell. 2019 · 2019
Earlier work this paper cites.
Unmasking Contextual Stereotypes: Measuring and Mitigating BERT’s Gender Bias. In GeBNLP
Marion Bartl, Malvina Nissim, and Albert Gatt. 2020 · 2020
Earlier work this paper cites.
Decolonising speech and language technology. In ICCL . 3504–3519
Steven Bird. 2020 · 2020
Earlier work this paper cites.
Language (Technology) is Power: A Critical Survey of “Bias” in NLP. In ACL . 5454–5476
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Earlier work this paper cites.
White Paper: On Artificial Intelligence - A European approach to excellence and trust, COM(2020) 65 final
European Commission. 2020 · 2020
Earlier work this paper cites.
On measuring and mitigating biased inferences of word embeddings. In AAAI , Vol. 34. 7659–7666
Sunipa Dev, Tao Li, Jeff M Phillips, and Vivek Srikumar. 2020 · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. In Findings of EMNLP . ACL, Online, 3356–3369
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Mitigating Gender Bias Amplification in Distribution by Posterior Regularization. In ACL . 2936–2942
Shengyu Jia, Tao Meng, Jieyu Zhao, and Kai-Wei Chang. 2020 · 2020
Earlier work this paper cites.
End-to-End Bias Mitigation by Modelling Biases in Corpora. In ACL . 8706–8716
Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson. 2020 · 2020
Earlier work this paper cites.
Racial disparities in automated speech recognition
Allison Koenecke, Andrew Nam, Emily Lake, Joe Nudell, Minnie Quartey, Zion Mengesha, Connor Toups, et al · 2020
Earlier work this paper cites.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In ACL . 7871–7880
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, et al · 2020
Earlier work this paper cites.
UnQovering Stereotyping Biases via Underspecified Questions. In Findings of EMNLP . Association for Computational Linguistics
Tao Li, Tushar Khot, Daniel Khashabi, Ashish Sabharwal, and Vivek Srikumar. 2020 · 2020
Earlier work this paper cites.
Towards Debiasing Sentence Representations. In ACL . 5502–5515
Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2020 · 2020
Earlier work this paper cites.
Does Gender Matter? Towards Fairness in Dialogue Systems. In ICCL . 4403–4416
Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang. 2020 · 2020
Earlier work this paper cites.
End-to-End Bias Mitigation by Modelling Biases in Corpora. In ACL . 8706–8716
Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson. 2020 · 2020
Earlier work this paper cites.
The “letter”’ and the “spirit” of comparative law in the time of “artificial intelligence” and other oxymora
Rostam J Neuwirth. 2020 · 2020
Earlier work this paper cites.
Encyclopedia of World problems and Human Potential, Limited access to society’s resources
The Union of International Associations. 2020 · 2020
Earlier work this paper cites.
Reducing Non-Normative Text Generation from Language Models. In ICNLG . 374–383
Xiangyu Peng, Siyan Li, Spencer Frazier, and Mark Riedl. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, et al · 2020
Earlier work this paper cites.
Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection. In ACL . 7237–7256
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020 · 2020
Earlier work this paper cites.
Masked Language Model Scoring. In ACL . 2699–2712
Julian Salazar, Davis Liang, Toan Q Nguyen, and Katrin Kirchhoff. 2020 · 2020
Earlier work this paper cites.
Intra-processing methods for debiasing neural networks
Yash Savani, Colin White, and Naveen Sundar Govindarajulu. 2020 · 2020
Earlier work this paper cites.
Towards Controllable Biases in Language Generation. In Findings of EMNLP . Association for Computational Linguistics, 3239–3254
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2020 · 2020
Earlier work this paper cites.
Model AI Governance Framework (Second Edition)
Personal Data Protection Commission (PDPC) (Singapore). 2020 · 2020
Earlier work this paper cites.
Towards Debiasing NLU Models from Unknown Biases. In EMNLP . Association for Computational Linguistics, 7597–7610
Prasetya Ajie Utama, Nafise Sadat Moosavi, and Iryna Gurevych. 2020 · 2020
Earlier work this paper cites.
Measuring and reducing gendered correlations in pre-trained models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020 · 2020
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020 · 2020
Earlier work this paper cites.
Persistent anti-muslim bias in large language models. In AAAI/ACM Conference on AI, Ethics, and Society . 298–306
Abubakar Abid, Maheen Farooqi, and James Zou. 2021 · 2021
Earlier work this paper cites.
Mitigating Language-Dependent Ethnic Bias in BERT. In EMNLP . Association for Computational Linguistics, 533–549
Jaimeen Ahn and Alice Oh. 2021 · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?. In ACM FAccT . 610–623
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Earlier work this paper cites.
Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets. In ACL-IJCNLP . Online, 1004–1015
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021 · 2021
Earlier work this paper cites.
Reviewing the evidence: heuristics and biases
Laura Bojke, Marta Soares, Karl Claxton, Abigail Colson, et al · 2021
Cited alongside, same era.
On the Opportunities and Risks of Foundation Models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, et al · 2021
Cited alongside, same era.
FairFil: Contrastive Neural Debiasing Method for Pretrained Text Encoders. In ICLR
Pengyu Cheng, Weituo Hao, Siyang Yuan, Shijing Si, and Lawrence Carin. 2021 · 2021
Cited alongside, same era.
Corporate governance of artificial intelligence in the public interest
Peter Cihon, Jonas Schuett, and Seth D Baum. 2021 · 2021
Cited alongside, same era.
A Novel Estimator of Mutual Information for Learning to Disentangle Textual Representations. In ACL-IJCNLP . Association for Computational Linguistics, 6539–6550
Pierre Colombo, Pablo Piantanida, and Chloé Clavel. 2021 · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, et al · 2022
Later among the works it cites.
Don’t Just Clean It, Proxy Clean It: Mitigating Bias by Proxy in Pre-Trained Models. In Findings of EMNLP . Association for Computational Linguistics, 5073–5085
Swetasudha Panda, Ari Kobren, Michael Wick, and Qinlan Shen. 2022 · 2022
Later among the works it cites.
Incorporating Subjectivity into Gendered Ambiguous Pronoun (GAP) Resolution using Style Transfer. In GeBNLP . 273–281
Kartikey Pant and Tanvi Dadu. 2022 · 2022
Later among the works it cites.
Perturbation Augmentation for Fairer NLP. In EMNLP . Association for Computational Linguistics, 9496–9521
Rebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith, Douwe Kiela, and Adina Williams. 2022 · 2022
Later among the works it cites.
First the Worst: Finding Better Gender Translations During Beam Search. In Findings of ACL . 3814–3823
Danielle Saunders, Rosie Sallis, and Bill Byrne. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Proposal for a Regulation Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act), COM (2021) 206 final [AI Act]
European Commission. 2021 · 2021
Cited alongside, same era.
Book Review: Race after technology: Abolitionist tools for the New Jim Code
Madeleine Crutchley. 2021 · 2021
Cited alongside, same era.
OSCaR: Orthogonal Subspace Correction and Rectification of Biases in Word Embeddings. In EMNLP . Association for Computational Linguistics, 5034–5050
Sunipa Dev, Tao Li, Jeff M Phillips, and Vivek Srikumar. 2021 · 2021
Cited alongside, same era.
Bold: Dataset and metrics for measuring biases in open-ended language generation. In ACM FAccT . 862–872
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021 · 2021
Cited alongside, same era.
He is very intelligent, she is very beautiful? on mitigating social biases in language modelling and generation. In Findings of ACL-IJCNLP . 4534–4545
Aparna Garimella, Akhash Amarnath, Kiran Kumar, Akash Pramod Yalla, et al · 2021
Cited alongside, same era.
Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases. In AAAI/ACM Conference on AI, Ethics, and Society . ACM, 122–133
Wei Guo and Aylin Caliskan. 2021 · 2021
Cited alongside, same era.
Detect and Perturb: Neutral Rewriting of Biased and Sensitive Text via Gradient-based Decoding. In Findings of EMNLP . ACL, 4173–4181
Zexue He, Bodhisattwa Prasad Majumder, and Julian McAuley. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Large pre-trained language models contain human-like biases of what is right and wrong to do
Patrick Schramowski, Cigdem Turan, Nico Andersen, et al · 2022
Later among the works it cites.
Quantifying Social Biases Using Templates is Unreliable. In Workshop on Trustworthy and Socially Responsible Machine Learning, NeurIPS
Preethi Seshadri, Pouya Pezeshkpour, and Sameer Singh. 2022 · 2022
Later among the works it cites.
Blenderbot 3: a deployed conversational agent that continually learns to responsibly engage
Kurt Shuster, Jing Xu, Mojtaba Komeili, Da Ju, Eric Michael Smith, Stephen Roller, et al · 2022
Later among the works it cites.
“I’m sorry to hear that”: Finding New Biases in Language Models with a Holistic Descriptor Dataset. In EMNLP . ACL, 9180–9211
Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022 · 2022
Later among the works it cites.
ChatGPT: Optimizing language models for dialogue
Team OpenAI 2022 · 2022
Later among the works it cites.
Text Style Transfer for Bias Mitigation using Masked Language Modeling. In NAACL: HLT-SRW . Association for Computational Linguistics, 163–171
Ewoenam Kwaku Tokpo and Toon Calders. 2022 · 2022
Later among the works it cites.
Harnessing Indigenous Tweets: The Reo Māori Twitter Corpus
David Trye, Te Taka Keegan, Paora Mato, and Mark Apperley. 2022a · 2022
Later among the works it cites.
A Hybrid Architecture for Labelling Bilingual Māori-English Tweets. In Findings of AACL-IJCNLP 2022 . ACL, 119–130
David Trye, Vithya Yogarajan, Jemma König, Te Taka Keegan, David Bainbridge, and Mark Apperley. 2022b · 2022
Later among the works it cites.
Recommendation on the Ethics of Artificial Intelligence (Paris)
Scientific United Nations Educational and Cultural Organization (UNESCO). 2022 · 2022
Later among the works it cites.
Pay Attention to Your Tone: Introducing a New Dataset for Polite Language Rewrite
Xun Wang, Tao Ge, Allen Mao, Yuki Li, Furu Wei, and Si-Qing Chen. 2022 · 2022
Later among the works it cites.
Social bias, discrimination and inequity in healthcare: mechanisms, implications and recommendations
Craig S Webster, Saana Taylor, Courtney Thomas, and Jennifer M Weller. 2022 · 2022
Later among the works it cites.
Lessons learned from developing a COVID-19 algorithm governance framework in Aotearoa New Zealand
Daniel Wilson, Frith Tweedie, Juliet Rumball-Smith, Kevin Ross, et al · 2022
Later among the works it cites.
Data and Model Bias in Artificial Intelligence for Healthcare Applications in New Zealand
Vithya Yogarajan, Gillian Dobbie, Sharon Leitch, Te Taka Keegan, Joshua Bensemann, Michael Witbrock, et al · 2022
Later among the works it cites.
Towards Identifying Social Bias in Dialog Systems: Framework, Dataset, and Benchmark. In Findings of EMNLP . 3576–3591
Jingyan Zhou, Jiawen Deng, Fei Mi, Yitong Li, Yasheng Wang, et al · 2022
Later among the works it cites.
OMPCSA online resource hub
The Prime Minister’s Chief Science Advisor. 2023 · 2023
Closest in time.
Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting Model. In ACL . 4486–4506
Chantal Amrhein, Florian Schottmann, Rico Sennrich, and Samuel Läubli. 2023 · 2023
Closest in time.
Holistic Evaluation of Language Models
Rishi Bommasani, Percy Liang, and Tony Lee. 2023 · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, et al · 2023
Closest in time.
On the Independence of Association Bias and Empirical Fairness in Language Models. In FAccT . ACM, 370–378
Laura Cabello, Anna Katrine Jørgensen, and Anders Søgaard. 2023 · 2023
Closest in time.
Explained: The Digital India Act 2023
Sanhita Chauriha. 2023 · 2023
Closest in time.
Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions. In ACL . 575–593
John Joon Young Chung, Ece Kamar, and Saleema Amershi. 2023 · 2023
Closest in time.
Evaluation of GPT-3.5 and GPT-4 for supporting real-world information needs in healthcare delivery
Debadutta Dash, Rahul Thapa, Juan M Banda, Akshay Swaminathan, et al · 2023
Closest in time.
Building Stereotype Repositories with LLMs and Community Engagement for Scale and Depth
Sunipa Dev, Akshita Jha, Jaya Goyal, Dinesh Tewari, et al · 2023
Closest in time.
Queer people are people first: Deconstructing sexual identity stereotypes in large language models
Harnoor Dhingra, Preetiha Jayashanker, Sayali Moghe, and Emma Strubell. 2023 · 2023
Closest in time.
Improving gender fairness of pre-trained language models without catastrophic forgetting. In ACL . 1249–126
Zahra Fatemi, Chen Xing, Wenhao Liu, and Caiming Xiong. 2023 · 2023
Closest in time.
WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models. In ACL . 9126–9140
Virginia Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May. 2023 · 2023
Closest in time.
Emilio Ferrara. 2023 · 2023
Closest in time.
Bias and Fairness in Large Language Models: A Survey
Isabel Gallegos, Ryan Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, et al · 2023
Closest in time.
Gender-tuning: Empowering Fine-tuning for Debiasing Pre-trained Language Models. In Findings of ACL . 5448–5458
Somayeh Ghanbarzadeh, Yan Huang, Hamid Palangi, Radames Cruz Moreno, and Hamed Khanpour. 2023 · 2023
Closest in time.
Detoxifying Text with MaRCo: Controllable Revision with Experts and Anti-Experts. In ACL . 228–242
Skyler Hallinan, Alisa Liu, Yejin Choi, and Maarten Sap. 2023 · 2023
Closest in time.
Modular and on-demand bias mitigation with attribute-removal subnetworks. In Findings of ACL . 6192–6214
Lukas Hauzenberger, Shahed Masoudian, Deepak Kumar, Markus Schedl, and Navid Rekabsaz. 2023 · 2023
Closest in time.
Executive Order on Safe, Secure, and Trustworthy Artificial Intelligence
The White House. 2023 · 2023
Closest in time.
TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models
Yue Huang, Qihui Zhang, Lichao Sun, et al · 2023
Closest in time.
Shielded Representations: Protecting Sensitive Attributes Through Iterative Gradient-Based Projection. In Findings of ACL . 5961–597
Shadi Iskander, Kira Radinsky, and Yonatan Belinkov. 2023 · 2023
Closest in time.
The development of a labelled te reo Māori–English bilingual database for language technology
Jesin James, Isabella Shields, Vithya Yogarajan, Peter Keegan, Catherine Watson, et al · 2023
Closest in time.
Learn What NOT to Learn: Towards Generative Safety in Chatbots
Leila Khalatbari, Yejin Bang, Dan Su, Willy Chung, Saeed Ghadimi, Hossein Sameti, and Pascale Fung. 2023 · 2023
Closest in time.
Critic-Guided Decoding for Controlled Text Generation. In Findings of ACL . 4598–4612
Minbeom Kim, Hwanhee Lee, Kang Min Yoo, Joonsuk Park, Hwaran Lee, and Kyomin Jung. 2023 · 2023
Closest in time.
A Survey on Fairness in Large Language Models
Yingji Li, Mengnan Du, Rui Song, Xin Wang, and Ying Wang. 2023a · 2023
Closest in time.
Yunqi Li and Yongfeng Zhang. 2023 · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in NLP
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023 · 2023
Closest in time.
Using In-Context Learning to Improve Dialogue Safety
Nicholas Meade, Spandana Gella, Devamanyu Hazarika, Prakhar Gupta, Di Jin, Siva Reddy, Yang Liu, and Dilek Hakkani-Tür. 2023 · 2023
Closest in time.
Te Kāhui Raraunga
Māori Data Governance Model. June, 2023 · 2023
Closest in time.
Biases in Large Language Models: Origins, Inventory and Discussion
Roberto Navigli, Simone Conia, and Björn Ross. 2023 · 2023
Closest in time.
The EU artificial intelligence act: regulating subliminal AI systems
Rostam J Neuwirth. 2023a · 2023
Closest in time.
Prohibited artificial intelligence practices in the proposed EU artificial intelligence act (AIA)
Rostam J Neuwirth. 2023b · 2023
Closest in time.
Interim Administrative Measures for Generative Artificial Intelligence (AI) Services
Cybersecurity Administration of China (CAC). 2023 · 2023
Closest in time.
P7003 Algorithmic Bias Considerations
Institute of Electrical and Electronics Engineers (IEEE) Standards Association. 2023 · 2023
Closest in time.
Consolidated Working Draft of the Framework Convention on Artificial Intelligence, Human Rights, Democracy and the Rule of Law, CAI(2023)18
Council of Europe Committee on Artificial Intelligence (CAI). 2023 · 2023
Closest in time.
BLIND: Bias removal with no demographics. In ACL . 8801–8821
Hadas Orgad and Yonatan Belinkov. 2023 · 2023
Closest in time.
Never Too Late to Learn: Regularizing Gender Bias in Coreference Resolution. In ACM-ICWSDM . 15–23
SunYoung Park, Kyuri Choi, Haeun Yu, and Youngjoong Ko. 2023 · 2023
Closest in time.
A Trip Towards Fairness: Bias and De-Biasing in Large Language Models
Leonardo Ranaldi, Elena Sofia Ruzzetti, Davide Venditti, Dario Onorati, and Fabio Massimo Zanzotto. 2023 · 2023
Closest in time.
The Tail Wagging the Dog: Dataset Construction Biases of Social Bias Benchmarks. In ACL . Toronto, Canada, 1373–1386
Nikil Selvam, Sunipa Dev, Daniel Khashabi, Tushar Khot, and Kai-Wei Chang. 2023 · 2023
Closest in time.
Learning to Generate Equitable Text in Dialogue from Biased Training Data. In ACL . 2898–2917
Anthony Sicilia and Malihe Alikhani. 2023 · 2023
Closest in time.
Language Models Get a Gender Makeover: Mitigating Gender Bias with Few-Shot Data Interventions. In ACL . 340––351
Himanshu Thakur, Atishay Jain, Praneetha Vaddamanu, Paul Pu Liang, and Louis-Philippe Morency. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, et al · 2023
Closest in time.
Nationality Bias in Text Generation. In EACL . Association for Computational Linguistics, 116–122
Pranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao Huang, and Shomir Wilson. 2023 · 2023
Closest in time.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, et al · 2023
Closest in time.
Multi-Target Multiplicity: Flexibility and Fairness in Target Specification under Resource Constraints. In ACM FAccT . 297–311
Jamelle Watson-Daniels, Solon Barocas, Jake M Hofman, and Alexandra Chouldechova. 2023 · 2023
Closest in time.
Adept: A debiasing prompt framework. In AAAI , Vol. 37. 10780–10788
Ke Yang, Charles Yu, Yi R Fung, Manling Li, and Heng Ji. 2023 · 2023
Closest in time.
Should We Attend More or Less? Modulating Attention for Fairness
Abdelrahman Zayed, Goncalo Mordido, Samira Shabanian, and Sarath Chandar. 2023 · 2023
Closest in time.
Click: Controllable Text Generation with Sequence Likelihood Contrastive Learning. In Findings of ACL . 1022–1040
Chujie Zheng, Pei Ke, Zheng Zhang, and Minlie Huang. 2023 · 2023
Closest in time.
Causal-debias: Unifying debiasing in pretrained language models and fine-tuning via causal invariant learning. In ACL . 4227–4241
Fan Zhou, Yuzhou Mao, Liu Yu, Yi Yang, and Ting Zhong. 2023 · 2023
Closest in time.
Theories of “gender” in nlp bias research. In FAccT . ACM, 2083–2102
Hannah Devinney, Jenny Björklund, and Henrik Björklund. 2022 · 2083
Closest in time.
BBQ: A hand-built bias benchmark for question answering. In Findings of ACL . 2086–2105
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022 · 2086
Closest in time.