Fetching the paper…
Reading the bibliography…
Recent progress in large language models (LLMs) has led to their widespread adoption in various domains.
A framework for understanding unintended consequences of machine learning
Harini Suresh and John V. Guttag. 2019 · 1901
Earlier work this paper cites.
Language (technology) is power: A critical survey of "bias" in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna M. Wallach. 2020 · 2005
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Earlier work this paper cites.
What’s in a name? Reducing bias in bios without access to protected attributes
Alexey Romanov, Maria De-Arteaga, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, Anna Rumshisky, and Adam Kalai. 2019 · 2019
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Unqovering stereotyping biases via underspecified questions
Tao Li, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Vivek Srikumar. 2020 · 2020
Earlier work this paper cites.
Unique names in china: Insights from research in japan—commentary: Increasing need for uniqueness in contemporary china: Empirical evidence
Yuji Ogihara. 2020 · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Earlier work this paper cites.
Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021 · 2021
Earlier work this paper cites.
BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021 · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel. 2021 · 2021
Earlier work this paper cites.
Openai claims to have mitigated bias and toxicity in gpt-3
Kyle Wiggers. 2021 · 2021
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 · 2022
Cited alongside, same era.
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022 · 2022
Cited alongside, same era.
A pathway towards responsible ai generated content
Chen Chen, Jie Fu, and Lingjuan Lyu. 2023 · 2023
Cited alongside, same era.
Safe rlhf: Safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. 2023 · 2023
Cited alongside, same era.
Bias incidents against muslims, jews on the rise in us amid middle east war, new data shows
Meredith Deliso. 2023 · 2023
I tried the ai novel-writing tool everyone hates, and it’s better than i expected
Adi Robertson. 2023 · 2023
Later among the works it cites.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2023 · 2023
Later among the works it cites.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Paul Röttger, Hannah Kirk, Bertram Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2023 · 2023
Later among the works it cites.
The unequal opportunities of large language models: Examining demographic biases in job recommendations by chatgpt and llama
Abel Salinas, Parth Shah, Yuzhong Huang, Robert McCormack, and Fred Morstatter. 2023 · 2023
Later among the works it cites.
Sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
How generative ai – a technology catalyst – is revolutionizing healthcare
Bill Frist. 2023 · 2023
Cited alongside, same era.
Chatgpt is going to change education, not destroy it
Will Douglas Heaven. 2023 · 2023
Cited alongside, same era.
Khyati Khandelwal, Manuel Tonneau, Andrew M. Bean, Hannah Rose Kirk, and Scott A. Hale. 2023 · 2023
Cited alongside, same era.
Gpt-4 beats 90% of lawyers trying to pass the bar
John Koetsier. 2023 · 2023
Cited alongside, same era.
Introducing code llama, an ai tool for coding
Meta. 2023 · 2023
Cited alongside, same era.
OpenAI. 2023 · 2023
Cited alongside, same era.
Chatgpt: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope
Partha Pratim Ray. 2023 · 2023
Cited alongside, same era.
Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla, Jess Gallegos, Andrew Smart, Emilio Garcia, and Gurleen Virk. 2023 · 2023
Later among the works it cites.
Ai storytellers: How generative models are penning bestselling novels
Times Now Digital. 2023 · 2023
Later among the works it cites.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. 2023 · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023 · 2023
Later among the works it cites.
Unveiling the implicit toxicity in large language models
Jiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang, Chengfei Li, Jinfeng Bai, and Minlie Huang. 2023 · 2023
Later among the works it cites.
Claude 2.0
Anthropic. 2024 · 2024
Closest in time.
LLM trustworthy leaderboard
Secure Learning Lab. 2024 · 2024
Closest in time.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022 · 2086
Closest in time.