Fetching the paper…
Reading the bibliography…
This research introduces STAR, a sociotechnical framework that improves on current best practices for red teaming safety of large language models.
Scikit-learn: Machine learning in python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay. 2011 · 2011
Earlier work this paper cites.
Notional supply chain risk management practices for federal information systems
Jon Boyens, Celia Paulsen, Nadya Bartol, Rama Moorthy, and Stephanie Shankles. 2012 · 2012
Earlier work this paper cites.
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty. 2015 · 2015
Earlier work this paper cites.
Beat the machine: Challenging humans to find a predictive model’s “unknown unknowns”
Joshua Attenberg, Panos Ipeirotis, and Foster Provost. 2015 · 2015
Earlier work this paper cites.
Fairness and Abstraction in Sociotechnical Systems
Andrew D. Selbst, Danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi. 2019 · 2019
Earlier work this paper cites.
UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
Leland McInnes, John Healy, and James Melville. 2020 · 2020
Earlier work this paper cites.
The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With Reality
Mitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S. Bernstein. 2021 · 2021
Earlier work this paper cites.
Understanding international perceptions of the severity of harmful content online
Jialun Aaron Jiang, Morgan Klaus Scheuerman, Casey Fiesler, and Jed R. Brubaker. 2021 · 2021
Earlier work this paper cites.
Bot-Adversarial Dialogue for safe conversational agents
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2021 · 2021
Earlier work this paper cites.
Toward user-driven algorithm auditing: Investigating users’ strategies for uncovering harmful algorithmic behavior
Alicia DeVos, Aditi Dhabalia, Hong Shen, Kenneth Holstein, and Motahhare Eslami. 2022 · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark. 2022 · 2022
Earlier work this paper cites.
Is your toxicity my toxicity? Exploring the impact of rater identity on toxicity annotation
Nitesh Goyal, Ian D. Kivlichan, Rachel Rosen, and Lucy Vasserman. 2022 · 2022
Earlier work this paper cites.
Red Teaming Language Models with Language Models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 · 2022
Cited alongside, same era.
The “problem” of human label variation: On ground truth in data, modeling and evaluation
Barbara Plank. 2022 · 2022
Cited alongside, same era.
LaMDA: Language Models for Dialog Applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le. 2022 · 2022
Cited alongside, same era.
Challenges in evaluating AI systems
Anthropic. 2023 · 2023
Cited alongside, same era.
Generative ai prohibited use policy
Google Privacy & Terms. 2023 · 2023
Later among the works it cites.
Fact sheet: Biden-Harris administration announces new actions to promote responsible AI innovation that protects Americans’ rights and safety
White House. 2023 · 2023
Later among the works it cites.
Red teaming ChatGPT via jailbreaking: Bias, robustness, reliability and toxicity
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing. 2023 · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023 · 2023
Later among the works it cites.
Stela: a community-centred approach to norm elicitation for ai alignment
Stevie Bergman, Nahema Marchal, John Mellor, Shakir Mohamed, Iason Gabriel, and William Isaac. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
DICES Dataset: Diversity in Conversational AI Evaluation for Safety
Lora Aroyo, Alex S. Taylor, Mark Diaz, Christopher M. Homan, Alicia Parrish, Greg Serapio-Garcia, Vinodkumar Prabhakaran, and Ding Wang. 2023 · 2023
Cited alongside, same era.
Representation in AI Evaluations
A. Stevie Bergman, Lisa Anne Hendricks, Maribeth Rauh, Boxi Wu, William Agnew, Markus Kunesch, Isabella Duan, Iason Gabriel, and William Isaac. 2023 · 2023
Cited alongside, same era.
Living guidelines for generative AI — why scientists must oversee its use
Claudi L. Bockting, Eva A. M. Van Dis, Robert Van Rooij, Willem Zuidema, and Johan Bollen. 2023 · 2023
Cited alongside, same era.
Content categories
Google Cloud. 2023 · 2023
Cited alongside, same era.
Christopher M. Homan, Greg Serapio-Garcia, Lora Aroyo, Mark Diaz, Alicia Parrish, Vinodkumar Prabhakaran, Alex S. Taylor, and Ding Wang. 2023 · 2023
Cited alongside, same era.
Gpt-4v(ision) system card
OpenAI. 2023 · 2023
Cited alongside, same era.
Adversarial Nibbler: A data-centric challenge for improving the safety of text-to-image models
Alicia Parrish, Hannah Rose Kirk, Jessica Quaye, Charvi Rastogi, Max Bartolo, Oana Inel, Juan Ciro, Rafael Mosquera, Addison Howard, Will Cukierski, D. Sculley, Vijay Janapa Reddi, and Lora Aroyo. 2023 · 2023
Cited alongside, same era.
AART: AI-assisted red-teaming with diverse data generation for new LLM-powered applications
Bhaktipriya Radharapu, Kevin Robinson, Lora Aroyo, and Preethi Lahoti. 2023 · 2023
Cited alongside, same era.
Closest in time.
Black-box access is insufficient for rigorous AI audits
Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin Von Hagen, Silas Alberti, Alan Chan, Qinyi Sun, Michael Gerovitch, David Bau, Max Tegmark, David Krueger, and Dylan Hadfield-Menell. 2024 · 2024
Closest in time.
Red-teaming for generative AI: Silver bullet or security theater?
Michael Feffer, Anusha Sinha, Zachary C. Lipton, and Hoda Heidari. 2024 · 2024
Closest in time.
Gecko: Versatile Text Embeddings Distilled from Large Language Models
Jinhyuk Lee, Zhuyun Dai, Xiaoqi Ren, Blair Chen, Daniel Cer, Jeremy R. Cole, Kai Hui, Michael Boratko, Rajvi Kapadia, Wen Ding, Yi Luan, Sai Meher Karthik Duddu, Gustavo Hernandez Abrego, Weiqiang Shi, Nithi Gupta, Aditya Kusupati, Prateek Jain, Siddhartha Reddy Jonnalagadda, Ming-Wei Chang, and Iftekhar Naim. 2024 · 2024
Closest in time.
Taishi Nakamura, Mayank Mishra, Simone Tedeschi, Yekun Chai, Jason T. Stillerman, Felix Friedrich, Prateek Yadav, Tanmay Laud, Vu Minh Chien, Terry Yue Zhuo, Diganta Misra, Ben Bogin, Xuan-Son Vu, Marzena Karpinska, Arnav Varma Dantuluri, Wojciech Kusa, Tommaso Furlanello, Rio Yokota, Niklas Muennighoff, Suhas Pai, Tosin Adewumi, Veronika Laippala, Xiaozhe Yao, Adalberto Junior, Alpay Ariyak, Aleksandr Drozd, Jordan Clive, Kshitij Gupta, Liangyu Chen, Qi Sun, Ken Tsui, Noah Persaud, Nour Fahmy, Tianlong Chen, Mohit Bansal, Nicolo Monti, Tai Dang, Ziyang Luo, Tien-Tung Bui, Roberto Navigli, Virendra Mehta, Matthew Blumberg, Victor May, Huu Nguyen, and Sampo Pyysalo. 2024 · 2024
Closest in time.
Diversity-aware annotation for conversational AI safety
Alicia Parrish, Vinodkumar Prabhakaran, Lora Aroyo, Mark Díaz, Christopher M. Homan, Greg Serapio-García, Alex S. Taylor, and Ding Wang. 2024 · 2024
Closest in time.
Rainbow Teaming: Open-ended generation of diverse adversarial prompts
Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H. Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, Tim Rocktäschel, and Roberta Raileanu. 2024 · 2024
Closest in time.
Low-resource languages jailbreak GPT-4
Zheng-Xin Yong, Cristina Menghini, and Stephen H. Bach. 2024 · 2024
Closest in time.