Fetching the paper…
Reading the bibliography…
This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group.
Speech act theory and pragmatics, 1980
Manfred Bierwisch John R. Searle, Ferenc Kiefer · 1980
Earlier work this paper cites.
Dignity and defamation: The visibility of hate
Jeremy Waldron · 2009
Earlier work this paper cites.
A 61-million-person experiment in social influence and political mobilization
Robert M Bond, Christopher J Fariss, Jason J Jones, Adam DI Kramer, Cameron Marlow, Jaime E Settle, and James H Fowler · 2012
Earlier work this paper cites.
Rise of concerns about ai: reflections and directions
Thomas G Dietterich and Eric J Horvitz · 2015
Earlier work this paper cites.
Evidencing the harms of hate speech
Katharine Gelber and Luke McNamara · 2015
Earlier work this paper cites.
The ladder: A reliable leaderboard for machine learning competitions
Avrim Blum and Moritz Hardt · 2015
Earlier work this paper cites.
Concrete problems in ai safety, 2016
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Decoding social media speak: developing a speech act theory research agenda, 2016
Ko de Ruyter Stephan Ludwig · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Tauman Kalai · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Steps toward robust artificial intelligence
Thomas G. Dietterich · 2017
Earlier work this paper cites.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation, 2018
Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, Peter Eckersley, Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff, Bobby Filar, Hyrum Anderson, Heather Roff, Gregory C. Allen, Jacob Steinhardt, Carrick Flynn, Seán Ó hÉigeartaigh, Simon Beard, Haydn Belfield, Sebastian Farquhar, Clare Lyle, Rebecca Crootof, Owain Evans, Michael Page, Joanna Bryson, Roman Yampolskiy, and Dario Amodei · 2018
Earlier work this paper cites.
Sharing as speech act, 2018
Emanuele Arielli · 2018
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng · 2019
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston · 2019
Earlier work this paper cites.
Entanglements and exploits: Sociotechnical security as an analytic framework
Matt Goerzen, Elizabeth Anne Watkins, and Gabrielle Lim · 2019
Earlier work this paper cites.
Release strategies and the social impacts of language models, 2019
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, Miles McCain, Alex Newhouse, Jason Blazakis, Kris McGuffie, and Jasmine Wang · 2019
Earlier work this paper cites.
Ai-mediated communication: How the perception that profile text was written by ai affects trustworthiness
Maurice Jakesch, Megan French, Xiao Ma, Jeffrey T. Hancock, and Mor Naaman · 2019
Earlier work this paper cites.
Algorithmic extremism: Examining youtube’s rabbit hole of radicalization, 2019
Mark Ledwich and Anna Zaitsev · 2019
Earlier work this paper cites.
Fairness and abstraction in sociotechnical systems
Andrew D Selbst, Danah Boyd, Sorelle A Friedler, Suresh Venkatasubramanian, and Janet Vertesi · 2019
Earlier work this paper cites.
Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai, 2019
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador García, Sergio Gil-López, Daniel Molina, Richard Benjamins, Raja Chatila, and Francisco Herrera · 2019
Earlier work this paper cites.
The oxford handbook of terrorism, 2019
Erica Chenoweth, Richard English, Andreas Gofas, and Stathis N. Kalyvas · 2019
Earlier work this paper cites.
Terrorism and ideology: Cracking the nut, 2019
John Horgan Donald Holbrook · 2019
Earlier work this paper cites.
Mlperf inference benchmark, 2020
Vijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson, Guenther Schmuelling, Carole-Jean Wu, Brian Anderson, Maximilien Breughe, Mark Charlebois, William Chou, Ramesh Chukka, Cody Coleman, Sam Davis, Pan Deng, Greg Diamos, Jared Duke, Dave Fick, J. Scott Gardner, Itay Hubara, Sachin Idgunji, Thomas B. Jablin, Jeff Jiao, Tom St. John, Pankaj Kanwar, David Lee, Jeffery Liao, Anton Lokhmotov, Francisco Massa, Peng Meng, Paulius Micikevicius, Colin Osborne, Gennady Pekhimenko, Arun Tejusve Raghunath Rajan, Dilip Sequeira, Ashish Sirasao, Fei Sun, Hanlin Tang, Michael Thomson, Frank Wei, Ephrem Wu, Lingjie Xu, Koichi Yamada, Bing Yu, George Yuan, Aaron Zhong, Peizhao Zhang, and Yuchen Zhou · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Suicide and self-harm content on instagram: A systematic scoping review
Jacobo Picardo, Sarah K. McKenzie, Sunny Collings, and Gabrielle Jenkin · 2020
Earlier work this paper cites.
Toward trustworthy ai development: Mechanisms for supporting verifiable claims, 2020
Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner, Ruth Fong, Tegan Maharaj, Pang Wei Koh, Sara Hooker, Jade Leung, Andrew Trask, Emma Bluemke, Jonathan Lebensold, Cullen O’Keefe, Mark Koren, Théo Ryffel, JB Rubinovitz, Tamay Besiroglu, Federica Carugati, Jack Clark, Peter Eckersley, Sarah de Haas, Maritza Johnson, Ben Laurie, Alex Ingerman, Igor Krawczuk, Amanda Askell, Rosario Cammarota, Andrew Lohn, David Krueger, Charlotte Stix, Peter Henderson, Logan Graham, Carina Prunkl, Bianca Martin, Elizabeth Seger, Noa Zilberman, Seán Ó hÉigeartaigh, Frens Kroeger, Girish Sastry, Rebecca Kagan, Adrian Weller, Brian Tse, Elizabeth Barnes, Allan Dafoe, Paul Scharre, Ariel Herbert-Voss, Martijn Rasser, Shagun Sodhani, Carrick Flynn, Thomas Krendl Gilbert, Lisa Dyer, Saif Khan, Yoshua Bengio, and Markus Anderljung · 2020
Earlier work this paper cites.
Software verification and validation of safe autonomous cars: A systematic literature review
Nijat Rajabli, Francesco Flammini, Roberto Nardone, and Valeria Vittorini · 2020
Earlier work this paper cites.
Algorithmic content moderation: Technical and political challenges in the automation of platform governance
Robert Gorwa, Reuben Binns, and Christian Katzenbach · 2020
Earlier work this paper cites.
Ml4h auditing: From paper to practice
Luis Oala, Jana Fehr, Luca Gilli, Pradeep Balachandran, Alixandro Werneck Leite, Saul Calderon-Ramirez, Danny Xie Li, Gabriel Nobis, Erick Alejandro Muñoz Alvarado, Giovanna Jaramillo-Gutierrez, Christian Matek, Arun Shroff, Ferath Kherif, Bruno Sanguinetti, and Thomas Wiegand · 2020
Earlier work this paper cites.
Mlperf training benchmark, 2020
Peter Mattson, Christine Cheng, Cody Coleman, Greg Diamos, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, David Brooks, Dehao Chen, Debojyoti Dutta, Udit Gupta, Kim Hazelwood, Andrew Hock, Xinyuan Huang, Atsushi Ike, Bill Jia, Daniel Kang, David Kanter, Naveen Kumar, Jeffery Liao, Guokai Ma, Deepak Narayanan, Tayo Oguntebi, Gennady Pekhimenko, Lillian Pentecost, Vijay Janapa Reddi, Taylor Robie, Tom St. John, Tsuguchika Tabaru, Carole-Jean Wu, Lingjie Xu, Masafumi Yamazaki, Cliff Young, and Matei Zaharia · 2020
Earlier work this paper cites.
Towards ecologically valid research on language user interfaces, 2020
Harm de Vries, Dzmitry Bahdanau, and Christopher Manning · 2020
Earlier work this paper cites.
Datasheets for datasets, 2021
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III au2, and Kate Crawford · 2021
Earlier work this paper cites.
Interactive storytelling for children: A case-study of design and development considerations for ethical conversational ai
Jennifer Chubb, Sondess Missaoui, Shauna Concannon, Liam Maloney, and James Alfred Walker · 2021
Earlier work this paper cites.
Training data leakage analysis in language models
Huseyin A Inan, Osman Ramadan, Lukas Wutschitz, Daniel Jones, Victor Rühle, James Withers, and Robert Sim · 2021
Earlier work this paper cites.
Bbq: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman · 2021
Earlier work this paper cites.
Honest: Measuring hurtful sentence completion in language models
Debora Nozza, Federico Bianchi, Dirk Hovy, et al · 2021
Earlier work this paper cites.
Bold: Dataset and metrics for measuring biases in open-ended language generation
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta · 2021
Earlier work this paper cites.
Towards understanding and mitigating social biases in language models
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2021
Earlier work this paper cites.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Earlier work this paper cites.
HateCheck: Functional tests for hate speech detection models
Paul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert · 2021
Earlier work this paper cites.
Beyond the individual: governing ai’s societal harm
Nathalie A Smuha · 2021
Earlier work this paper cites.
The 4cs: Classifying online risk to children, 2021
Sonia Livingstone and Mariya Stoilova · 2021
Earlier work this paper cites.
Understanding online hate: Vsp regulation and the broader context
Bertie Vidgen, Emily Burden, and Helen Margetts · 2021
Earlier work this paper cites.
Gender and representation bias in GPT-3 generated stories
Li Lucy and David Bamman · 2021
Earlier work this paper cites.
Artificial intelligence in communication impacts language and social relationships. arxiv
J Hohenstein, D DiFranzo, RF Kizilcec, Z Aghajari, H Mieczkowski, K Levy, M Naaman, J Hancock, and M Jung · 2021
Earlier work this paper cites.
Measurement and fairness
Abigail Z. Jacobs and Hanna Wallach · 2021
Earlier work this paper cites.
Governing ai safety through independent audits
Gregory Falco, Ben Shneiderman, Julia Badger, Ryan Carrier, Anton Dahbura, David Danks, Martin Eling, Alwyn Goodloe, Jerry Gupta, Christopher Hart, et al · 2021
Earlier work this paper cites.
Ethics-based auditing to develop trustworthy ai
Jakob Mökander and Luciano Floridi · 2021
Earlier work this paper cites.
Machine learning for health: algorithm auditing & quality control
Luis Oala, Andrew G Murchison, Pradeep Balachandran, Shruti Choudhary, Jana Fehr, Alixandro Werneck Leite, Peter G Goldschmidt, Christian Johner, Elora DM Schörverth, Rose Nakasi, et al · 2021
Earlier work this paper cites.
What will it take to fix benchmarking in natural language understanding?, 2021
Samuel R. Bowman and George E. Dahl · 2021
Earlier work this paper cites.
Dynabench: Rethinking benchmarking in nlp, 2021
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams · 2021
Earlier work this paper cites.
Unitary ai, detoxify, 2021
Unitary AI · 2021
Earlier work this paper cites.
The situation awareness framework for explainable ai (safe-ai) and human factors considerations for xai systems
Lindsay Sanneman and Julie A. Shah · 2022
Earlier work this paper cites.
The accidental taxonomist
Heather Hedden · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations, 2022
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2022
Cited alongside, same era.
Uncalibrated models can improve human-ai collaboration, 2022
Kailas Vodrahalli, Tobias Gerstenberg, and James Zou · 2022
Cited alongside, same era.
A new generation of perspective api: Efficient multilingual character-level transformers, 2022
Alyssa Lees, Vinh Q. Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman · 2022
Cited alongside, same era.
Into the laions den: Investigating hate in multimodal datasets, 2023
Abeba Birhane, Vinay Prabhu, Sang Han, Vishnu Naresh Boddeti, and Alexandra Sasha Luccioni · 2023
Later among the works it cites.
Dmlr: Data-centric machine learning research – past, present and future, 2023
Luis Oala, Manil Maskey, Lilith Bat-Leah, Alicia Parrish, Nezihe Merve Gürel, Tzu-Sheng Kuo, Yang Liu, Rotem Dror, Danilo Brajovic, Xiaozhe Yao, Max Bartolo, William A Gaviria Rojas, Ryan Hileman, Rainier Aliment, Michael W. Mahoney, Meg Risdal, Matthew Lease, Wojciech Samek, Debojyoti Dutta, Curtis G Northcutt, Cody Coleman, Braden Hancock, Bernard Koch, Girmaw Abebe Tadesse, Bojan Karlaš, Ahmed Alaa, Adji Bousso Dieng, Natasha Noy, Vijay Janapa Reddi, James Zou, Praveen Paritosh, Mihaela van der Schaar, Kurt Bollacker, Lora Aroyo, Ce Zhang, Joaquin Vanschoren, Isabelle Guyon, and Peter Mattson · 2023
Later among the works it cites.
Dices dataset: Diversity in conversational ai evaluation for safety, 2023
Lora Aroyo, Alex S. Taylor, Mark Diaz, Christopher M. Homan, Alicia Parrish, Greg Serapio-Garcia, Vinodkumar Prabhakaran, and Ding Wang · 2023
Later among the works it cites.
The shifted and the overlooked: A task-oriented investigation of user-gpt interactions, 2023
Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, and Jiawei Han · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hatemoji: A test suite and adversarially-generated dataset for benchmarking and detecting emoji-based hate
Hannah Kirk, Bertie Vidgen, Paul Rottger, Tristan Thrush, and Scott Hale · 2022
Cited alongside, same era.
Critical perspectives: A benchmark revealing pitfalls in PerspectiveAPI
Lucas Rosenblatt, Lorena Piedras, and Julia Wilkins · 2022
Cited alongside, same era.
Chatbots and mental health: insights into the safety of generative ai
Julian De Freitas, Ahmet Kaan Uğuralp, Zeliha Oğuz-Uğuralp, and Stefano Puntoni · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark · 2022
Cited alongside, same era.
Current and near-term ai as a potential existential risk factor
Benjamin S. Bucknall and Shiri Dori-Hacohen · 2022
Cited alongside, same era.
How to certify machine learning based safety-critical systems? a systematic literature review
Florian Tambon, Gabriel Laberge, Le An, Amin Nikanjam, Paulina Stevia Nouwou Mindom, Yann Pequignot, Foutse Khomh, Giulio Antoniol, Ettore Merlo, and François Laviolette · 2022
Cited alongside, same era.
Chatgpt: Fundamentals, applications and social impacts
Malak Abdullah, Alia Madain, and Yaser Jararweh · 2022
Cited alongside, same era.
Jacob Metcalf, Emanuel Moss, Ranjit Singh, Emnet Tafese, and Elizabeth Anne Watkins · 2022
Cited alongside, same era.
Toward comprehensive risk assessments and assurance of ai-based systems
Heidy Khlaaf · 2023
Later among the works it cites.
Palm 2 technical report, 2023
Rohan Anil et al · 2023
Later among the works it cites.
Executive order on the safe, secure, and trustworthy development and use of artificial intelligence, 2023
Joseph R Biden · 2023
Later among the works it cites.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation, 2023
Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang · 2023
Later among the works it cites.
Red-teaming large language models using chain of utterances for safety-alignment, 2023
Rishabh Bhardwaj and Soujanya Poria · 2023
Later among the works it cites.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset, 2023
Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Chi Zhang, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Later among the works it cites.
Auditnlg: Auditing generative ai language modeling for trustworthiness, 2023
SalesForce · 2023
Later among the works it cites.
A holistic approach to undesired content detection in the real world, 2023
Todor Markov, Chong Zhang, Sandhini Agarwal, Tyna Eloundou, Teddy Lee, Steven Adler, Angela Jiang, and Lilian Weng · 2023
Later among the works it cites.
Foundation metrics for evaluating effectiveness of healthcare conversations powered by generative ai
Mahyar Abbasian, Elahe Khatibi, Iman Azimi, David Oniani, Zahra Shakeri Hossein Abad, Alexander Thieme, Ram Sriram, Zhongqi Yang, Yanshan Wang, Bryant Lin, et al · 2024
Closest in time.
Not my voice! a taxonomy of ethical and safety harms of speech generators, 2024
Wiebke Hutiri, Oresiti Papakyriakopoulos, and Alice Xiang · 2024
Closest in time.
Iso/iec/ieee 24748-7000:2022. systems and software engineering life cycle management part 7000: Standard model process for addressing ethical concerns during system design, 2024
ISO/IEC/IEEE · 2024
Closest in time.
Benchmark probing: Investigating data leakage in large language models
Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan · 2024
Closest in time.
Private benchmarking to prevent contamination and improve comparative evaluation of llms, 2024
Nishanth Chandran, Sunayana Sitaram, Divya Gupta, Rahul Sharma, Kashish Mittal, and Manohar Swaminathan · 2024
Closest in time.
Simplesafetytests: a test suite for identifying critical safety risks in large language models, 2024
Bertie Vidgen, Nino Scherrer, Hannah Rose Kirk, Rebecca Qian, Anand Kannappan, Scott A. Hale, and Paul Röttger · 2024
Closest in time.
Acceptable use policies for foundation models, 2024
Kevin Klyman · 2024
Closest in time.
Evaluating the moral beliefs encoded in llms
Nino Scherrer, Claudia Shi, Amir Feder, and David Blei · 2024
Closest in time.
A strongreject for empty jailbreaks
Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, et al · 2024
Closest in time.
Rainbow teaming: Open-ended generation of diverse adversarial prompts, 2024
Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H. Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, Tim Rocktäschel, and Roberta Raileanu · 2024
Closest in time.
Gradient-based language model red teaming, 2024
Nevan Wichers, Carson Denison, and Ahmad Beirami · 2024
Closest in time.
Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions, 2024
Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou · 2024
Closest in time.
Realchat-1m: A large-scale real-world LLM conversation dataset
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang · 2024
Closest in time.
(inthe)wildchat: 570k chatGPT interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng · 2024
Closest in time.
Preventing harm from non-conscious bias in medical generative ai
Janna Hastings · 2024
Closest in time.
Unequal Risk, Unequal Reward: How Gen AI disproportionately harms countries
Barani Maung and Keegan McBride · 2024
Closest in time.
Cognitive bias in high-stakes decision-making with llms
Jessica Echterhoff, Yao Liu, Abeer Alessa, Julian McAuley, and Zexue He · 2024
Closest in time.
Genai against humanity: nefarious applications of generative artificial intelligence and large language models
Emilio Ferrara · 2024
Closest in time.
Abusegpt: Abuse of generative ai chatbots to create smishing campaigns, 2024
Ashfak Md Shibli, Mir Mehedi A. Pritom, and Maanak Gupta · 2024
Closest in time.
The terrifying a.i. scam that uses your loved one’s voice, 2024
The New Yorker · 2024
Closest in time.
Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b, 2024
Pranav Gade, Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladish · 2024
Closest in time.
The ethics of interaction: Mitigating security threats in llms, 2024
Ashutosh Kumar, Sagarika Singh, Shiv Vignesh Murty, and Swathy Ragupathy · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2024
Closest in time.
Two types of ai existential risk: Decisive and accumulative, 2024
Atoosa Kasirzadeh · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
Closest in time.
Evaluating frontier models for dangerous capabilities, 2024
Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Gregoire Deletang, Anian Ruoss, Seliem El-Sayed, Sasha Brown, Anca Dragan, Rohin Shah, Allan Dafoe, and Toby Shevlane · 2024
Closest in time.
Towards ai safety: A taxonomy for ai system evaluation, 2024
Boming Xia, Qinghua Lu, Liming Zhu, and Zhenchang Xing · 2024
Closest in time.
Croissant: A metadata format for ml-ready datasets, 2024
Mubashara Akhtar, Omar Benjelloun, Costanza Conforti, Joan Giner-Miguelez, Nitisha Jain, Michael Kuchnik, Quentin Lhoest, Pierre Marcenac, Manil Maskey, Peter Mattson, Luis Oala, Pierre Ruyssen, Rajat Shinde, Elena Simperl, Goeffry Thomas, Slava Tykhonov, Joaquin Vanschoren, Steffen Vogler, and Carole-Jean Wu · 2024
Closest in time.
NTIA Open Weights Response: Towards A Secure Open Society Powered By Personal AI, 2024
Context Fund Policy Working Group · 2024
Closest in time.
Causally estimating the effect of youtube’s recommender system using counterfactual bots
Homa Hosseinmardi, Amir Ghasemian, Miguel Rivera-Lanas, Manoel Horta Ribeiro, Robert West, and Duncan J. Watts · 2024
Closest in time.
A causal framework for ai regulation and auditing, 2024
Lee Sharkey, Clíodhna Ní Ghuidhir, Dan Braun, Jérémy Scheurer, Mikita Balesni, Lucius Bushnaq, Charlotte Stix, and Marius Hobbhahn · 2024
Closest in time.
Trustllm: Trustworthiness in large language models, 2024
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric Xing, Furong Huang, Hao Liu, Heng Ji, Hongyi Wang, Huan Zhang, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang, Mohit Bansal, James Zou, Jian Pei, Jian Liu, Jianfeng Gao, Jiawei Han, Jieyu Zhao, Jiliang Tang, Jindong Wang, Joaquin Vanschoren, John Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang, Lifang He, Lifu Huang, Michael Backes, Neil Zhenqiang Gong, Philip S. Yu, Pin-Yu Chen, Quanquan Gu, Ran Xu, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen, Tianming Liu, Tianyi Zhou, William Wang, Xiang Li, Xiangliang Zhang, Xiao Wang, Xing Xie, Xun Chen, Xuyu Wang, Yan Liu, Yanfang Ye, Yinzhi Cao, Yong Chen, and Yue Zhao · 2024
Closest in time.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models, 2024
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li · 2024
Closest in time.
Jailbreaking leading safety-aligned llms with simple adaptive attacks, 2024
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Closest in time.
Ignore this title and hackaprompt: Exposing systemic vulnerabilities of llms through a global scale prompt hacking competition, 2024
Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, and Jordan Boyd-Graber · 2024
Closest in time.
How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms, 2024
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi · 2024
Closest in time.
Generative ai red teaming challenge: transparency report, 2024
Victor Storchan, Ravin Kumar, Rumman Chowdhury, Seraphina Goldfarb-Tarrant, and Sven Cattell · 2024
Closest in time.
Metr example task suite, public
Megan Kinniment, Brian Goodrich, Max Hasin, Ryan Bloom, Haoxing Du, Lucas Jun Koba Sato, Daniel Ziegler, Timothee Chauvin, Thomas Broadley, Tao R. Lin, Ted Suzman, Francisco Carvalho, Michael Chen, Niels Warncke, Bart Bussmann, Axel Højmark, Chris MacLeod, and Elizabeth Barnes · 2024
Closest in time.
Activefence safety api, 2024
ActiveFence · 2024
Closest in time.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment, 2024
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li · 2024
Closest in time.
Mapping global dynamics of benchmark creation and saturation in artificial intelligence
Simon Ott, Adriano Barbosa-Silva, Kathrin Blagec, Jan Brauner, and Matthias Samwald · 2041
Closest in time.