Fetching the paper…
Reading the bibliography…
Foundation models (FMs) provide societal benefits but also amplify risks.
The claude 3 model family: Opus, sonnet, haiku, 2004
Anthropic · 2004
Earlier work this paper cites.
Regulation (EU) 2016/679 of the European Parliament and of the Council
European Parliament and Council of the European Union · 2016
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Earlier work this paper cites.
Provisions on the management of algorithmic recommendations in internet information services
Cyberspace Administration of China · 2021
Earlier work this paper cites.
Provisions on the administration of deep synthesis internet information services
Cyberspace Administration of China · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Earlier work this paper cites.
Introducing ChatGPT
OpenAI · 2022
Earlier work this paper cites.
Taxonomy of risks posed by language models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Frontier ai regulation: Managing emerging risks to public safety
Markus Anderljung, Joslyn Barnhart, Jade Leung, Anton Korinek, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, et al · 2023
Earlier work this paper cites.
Anthropic acceptable use policy
Anthropic · 2023
Earlier work this paper cites.
Introducing Claude
Anthropic · 2023
Earlier work this paper cites.
Baidu ernie user agreement
Baidu · 2023
Earlier work this paper cites.
Managing ai risks in an era of rapid progress
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, et al · 2023
Earlier work this paper cites.
Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence
Joseph Biden · 2023
Earlier work this paper cites.
Interim measures for the management of generative artificial intelligence services
Cyberspace Administration of China · 2023
Earlier work this paper cites.
Deepseek license agreement
DeepSeek · 2023
Earlier work this paper cites.
Deepseek user agreement
DeepSeek · 2023
Earlier work this paper cites.
Multilingual jailbreak challenges in large language models
Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing · 2023
Earlier work this paper cites.
Generative ai in health care and liability risks for physicians and safety concerns for patients
Mindy Duffourc and Sara Gerke · 2023
Earlier work this paper cites.
Gemini: A family of highly capable multimodal models
Gemini Team · 2023
Earlier work this paper cites.
Google generative ai prohibited use policy
Google · 2023
Earlier work this paper cites.
An overview of catastrophic ai risks, 2023
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside · 2023
Earlier work this paper cites.
An empirical study of metrics to measure representational harms in pre-trained language models
Saghar Hosseini, Hamid Palangi, and Ahmed Hassan Awadallah · 2023
Earlier work this paper cites.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang · 2023
Cited alongside, same era.
Meta llama-2’s acceptable use policy
Meta · 2023
Cited alongside, same era.
Scientific and technological ethics review regulation (trial)
Ministry of Science and Technology of Cina · 2023
Cited alongside, same era.
GPT-4V(ision) system card
OpenAI · 2023
Cited alongside, same era.
New models and developer products announced at devday, 2023
OpenAI · 2023
Cited alongside, same era.
Openai usage policies (pre-jan 10, 2024)
OpenAI · 2023
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
Gemini Team · 2024
Closest in time.
Command r: Retrieval-augmented generation at production scale
Aidan Gomez · 2024
Closest in time.
Introducing command r+: A scalable llm built for business
Aidan Gomez · 2024
Closest in time.
Google gemma prohibited use policy
Google · 2024
Closest in time.
Acceptable use policies for foundation models: Considerations for policymakers and developers
Kevin Klyman · 2024
Closest in time.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy · 2023
Cited alongside, same era.
La plateforme
Mistral AI Team · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
The dangers of generative artificial intelligence, 2023
Luke Tredinnick and Claire Laybats · 2023
Cited alongside, same era.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al · 2023
Cited alongside, same era.
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al · 2024
Closest in time.
Introducing meta llama 3
Meta · 2024
Closest in time.
Meta ais terms of service
Meta · 2024
Closest in time.
Mistral’s legal terms and conditions
Mistral · 2024
Closest in time.
Hello gpt-4o, 2024
OpenAI · 2024
Closest in time.
Openai model spec
OpenAI · 2024
Closest in time.
Openai usage policies
OpenAI · 2024
Closest in time.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2024
Closest in time.
Rainbow teaming: Open-ended generation of diverse adversarial prompts, 2024
Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H. Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, Tim Rocktäschel, and Roberta Raileanu · 2024
Closest in time.
“do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models, 2024
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2024
Closest in time.
Stability’s acceptable use policy
Stability · 2024
Closest in time.
Cheaper, better, faster, stronger
Mistral AI Team · 2024
Closest in time.
Introducing dbrx: A new state-of-the-art open llm
Mosaic Research Team · 2024
Closest in time.
Introducing qwen1.5
Qwen Team · 2024
Closest in time.
Sorry-bench: Systematically evaluating large language model safety refusal behaviors
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, et al · 2024
Closest in time.
Ai risk categorization decoded (air 2024): From government regulations to corporate policies
Yi Zeng, Kevin Klyman, Andy Zhou, Yu Yang, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, and Bo Li · 2024
Closest in time.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi · 2024
Closest in time.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2024
Closest in time.