Fetching the paper…
Reading the bibliography…
Recent advancements in AI safety have led to increased efforts in training and red-teaming large language models (LLMs) to mitigate unsafe content generation.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 1908
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Earlier work this paper cites.
A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily, 2023
Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang · 2023
Earlier work this paper cites.
The trojan detection challenge 2023 (llm edition)
Center for AI Safety · 2023
Earlier work this paper cites.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks, 2023
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto · 2023
Earlier work this paper cites.
Low-resource languages jailbreak GPT-4
Zheng Xin Yong, Cristina Menghini, and Stephen Bach · 2023
Cited alongside, same era.
Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity, 2023
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models, 2023
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Large language models for mathematical reasoning: Progresses and challenges
Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin · 2024
Cited alongside, same era.
Meta-llama-3-8b-instruct
Meta AI · 2024
Cited alongside, same era.
Anthropic models
Anthropic · 2024
Cited alongside, same era.
Large language models for mathematicians, 2024
Simon Frieder, Julius Berner, Philipp Petersen, and Thomas Lukasiewicz · 2024
Closest in time.
Gemini models
Google · 2024
Closest in time.
Artprompt: Ascii art-based jailbreak attacks against aligned llms, 2024
Fengqing Jiang, Zhangchen Xu, Luyao Niu, Zhen Xiang, Bhaskar Ramasubramanian, Bo Li, and Radha Poovendran · 2024
Closest in time.
Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction
Tong Liu, Zhe Zhao, Yinpeng Dong, Guozhu Meng, and Kai Chen · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Masterkey: Automated jailbreaking of large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu · 2024
Cited alongside, same era.
Huggingface sentence-transformers all-minilm-l6-v2
Hugging Face · 2024
Cited alongside, same era.
Large language models are neurosymbolic reasoners
Meng Fang, Shilong Deng, Yudi Zhang, Zijing Shi, Ling Chen, Mykola Pechenizkiy, and Jun Wang · 2024
Cited alongside, same era.
Multilingual jailbreak challenges in large language models
Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing
Cited in the paper.
Openai models
OpenAI
Cited in the paper.
OpenAI
Cited in the paper.
Closest in time.
Gpt-4o-2024-05-13
OpenAI · 2024
Closest in time.
Tricking llms into disobedience: Formalizing, analyzing, and detecting jailbreaks
Abhinav Sukumar Rao, Atharva Roshan Naik, Sachin Vashistha, Somak Aditya, and Monojit Choudhury · 2024
Closest in time.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts, 2024
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing · 2024
Closest in time.