Fetching the paper…
Reading the bibliography…
Large language models trained for safety and harmlessness remain susceptible to adversarial misuse, as evidenced by the prevalence of "jailbreak" attacks on early releases of ChatGPT that elicit undesired behavior.
Safety Through Design
W.C. Christensen and F.A. Manuele · 1999
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Adversarial attacks and defences: A survey
Anirban Chakraborty, Manaar Alam, Vishal Dey, Anupam Chattopadhyay, and Debdeep Mukhopadhyay · 2018
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan · 2020
Earlier work this paper cites.
Adversarial attacks on deep-learning models in natural language processing: A survey
Wei Emma Zhang, Quan Z Sheng, Ahoud Alhazmi, and Chenliang Li · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Process for adapting language models to society (PALMS) with values-targeted datasets
Irene Solaiman and Christy Dennison · 2021
Earlier work this paper cites.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Earlier work this paper cites.
On the impossible safety of large AI models
El-Mahdi El-Mhamdi, Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Lê-Nguyên Hoang, Rafael Pinot, Sébastien Rouault, and John Stephan · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Earlier work this paper cites.
All the news that’s fit to fabricate: Ai-generated text as a tool of media misinformation
Sarah Kreps, R. Miles McCain, and Miles Brundage · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Red teaming language models with language models
Ethan Perez, Saffron Huang, H. Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Cited alongside, same era.
On second thought, let’s not think step by step! Bias and toxicity in zero-shot reasoning
Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang · 2022
Cited alongside, same era.
DAN is my new friend
walkerspider · 2022
Cited alongside, same era.
Exploring the limits of domain-adaptive training for detoxifying large-scale language models
Boxin Wang, Wei Ping, Chaowei Xiao, Peng Xu, Mostofa Patwary, Mohammad Shoeybi, Bo Li, Anima Anandkumar, and Bryan Catanzaro · 2022
Cited alongside, same era.
Josh A Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, and Katerina Sedova · 2023
Closest in time.
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Closest in time.
A two sentence jailbreak for GPT-4 and Claude & why nobody knows how to fix it
Alexey Guzey · 2023
Closest in time.
Large language models can be used to effectively scale spear phishing campaigns
Julian Hazell · 2023
Closest in time.
Automatically auditing large language models via discrete optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
“Thread of known ChatGPT jailbreaks”
Zack Witten · 2022
Cited alongside, same era.
Universal LLM jailbreak: ChatGPT, GPT-4, Bard, Bing, Anthropic, and beyond
Adversa · 2023
Cited alongside, same era.
“Another jailbreak for GPT4: Talk to it in Morse code”
Boaz Barak · 2023
Cited alongside, same era.
“Deploying GPT-4 subject to adversarial pressures of real world has been a great practice run for practical AI alignment. Just getting started, but encouraged by degree of alignment we’ve achieved so far (and the engineering process we’ve been maturing to improve issues).”
Greg Brockman · 2023
Cited alongside, same era.
Introducing ChatGPT and Whisper APIs
Greg Brockman, Atty Eleti, Elie Georges, Joanne Jang, Logan Kilpatrick, Rachel Lim, Luke Miller, and Michelle Pokrass · 2023
Cited alongside, same era.
The hacking of ChatGPT is just getting started
Matt Burgess · 2023
Cited alongside, same era.
Erik Jones, Anca Dragan, Aditi Raghunathan, and Jacob Steinhardt · 2023
Closest in time.
Exploiting programmatic behavior of LLMs: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto · 2023
Closest in time.
Pretraining language models with human preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez · 2023
Closest in time.
Multi-step jailbreaking privacy attacks on ChatGPT
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, and Yangqiu Song · 2023
Closest in time.
Analyzing leakage of personally identifiable information in language models
Nils Lukas, Ahmed Salem, Robert Sim, Shruti Tople, Lukas Wutschitz, and Santiago Zanella-Béguelin · 2023
Closest in time.
Jailbreaking ChatGPT on release day
Zvi Mowshowitz · 2023
Closest in time.
Mechanistic interpretability quickstart guide
Neel Nanda · 2023
Closest in time.
“New jailbreak based on virtual functions smuggle”
Nin_kat · 2023
Closest in time.
“The new jailbreak is so fun”
Roman Semenov · 2023
Closest in time.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan · 2023
Closest in time.
You can use GPT-4 to create prompt injections against GPT-4
WitchBOT · 2023
Closest in time.
Fundamental limitations of alignment in large language models
Yotam Wolf, Noam Wies, Yoav Levine, and Amnon Shashua · 2023
Closest in time.