Fetching the paper…
Reading the bibliography…
Despite efforts to align large language models to produce harmless responses, they are still vulnerable to jailbreak prompts that elicit unrestricted behaviour.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh · 2020
Earlier work this paper cites.
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Rlprompt: Optimizing discrete text prompts with reinforcement learning
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric P Xing, and Zhiting Hu · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Red Teaming Language Models with Language Models, February 2022
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Earlier work this paper cites.
Model card and evaluations for claude models, July 2023
Anthropic · 2023
Earlier work this paper cites.
Are aligned neural networks adversarially aligned?, June 2023
Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, and Ludwig Schmidt · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Toxicity in ChatGPT: Analyzing Persona-assigned Language Models, April 2023
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan · 2023
Cited alongside, same era.
Personas as a way to model truthfulness in language models, 2023
Nitish Joshi, Javier Rando, Abulhair Saparov, Najoung Kim, and He He · 2023
Closest in time.
ChatGPT "DAN" (and other "Jailbreaks"), August 2023
Lee Kiho · 2023
Closest in time.
Generative Agents: Interactive Simulacra of Human Behavior, August 2023
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein · 2023
Closest in time.
Role-play with large language models
Murray Shanahan, Kyle McDonell, and Laria Reynolds · 2023
Closest in time.
Jailbroken: How Does LLM Safety Training Fail?, July 2023
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
PICT: A Zero-Shot Prompt Template to Automate Evaluation, 2023
Quentin Feuillade-Montixi · 2023
Cited alongside, same era.
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Cited alongside, same era.
Automatically auditing large language models via discrete optimization
Erik Jones, Anca Dragan, Aditi Raghunathan, and Jacob Steinhardt · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al
Cited in the paper.
Explore, Establish, Exploit: Red Teaming Language Models from Scratch, June 2023b
Stephen Casper, Jason Lin, Joe Kwon, Gatlen Culp, and Dylan Hadfield-Menell
Cited in the paper.
Autodan: Generating stealthy jailbreak prompts on aligned large language models, 2023a
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao
Cited in the paper.
Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study, May 2023b
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu
Cited in the paper.
Usage policies, 2023a
OpenAI
Cited in the paper.
Yotam Wolf, Noam Wies, Oshri Avnery, Yoav Levine, and Amnon Shashua · 2023
Closest in time.
Universal and Transferable Adversarial Attacks on Aligned Language Models, July 2023
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Closest in time.