Fetching the paper…
Reading the bibliography…
Many publicly available language models have been safety tuned to reduce the likelihood of toxic or liability-inducing text.
“AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts”, 2020
Taylor Shin et al · 2010
Earlier work this paper cites.
“Explaining and harnessing adversarial examples”
Ian Goodfellow, Jonathon Shlens and Christian Szegedy · 2014
Earlier work this paper cites.
“Fitnets: Hints for thin deep nets”
Adriana Romero et al · 2014
Earlier work this paper cites.
“Distilling the Knowledge in a Neural Network”, 2015
Geoffrey Hinton, Oriol Vinyals and Jeff Dean · 2015
Earlier work this paper cites.
“Deep learning with differential privacy”
Martin Abadi et al · 2016
Earlier work this paper cites.
“Towards evaluating the robustness of neural networks”
Nicholas Carlini and David Wagner · 2017
Earlier work this paper cites.
“Towards deep learning models resistant to adversarial attacks”
Aleksander Madry et al · 2017
Earlier work this paper cites.
“Adversarial examples are not bugs, they are features”
Andrew Ilyas et al · 2019
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Earlier work this paper cites.
“Red teaming language models with language models”
Ethan Perez et al · 2022
Earlier work this paper cites.
Weijia Shi et al · 2022
Earlier work this paper cites.
“Detecting language model attacks with perplexity”
Gabriel Alon and Michael Kamfonas · 2023
Earlier work this paper cites.
“Many-shot Jailbreaking”, 2023
Cem Anil et al · 2023
Earlier work this paper cites.
“Jailbreaking black box large language models in twenty queries”
Patrick Chao et al · 2023
Earlier work this paper cites.
“Prompting4debugging: Red-teaming text-to-image diffusion models by finding problematic prompts”
Zhi-Yi Chin et al · 2023
Earlier work this paper cites.
“Scaling laws for adversarial attacks on language model activations”
Stanislav Fort · 2023
Earlier work this paper cites.
“Codeattack: Code-based adversarial attacks for pre-trained programming language models”
Akshita Jha and Chandan Reddy · 2023
Earlier work this paper cites.
“Automatically Auditing Large Language Models via Discrete Optimization”, 2023
Erik Jones, Anca Dragan, Aditi Raghunathan and Jacob Steinhardt · 2023
Cited alongside, same era.
“Exploiting programmatic behavior of llms: Dual-use through standard security attacks”
Daniel Kang et al · 2023
Cited alongside, same era.
“Open sesame! universal black box jailbreaking of large language models”
Raz Lapid, Ron Langberg and Moshe Sipper · 2023
Cited alongside, same era.
“LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B (arXiv: 2310.20624). arXiv”, 2023
S Lermen, C Rogers-Smith and J Ladish · 2023
Cited alongside, same era.
“Autodan: Generating stealthy jailbreak prompts on aligned large language models”
“Cold-attack: Jailbreaking llms with stealthiness and controllability”
Xingang Guo et al · 2024
Closest in time.
“llama3-jailbreak”, 2024
Haize · 2024
Closest in time.
“Query-Based Adversarial Prompt Generation”
Jonathan Hayase et al · 2024
Closest in time.
“ChatGPT DAN, Jailbreaks prompt”, 2024
Kiho Lee · 2024
Closest in time.
Zeyi Liao and Huan Sun · 2024
Closest in time.
“The Trojan Detection Challenge 2023 (LLM Edition)”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiaogeng Liu, Nan Xu, Muhao Chen and Chaowei Xiao · 2023
Cited alongside, same era.
“Eliciting Language Model Behaviors using Reverse Language Models”
Jacob Pfau et al · 2023
Cited alongside, same era.
“Scalable and transferable black-box jailbreaks for language models via persona modulation”
Rusheb Shah et al · 2023
Cited alongside, same era.
“Llama 2: Open foundation and fine-tuned chat models”
Hugo Touvron et al · 2023
Cited alongside, same era.
“NeurIPS 2023 Trojan Detection Competition” Accessed: 2024-07-09, https://trojandetection.ai/ , 2023
Trojan Detection Competition · 2023
Cited alongside, same era.
“Low-resource languages jailbreak gpt-4”
Zheng-Xin Yong, Cristina Menghini and Stephen Bach · 2023
Cited alongside, same era.
“Removing rlhf protections in gpt-4 via fine-tuning”
Qiusi Zhan et al · 2023
Cited alongside, same era.
“AutoDAN: Automatic and Interpretable Adversarial Attacks on Large Language Models”
Sicheng Zhu et al · 2023
Cited alongside, same era.
Mantas Mazeika et al · 2024
Closest in time.
“Improving Zero-shot Generalization of Learned Prompts via Unsupervised Knowledge Distillation”
Marco Mistretta, Alberto Baldrati, Marco Bertini and Andrew Bagdanov · 2024
Closest in time.
“Advprompter: Fast adaptive adversarial prompting for llms”
Anselm Paulus et al · 2024
Closest in time.
“Direct preference optimization: Your language model is secretly a reward model”
Rafael Rafailov et al · 2024
Closest in time.
“Fast Adversarial Attacks on Language Models In One GPU Minute”
Vinu Sadasivan et al · 2024
Closest in time.
“All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks”
Kazuhiro Takemoto · 2024
Closest in time.
“Fluent dreaming for language models”
T Thompson, Zygimantas Straznickas and Michael Sklar · 2024
Closest in time.
“Breaking Circuit Breakers”, 2024
T. Thompson and Michael Sklar · 2024
Closest in time.
Hao Wang, Hao Li, Minlie Huang and Lei Sha · 2024
Closest in time.
“Jailbroken: How does llm safety training fail?”
Alexander Wei, Nika Haghtalab and Jacob Steinhardt · 2024
Closest in time.
“Survey on Knowledge Distillation for Large Language Models: Methods, Evaluation, and Application”
Chuanpeng Yang et al · 2024
Closest in time.
“Improving Alignment and Robustness with Short Circuiting”
Andy Zou et al · 2024
Closest in time.