Fetching the paper…
Reading the bibliography…
Jailbreak attacks aim to bypass the LLMs' safeguards.
Heuristic “optimization”: Why, when, and how to use it
Stelios H Zanakis and James R Evans · 1981
Earlier work this paper cites.
Heuristics: intelligent search strategies for computer problem solving
PictureJudea Pearl · 1984
Earlier work this paper cites.
Feedback control techniques for gradient based learning
Yongong Tan, Xuanju Dang, and Chun-Yi-Su · 2000
Earlier work this paper cites.
Bayesian Learning via Stochastic Gradient Langevin Dynamics
Max Welling and Yee Whye Teh · 2011
Earlier work this paper cites.
Content Analysis: An Introduction to Its Methodology
Klaus Krippendorff · 2018
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits · 2020
Earlier work this paper cites.
A General Language Assistant as a Laboratory for Alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan · 2021
Earlier work this paper cites.
Extracting Training Data from Large Language Models
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel · 2021
Earlier work this paper cites.
BadNL: Backdoor Attacks Against NLP Models with Semantic-preserving Improvements
Xiaoyi Chen, Ahmed Salem, Michael Backes, Shiqing Ma, Qingni Shen, Zhonghai Wu, and Yang Zhang · 2021
Earlier work this paper cites.
WebGPT: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman · 2021
Earlier work this paper cites.
Detecting AI Trojans Using Meta Neural Analysis
Xiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov, Carl A. Gunter, and Bo Li · 2021
Earlier work this paper cites.
Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures
Eugene Bagdasaryan and Vitaly Shmatikov · 2022
Earlier work this paper cites.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Bad Characters: Imperceptible NLP Attacks
Nicholas Boucher, Ilia Shumailov, Ross Anderson, and Nicolas Papernot · 2022
Earlier work this paper cites.
A Holistic Approach to Undesired Content Detection in the Real World
Todor Markov, Chong Zhang, Sandhini Agarwal, Tyna Eloundou, Teddy Lee, Steven Adler, Angela Jiang, and Lilian Weng · 2022
Earlier work this paper cites.
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
DeepPhish: Understanding User Trust Towards Artificially Generated Profiles in Online Social Networks
Jaron Mink, Licheng Luo, Natã M. Barbosa, Olivia Figueira, Yang Wang, and Gang Wang · 2022
Earlier work this paper cites.
Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks
Fatemehsadat Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick, and Reza Shokri · 2022
Earlier work this paper cites.
https://chat.openai.com/chat , 2022
OpenAI · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Red Teaming Language Models with Language Models
Ethan Perez, Saffron Huang, H. Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Earlier work this paper cites.
Ignore Previous Prompt: Attack Techniques For Language Models
Fábio Perez and Ian Ribeiro · 2022
Earlier work this paper cites.
Truth Serum: Poisoning Machine Learning Models to Reveal Their Secrets
Florian Tramèr, Reza Shokri, Ayrton San Joaquin, Hoang Le, Matthew Jagielski, Sanghyun Hong, and Nicholas Carlini · 2022
Earlier work this paper cites.
Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Sahar Abdelnabi, Kai Greshake, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Earlier work this paper cites.
Detecting Language Model Attacks with Perplexity
Gabriel Alon and Michael Kamfonas · 2023
Earlier work this paper cites.
http://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm , 2023
CAC · 2023
Earlier work this paper cites.
Jailbreaking Black Box Large Language Models in Twenty Queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong · 2023
Earlier work this paper cites.
Jailbreaker: Automated Jailbreak Across Multiple Large Language Model Chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu · 2023
Earlier work this paper cites.
Multilingual Jailbreak Challenges in Large Language Models
Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing · 2023
Earlier work this paper cites.
A Pro-Innovation Approach to AI Regulation
DSIT · 2023
Earlier work this paper cites.
https://ai.google/discover/palm2/ , 2023
Google · 2023
Cited alongside, same era.
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Cited alongside, same era.
Large Language Models Can Be Used To Effectively Scale Spear Phishing Campaigns
Julian Hazell · 2023
Cited alongside, same era.
MGTBench: Benchmarking Machine-Generated Text Detection
Xinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang · 2023
Cited alongside, same era.
Synthetic Lies: Understanding AI-Generated Misinformation and Evaluating Algorithmic and Human Solutions
Jiawei Zhou, Yixuan Zhang, Qianni Luo, Andrea G. Parker, and Munmun De Choudhury · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
https://artificialintelligenceact.eu/ , 2024
EU AI Act · 2024
Closest in time.
https://aws.amazon.com/cn/machine-learning/responsible-ai/policy/ , 2024
Amazon · 2024
Closest in time.
https://aws.amazon.com/cn/aup/ , 2024
Amazon · 2024
Closest in time.
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiaowei Huang, Wenjie Ruan, Wei Huang, Gaojie Jin, Yi Dong, Changshun Wu, Saddek Bensalem, Ronghui Mu, Yi Qi, Xingyu Zhao, Kaiwen Cai, Yanghao Zhang, Sihao Wu, Peipei Xu, Dengyu Wu, Andre Freitas, and Mustafa A. Mustafa · 2023
Cited alongside, same era.
Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen · 2023
Cited alongside, same era.
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa · 2023
Cited alongside, same era.
Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto · 2023
Cited alongside, same era.
Certifying LLM Safety against Adversarial Prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Aaron Jiaxun Li, Soheil Feizi, and Himabindu Lakkaraju · 2023
Cited alongside, same era.
Multi-step Jailbreaking Privacy Attacks on ChatGPT
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, and Yangqiu Song · 2023
Cited alongside, same era.
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2023
Cited alongside, same era.
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Closest in time.
https://claude.ai/ , 2024
Anthropic · 2024
Closest in time.
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
Junjie Chu, Zeyang Sha, Michael Backes, and Yang Zhang · 2024
Closest in time.
DeepSeek-AI · 2024
Closest in time.
h4rm3l: A Dynamic Benchmark of Composable Jailbreak Attacks for LLM Safety Assessment
Moussa Koulako Bala Doumbouya, Ananjan Nandi, Gabriel Poesia, Davide Ghilardi, Anna Goldie, Federico Bianchi, Dan Jurafsky, and Christopher D. Manning · 2024
Closest in time.
https://policies.google.com/terms/generative-ai/use-policy?hl=en , 2024
Google · 2024
Closest in time.
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
Xingang Guo, Fangxu Yu, Huan Zhang, Lianhui Qin, and Bin Hu · 2024
Closest in time.
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
Xirui Li, Ruochen Wang, Minhao Cheng, Tianyi Zhou, and Cho-Jui Hsieh · 2024
Closest in time.
https://ai.meta.com/llama/use-policy/ , 2024
Meta · 2024
Closest in time.
Llama 3.1
Meta · 2024
Closest in time.
Llama Guard 2
Meta · 2024
Closest in time.
Llama Guard 3
Meta · 2024
Closest in time.
Prompt Guard
Meta · 2024
Closest in time.
https://learn.microsoft.com/en-us/legal/cognitive-services/openai/code-of-conduct , 2024
Microsoft · 2024
Closest in time.
https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/harm-categories?tabs=warning , 2024
Microsoft · 2024
Closest in time.
https://openai.com/policies/usage-policies , 2024
OpenAI · 2024
Closest in time.
OpenAI’s commitment to child safety: adopting safety by design principles
OpenAI · 2024
Closest in time.
AI Bill of Rights
OSTP · 2024
Closest in time.
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian · 2024
Closest in time.
LLM Jailbreak Attack versus Defense Techniques – A Comprehensive Study
Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Picek · 2024
Closest in time.
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Sibo Yi, Yule Liu, Zhen Sun, Tianshuo Cong, Xinlei He, Jiaxing Song, Ke Xu, and Qi Li · 2024
Closest in time.
https://status.deepseek.com , 2025
DeepSeek · 2025
Closest in time.
https://platform.deepseek.com , 2025
DeepSeek · 2025
Closest in time.
EU centre to prevent and combat child sexual abuse
EU · 2025
Closest in time.
https://github.com/ThuCCSLab/Awesome-LM-SSP/ , 2025
ThuCCSLab · 2025
Closest in time.
https://https://github.com/ydyjya/Awesome-LLM-Safety/ , 2025
Zhenhong Zhou · 2025
Closest in time.