Fetching the paper…
Reading the bibliography…
Large Language Diffusion Models (LLDMs) exhibit comparable performance to LLMs while offering distinct advantages in inference speed and mathematical reasoning tasks.The precise and rapid generation capabilities of LLDMs amplify concerns of harmful generations, while existing jailbreak methodologies designed for Large Language Models (LLMs) prove limited effectiveness against LLDMs and fail to expose safety vulnerabilities.Successful defense cannot definitively resolve harmful generation concerns, as it remains unclear whether LLDMs possess safety robustness or existing attacks are incompatible with diffusion-based architectures.To address this, we first reveal the vulnerability of LLDMs to jailbreak and demonstrate that attack failure in LLDMs stems from fundamental architectural differences.We present a PArallel Decoding jailbreak (PAD) for diffusion-based language models.
Denoising Diffusion Probabilistic Models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2006
Earlier work this paper cites.
A definition of cascading disasters and cascading effects: Going beyond the “toppling dominos” metaphor
Pescaroli, G.; and Alexander, D. 2015 · 2015
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; Joseph, N.; Kadavath, S.; Kernion, J.; Conerly, T.; El-Showk, S.; Elhage, N.; Hatfield-Dodds, Z.; Hernandez, D.; Hume, T.; Johnston, S.; Kravec, S.; Lovitt, L.; Nanda, N.; Olsson, C.; Amodei, D.; Brown, T.; Clark, J.; McCandlish, S.; Olah, C.; Mann, B.; and Kaplan, J. 2022 · 2022
Earlier work this paper cites.
DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models
He, Z.; Sun, T.; Wang, K.; Huang, X.; and Qiu, X. 2022 · 2022
Earlier work this paper cites.
Diffusion-LM Improves Controllable Text Generation
Li, X. L.; Thickstun, J.; Gulrajani, I.; Liang, P.; and Hashimoto, T. B. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022 · 2022
Earlier work this paper cites.
The carbon footprint of machine learning training will plateau, then shrink
Patterson, D.; Gonzalez, J.; Hölzle, U.; Le, Q.; Liang, C.; Munguia, L.-M.; Rothchild, D.; So, D. R.; Texier, M.; and Dean, J. 2022 · 2022
Earlier work this paper cites.
Structured Denoising Diffusion Models in Discrete State-Spaces
Austin, J.; Johnson, D. D.; Ho, J.; Tarlow, D.; and van den Berg, R. 2023 · 2023
Earlier work this paper cites.
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Inan, H.; Upasani, K.; Chi, J.; Rungta, R.; Iyer, K.; Mao, Y.; Tontchev, M.; Hu, Q.; Fuller, B.; Testuggine, D.; and Khabsa, M. 2023 · 2023
Earlier work this paper cites.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; and Chaplot, D. S. 2023 · 2023
Earlier work this paper cites.
Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
Sun, Z.; Shen, Y.; Zhou, Q.; Zhang, H.; Chen, Z.; Cox, D.; Yang, Y.; and Gan, C. 2023 · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; Millican, K.; et al. 2023 · 2023
Earlier work this paper cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; and Almahairi, A. 2023 · 2023
Earlier work this paper cites.
Many-shot Jailbreaking
Anil, C.; Durmus, E.; Panickssery, N.; and Sharma, M. e. a. 2024 · 2024
Earlier work this paper cites.
Bianchi, F.; Suzgun, M.; Attanasio, G.; Röttger, P.; Jurafsky, D.; Hashimoto, T.; and Zou, J. 2024 · 2024
Cited alongside, same era.
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
Cao, B.; Cao, Y.; Lin, L.; and Chen, J. 2024 · 2024
Cited alongside, same era.
Jailbreaking Black Box Large Language Models in Twenty Queries
Chao, P.; Robey, A.; Dobriban, E.; Hassani, H.; Pappas, G. J.; and Wong, E. 2024 · 2024
Cited alongside, same era.
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
DeepSeek-AI; Bi, X.; Chen, D.; Chen, G.; and Chen, S. 2024 · 2024
Cited alongside, same era.
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Yi, S.; Liu, Y.; Sun, Z.; Cong, T.; He, X.; Song, J.; Xu, K.; and Li, Q. 2024 · 2024
Later among the works it cites.
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
Yuan, Y.; Jiao, W.; Wang, W.; tse Huang, J.; He, P.; Shi, S.; and Tu, Z. 2024 · 2024
Later among the works it cites.
Zeng, Y.; Lin, H.; Zhang, J.; Yang, D.; Jia, R.; and Shi, W. 2024 · 2024
Later among the works it cites.
On Prompt-Driven Safeguarding for Large Language Models
Zheng, C.; Yin, F.; Zhou, H.; Meng, F.; Zhou, J.; Chang, K.-W.; Huang, M.; and Peng, N. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grattafiori, A.; Dubey, A.; Jauhri, A.; Pandey, A.; and Kadian, A. 2024 · 2024
Cited alongside, same era.
Gu, J.; Jiang, X.; Shi, Z.; Tan, H.; Zhai, X.; Xu, C.; Li, W.; Shen, Y.; Ma, S.; Liu, H.; et al. 2024 · 2024
Cited alongside, same era.
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
Guo, X.; Yu, F.; Zhang, H.; Qin, L.; and Hu, B. 2024 · 2024
Cited alongside, same era.
Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
Jia, X.; Pang, T.; Du, C.; Huang, Y.; Gu, J.; Liu, Y.; Cao, X.; and Lin, M. 2024 · 2024
Cited alongside, same era.
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Liu, X.; Xu, N.; Chen, M.; and Xiao, C. 2024 · 2024
Cited alongside, same era.
Llama Team, A. . M. 2024 · 2024
Cited alongside, same era.
OpenAI; Achiam, J.; Adler, S.; Agarwal, S.; and Ahmad, L. 2024 · 2024
Cited alongside, same era.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2024 · 2024
Cited alongside, same era.
Zhou, Z.; Xiang, J.; Chen, H.; Liu, Q.; Li, Z.; and Su, S. 2024 · 2024
Later among the works it cites.
Jailbreaking black box large language models in twenty queries
Chao, P.; Robey, A.; Dobriban, E.; Hassani, H.; Pappas, G. J.; and Wong, E. 2025 · 2025
Closest in time.
Gemini Diffusion
Google DeepMind. 2025 · 2025
Closest in time.
Chatbug: A common vulnerability of aligned llms induced by chat templates
Jiang, F.; Xu, Z.; Niu, L.; Lin, B. Y.; and Poovendran, R. 2025 · 2025
Closest in time.
Jin, H.; Chen, R.; Zhang, P.; Zhou, A.; Zhang, Y.; and Wang, H. 2025 · 2025
Closest in time.
dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
Liu, Z.; Yang, Y.; Zhang, Y.; Chen, J.; Zou, C.; Wei, Q.; Wang, S.; and Zhang, L. 2025 · 2025
Closest in time.
Qwen; Yang, A.; Yang, B.; Zhang, B.; and Hui, B. 2025 · 2025
Closest in time.
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
Wang, K.; Zhang, G.; Zhou, Z.; and et al., J. W. 2025 · 2025
Closest in time.
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
Wu, C.; Zhang, H.; Xue, S.; Liu, Z.; Diao, S.; Zhu, L.; Luo, P.; Han, S.; and Xie, E. 2025 · 2025
Closest in time.
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
You, Z.; Nie, S.; Zhang, X.; Hu, J.; Zhou, J.; Lu, Z.; Wen, J.-R.; and Li, C. 2025 · 2025
Closest in time.
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
Zhu, F.; Wang, R.; Nie, S.; Zhang, X.; Wu, C.; Hu, J.; Zhou, J.; Chen, J.; Lin, Y.; Wen, J.-R.; and Li, C. 2025 · 2025
Closest in time.