Fetching the paper…
Reading the bibliography…
Safety aligned Large Language Models (LLMs) are vulnerable to harmful fine-tuning attacks -- a few harmful data mixed in the fine-tuning dataset can break the LLMs's safety alignment.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Bianchi, F., Suzgun, M., Attanasio, G., Röttger, P., Jurafsky, D., Hashimoto, T., and Zou, J · 2023
Earlier work this paper cites.
Safe rlhf: Safe reinforcement learning from human feedback
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y · 2023
Earlier work this paper cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T · 2023
Earlier work this paper cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Earlier work this paper cites.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Sun, R., Wang, Y., and Yang, Y · 2023
Earlier work this paper cites.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b
Lermen, S., Rogers-Smith, C., and Ladish, J · 2023
Earlier work this paper cites.
Fine-tuning can cripple your foundation model; preserving features may be the solution
Mukhoti, J., Gal, Y., Torr, P. H., and Dokania, P. K · 2023
Earlier work this paper cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Earlier work this paper cites.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z · 2023
Earlier work this paper cites.
Alpaca: A strong, replicable instruction-following model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Earlier work this paper cites.
Pairwise proximal policy optimization: Harnessing relative feedback for llm alignment
Wu, T., Zhu, B., Zhang, R., Wen, Z., Ramchandran, K., and Jiao, J · 2023
Earlier work this paper cites.
Shadow alignment: The ease of subverting safely-aligned language models
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Earlier work this paper cites.
Selfee: Iterative self-revising llm empowered by self-feedback generation
Ye, S., Jo, Y., Kim, D., Kim, S., Hwang, H., and Seo, M · 2023
Earlier work this paper cites.
Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity
Yin, L., Wu, Y., Zhang, Z., Hsieh, C.-Y., Wang, Y., Jia, Y., Pechenizkiy, M., Liang, Y., Wang, Z., and Liu, S · 2023
Earlier work this paper cites.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Earlier work this paper cites.
Removing rlhf protections in gpt-4 via fine-tuning
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D · 2023
Earlier work this paper cites.
Bhardwaj, R., Anh, D. D., and Poria, S · 2024
Earlier work this paper cites.
Defending against unforeseen failure modes with latent adversarial training
Casper, S., Schulze, L., Patel, O., and Hadfield-Menell, D · 2024
Earlier work this paper cites.
Safety-aware fine-tuning of large language models
Choi, H. K., Du, X., and Li, Y · 2024
Cited alongside, same era.
Towards secure tuning: Mitigating security risks arising from benign instruction fine-tuning
Du, Y., Zhao, S., Cao, J., Ma, M., Zhao, D., Fan, F., Liu, T., and Qin, B · 2024
Cited alongside, same era.
Mimicking user data: On mitigating fine-tuning risks in closed large language models
Eiras, F., Petrov, A., Torr, P. H., Kumar, M. P., and Bibi, A · 2024
Cited alongside, same era.
The vllm safety paradox: Dual ease in jailbreak attack and defense
Guo, Y., Jiao, F., Nie, L., and Kankanhalli, M · 2024
Cited alongside, same era.
Covert malicious finetuning: Challenges in safeguarding llm adaptation
On the vulnerability of safety alignment in open-access llms
Yi, J., Ye, R., Chen, Q., Zhu, B., Chen, S., Lian, D., Sun, G., Xie, X., and Wu, F · 2024
Closest in time.
Locking down the finetuned llms safety
Zhu, M., Yang, L., Wei, Y., Zhang, N., and Zhang, Y · 2024
Closest in time.
Safety fine-tuning at (almost) no cost: A baseline for vision large language models
Zong, Y., Bohdal, O., Yu, T., Yang, Y., and Hospedales, T · 2024
Closest in time.
Fight fire with fire: Defending against malicious rl fine-tuning via reward neutralization
Cao, W · 2025
Closest in time.
Model tampering attacks enable more rigorous evaluations of llm capabilities
Che, Z., Casper, S., Kirk, R., Satheesh, A., Slocum, S., McKinney, L. E., Gandikota, R., Ewart, A., Rosati, D., Wu, Z., et al · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Halawi, D., Wei, A., Wallace, E., Wang, T. T., Haghtalab, N., and Steinhardt, J · 2024
Cited alongside, same era.
What’s in your" safe" data?: Identifying benign data that breaks safety
He, L., Xia, M., and Henderson, P · 2024
Cited alongside, same era.
Safe lora: the silver lining of reducing safety risks when fine-tuning large language models
Hsu, C.-Y., Tsai, Y.-L., Lin, C.-H., Chen, P.-Y., Yu, C.-M., and Huang, C.-Y · 2024
Cited alongside, same era.
What makes and breaks safety fine-tuning? mechanistic study
Jain, S., Lubana, E. S., Oksuz, K., Joy, T., Torr, P. H., Sanyal, A., and Dokania, P. K · 2024
Cited alongside, same era.
No two devils alike: Unveiling distinct mechanisms of fine-tuning attacks
Leong, C. T., Cheng, Y., Xu, K., Wang, J., Wang, H., and Li, W · 2024
Cited alongside, same era.
Keeping llms aligned after fine-tuning: The crucial role of prompt templates
Lyu, K., Zhao, H., Gu, X., Yu, D., Goyal, A., and Arora, S · 2024
Cited alongside, same era.
Navigating the safety landscape: Measuring risks in finetuning large language models
Peng, S., Chen, P.-Y., Hull, M., and Chau, D. H · 2024
Cited alongside, same era.
Towards understanding the fragility of multilingual llms against fine-tuning attacks
Poppi, S., Yong, Z.-X., He, Y., Chern, B., Zhao, H., Yang, A., and Chi, J · 2024
Cited alongside, same era.
Closest in time.
Fundamental safety-capability trade-offs in fine-tuning large language models
Chen, P.-Y., Shen, H., Das, P., and Chen, T · 2025
Closest in time.
On weaponization-resistant large language models with prospect theoretic alignment
Cheng, Z., Zhang, M., Sun, J., and Dai, W · 2025
Closest in time.
Fundamental limitations in defending llm finetuning apis
Davies, X., Winsor, E., Korbak, T., Souly, A., Kirk, R., de Witt, C. S., and Gal, Y · 2025
Closest in time.
Djuhera, A., Kadhe, S. R., Ahmed, F., Zawad, S., and Boche, H · 2025
Closest in time.
Fan, C., Jia, J., Zhang, Y., Ramakrishna, A., Hong, M., and Liu, S · 2025
Closest in time.
Safety misalignment against large language models
Gong, Y., Ran, D., He, X., Cong, T., Wang, A., and Wang, X · 2025
Closest in time.
Benign samples matter! fine-tuning on outlier benign samples severely breaks safety
Guan, Z., Hu, M., Zhu, R., Li, S., and Vullikanti, A · 2025
Closest in time.
Your task may vary: A systematic understanding of alignment and safety degradation when fine-tuning LLMs, 2025
Hsiung, L., Pang, T., Tang, Y.-C., Song, L., Ho, T.-Y., Chen, P.-Y., and Yang, Y · 2025
Closest in time.
No, of course i can! refusal mechanisms can be exploited using harmless fine-tuning data
Kazdan, J., Yu, L., Schaeffer, R., Cundy, C., Koyejo, S., and Krishnamurthy, D · 2025
Closest in time.
Detecting instruction fine-tuning attack on language models with influence function
Li, J · 2025
Closest in time.
Safety alignment shouldn’t be complicated, 2025
Li, J. and Kim, J.-E · 2025
Closest in time.
Salora: Safety-alignment preserved low-rank adaptation
Li, M., Si, W. M., Backes, M., Zhang, Y., and Wang, Y · 2025
Closest in time.
Lookahead tuning: Safer language models via partial answer previews
Liu, K., Wang, M., Luo, Y., Yuan, L., Sun, M., Zhang, N., Liang, L., Zhang, Z., Zhou, J., and Chen, H · 2025
Closest in time.
Safe delta: Consistently preserving safety when fine-tuning llms on diverse datasets
Lu, N., Liu, S., Wu, J., Chen, W., Zhang, Z., Ong, Y.-S., Wang, Q., and Tang, K · 2025
Closest in time.
Shape it up! restoring llm safety during finetuning, 2025
Peng, S., Chen, P.-Y., Chi, J., Lee, S., and Chau, D. H · 2025
Closest in time.
Mitigating fine-tuning risks in llms via safety-aware probing optimization
Wu, C., Zhang, Z., Wei, Z., Zhang, Y., and Sun, M · 2025
Closest in time.
Alleviating the fear of losing alignment in llm fine-tuning
Yang, K., Tao, G., Chen, X., and Xu, J · 2025
Closest in time.