Fetching the paper…
Reading the bibliography…
Recent research shows that Large Language Models (LLMs) are vulnerable to harmful fine-tuning attacks -- models lose their safety alignment ability after fine-tuning on a few harmful samples.
Convex optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
Griffith, S., Subramanian, K., Scholz, J., Isbell, C. L., and Thomaz, A. L · 2013
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Bianchi, F., Suzgun, M., Attanasio, G., Röttger, P., Jurafsky, D., Hashimoto, T., and Zou, J · 2023
Earlier work this paper cites.
Safe rlhf: Safe reinforcement learning from human feedback
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y · 2023
Earlier work this paper cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Earlier work this paper cites.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Sun, R., Wang, Y., and Yang, Y · 2023
Earlier work this paper cites.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b
Lermen, S., Rogers-Smith, C., and Ladish, J · 2023
Earlier work this paper cites.
Training socially aligned language models in simulated human society
Liu, R., Yang, R., Jia, C., Zhang, G., Zhou, D., Dai, A. M., Yang, D., and Vosoughi, S · 2023
Earlier work this paper cites.
Fine-tuning can cripple your foundation model; preserving features may be the solution
Mukhoti, J., Gal, Y., Torr, P. H., and Dokania, P. K · 2023
Earlier work this paper cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Earlier work this paper cites.
Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails
Rebedea, T., Dinu, R., Sreedhar, M., Parisien, C., and Cohen, J · 2023
Earlier work this paper cites.
Preference ranking optimization for human alignment
Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H · 2023
Earlier work this paper cites.
Pairwise proximal policy optimization: Harnessing relative feedback for llm alignment
Wu, T., Zhu, B., Zhang, R., Wen, Z., Ramchandran, K., and Jiao, J · 2023
Earlier work this paper cites.
Shadow alignment: The ease of subverting safely-aligned language models
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Earlier work this paper cites.
Selfee: Iterative self-revising llm empowered by self-feedback generation
Ye, S., Jo, Y., Kim, D., Kim, S., Hwang, H., and Seo, M · 2023
Earlier work this paper cites.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Cited alongside, same era.
Removing rlhf protections in gpt-4 via fine-tuning
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Defending against unforeseen failure modes with latent adversarial training
Casper, S., Schulze, L., Patel, O., and Hadfield-Menell, D · 2024
Cited alongside, same era.
Oml: Open, monetizable, and loyal ai
Cheng, Z., Contente, E., Finch, B., Golev, O., Hayase, J., Miller, A., Moshrefi, N., Nasery, A., Nailwal, S., Oh, S., et al · 2024
Navigating the safety landscape: Measuring risks in finetuning large language models
Peng, S., Chen, P.-Y., Hull, M., and Chau, D. H · 2024
Later among the works it cites.
Open problems in technical ai governance
Reuel, A., Bucknall, B., Casper, S., Fist, T., Soder, L., Aarne, O., Hammond, L., Ibrahim, L., Chan, A., Wills, P., et al · 2024
Later among the works it cites.
Seal: Safety-enhanced aligned llm fine-tuning via bilevel data selection
Shen, H., Chen, P.-Y., Das, P., and Chen, T · 2024
Later among the works it cites.
Open-ethical ai: Advancements in open-source human-centric neural language models
Sicari, S., Cevallos M, J. F., Rizzardi, A., and Coen-Porisini, A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Safety-aware fine-tuning of large language models
Choi, H. K., Du, X., and Li, Y · 2024
Cited alongside, same era.
Recent advances in attack and defense approaches of large language models
Cui, J., Xu, Y., Huang, Z., Zhou, S., Jiao, J., and Zhang, J · 2024
Cited alongside, same era.
Towards secure tuning: Mitigating security risks arising from benign instruction fine-tuning
Du, Y., Zhao, S., Cao, J., Ma, M., Zhao, D., Fan, F., Liu, T., and Qin, B · 2024
Cited alongside, same era.
Mimicking user data: On mitigating fine-tuning risks in closed large language models
Eiras, F., Petrov, A., Torr, P. H., Kumar, M. P., and Bibi, A · 2024
Cited alongside, same era.
Enhancing ai safety through the fusion of low rank adapters
Gudipudi, S. S., Vipparla, S., Singh, H., Goel, S., and Kumaraguru, P · 2024
Cited alongside, same era.
The vllm safety paradox: Dual ease in jailbreak attack and defense
Guo, Y., Jiao, F., Nie, L., and Kankanhalli, M · 2024
Cited alongside, same era.
Covert malicious finetuning: Challenges in safeguarding llm adaptation
Halawi, D., Wei, A., Wallace, E., Wang, T. T., Haghtalab, N., and Steinhardt, J · 2024
Cited alongside, same era.
Tamirisa, R., Bharathi, B., Phan, L., Zhou, A., Gatti, A., Suresh, T., Lin, M., Wang, J., Wang, R., Arel, R., et al · 2024
Later among the works it cites.
Hˆ 3 fusion: Helpful, harmless, honest fusion of aligned llms
Tekin, S. F., Ilhan, F., Huang, T., Hu, S., Yahn, Z., and Liu, L · 2024
Later among the works it cites.
Operationalizing a threat model for red-teaming large language models (llms)
Verma, A., Krishna, S., Gehrmann, S., Seshadri, M., Pradhan, A., Ault, T., Barrett, L., Rabinowitz, D., Doucette, J., and Phan, N · 2024
Later among the works it cites.
Mitigating fine-tuning jailbreak attack with backdoor enhanced alignment
Wang, J., Li, J., Li, Y., Qi, X., Chen, M., Hu, J., Li, Y., Li, B., and Xiao, C · 2024
Later among the works it cites.
Assessing the brittleness of safety alignment via pruning and low-rank modifications
Wei, B., Huang, K., Huang, Y., Xie, T., Qi, X., Xia, M., Mittal, P., Wang, M., and Henderson, P · 2024
Later among the works it cites.
Wu, D., Lu, X., Zhao, Y., and Qin, B · 2024
Later among the works it cites.
Emerging safety attack and defense in federated instruction tuning of large language models
Ye, R., Chai, J., Liu, X., Yang, Y., Wang, Y., and Chen, S · 2024
Later among the works it cites.
On the vulnerability of safety alignment in open-access llms
Yi, J., Ye, R., Chen, Q., Zhu, B., Chen, S., Lian, D., Sun, G., Xie, X., and Wu, F · 2024
Later among the works it cites.
Locking down the finetuned llms safety
Zhu, M., Yang, L., Wei, Y., Zhang, N., and Zhang, Y · 2024
Later among the works it cites.
Safety fine-tuning at (almost) no cost: A baseline for vision large language models
Zong, Y., Bohdal, O., Yu, T., Yang, Y., and Hospedales, T · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers
Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., Kolter, J. Z., Fredrikson, M., and Hendrycks, D · 2024
Later among the works it cites.
Probe before you talk: Towards black-box defense against backdoor unalignment for large language models
Anonymous · 2025
Closest in time.
Open problems in machine unlearning for ai safety
Barez, F., Fu, T., Prabhu, A., Casper, S., Sanyal, A., Bibi, A., O’Gara, A., Kirk, R., Bucknall, B., Fist, T., et al · 2025
Closest in time.
Aegis2. 0: A diverse ai safety dataset and risks taxonomy for alignment of llm guardrails
Ghosh, S., Varshney, P., Sreedhar, M. N., Padmakumar, A., Rebedea, T., Varghese, J. R., and Parisien, C · 2025
Closest in time.
Salora: Safety-alignment preserved low-rank adaptation
Li, M., Si, W. M., Backes, M., Zhang, Y., and Wang, Y · 2025
Closest in time.