Fetching the paper…
Reading the bibliography…
Recent studies have uncovered a troubling vulnerability in the fine-tuning stage of large language models (LLMs): even fine-tuning on entirely benign datasets can lead to a significant increase in the harmfulness of LLM outputs.
Shortcut learning in deep neural networks
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A · 2020
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
Pruthi, G., Liu, F., Kale, S., and Sundararajan, M · 2020
Earlier work this paper cites.
A new generation of perspective api: Efficient multilingual character-level transformers
Lees, A., Tran, V. Q., Tay, Y., Sorensen, J., Gupta, J., Metzler, D., and Vasserman, L · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Bianchi, F., Suzgun, M., Attanasio, G., Röttger, P., Jurafsky, D., Hashimoto, T., and Zou, J · 2023
Earlier work this paper cites.
Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023
Conover, M., Hayes, M., Mathur, A., Xie, J., Wan, J., Shah, S., Ghodsi, A., Wendell, P., Zaharia, M., and Xin, R · 2023
Earlier work this paper cites.
Shortcut learning of large language models in natural language understanding
Du, M., He, F., Zou, N., Tao, D., and Hu, X · 2023
Earlier work this paper cites.
Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b
Gade, P., Lermen, S., Rogers-Smith, C., and Ladish, J · 2023
Earlier work this paper cites.
Regulating chatgpt and other large generative ai models. arxiv
Hacker, P., Engel, A., and Mauer, M · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Earlier work this paper cites.
Datainf: Efficiently estimating data influence in lora-tuned llms and diffusion models
Kwon, Y., Wu, E., Wu, K., and Zou, J · 2023
Earlier work this paper cites.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b
Lermen, S., Rogers-Smith, C., and Ladish, J · 2023
Earlier work this paper cites.
A holistic approach to undesired content detection in the real world
Markov, T., Zhang, C., Agarwal, S., Nekoul, F. E., Lee, T., Adler, S., Jiang, A., and Weng, L · 2023
Earlier work this paper cites.
Fine-tuning can cripple your foundation model; preserving features may be the solution
Mukhoti, J., Gal, Y., Torr, P. H., and Dokania, P. K · 2023
Earlier work this paper cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Earlier work this paper cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Earlier work this paper cites.
Shadow alignment: The ease of subverting safely-aligned language models
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Earlier work this paper cites.
Removing rlhf protections in gpt-4 via fine-tuning
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Cited alongside, same era.
Chhabra, A., Li, B., Chen, J., Mohapatra, P., and Liu, H · 2024
Cited alongside, same era.
Safety-aware fine-tuning of large language models
Choi, H. K., Du, X., and Li, Y · 2024
Cited alongside, same era.
Covert malicious finetuning: Challenges in safeguarding llm adaptation, 2024
Seal: Safety-enhanced aligned llm fine-tuning via bilevel data selection
Shen, H., Chen, P.-Y., Das, P., and Chen, T · 2024
Later among the works it cites.
Tamper-resistant safeguards for open-weight llms
Tamirisa, R., Bharathi, B., Phan, L., Zhou, A., Gatti, A., Suresh, T., Lin, M., Wang, J., Wang, R., Arel, R., et al · 2024
Later among the works it cites.
Assessing the brittleness of safety alignment via pruning and low-rank modifications
Wei, B., Huang, K., Huang, Y., Xie, T., Qi, X., Xia, M., Mittal, P., Wang, M., and Henderson, P · 2024
Later among the works it cites.
Wu, D., Lu, X., Zhao, Y., and Qin, B · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Halawi, D., Wei, A., Wallace, E., Wang, T. T., Haghtalab, N., and Steinhardt, J · 2024
Cited alongside, same era.
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms
Han, S., Rao, K., Ettinger, A., Jiang, L., Lin, B. Y., Lambert, N., Choi, Y., and Dziri, N · 2024
Cited alongside, same era.
What is in your safe data? identifying benign data that breaks safety, 2024
He, L., Xia, M., and Henderson, P · 2024
Cited alongside, same era.
Safe lora: the silver lining of reducing safety risks when fine-tuning large language models
Hsu, C.-Y., Tsai, Y.-L., Lin, C.-H., Chen, P.-Y., Yu, C.-M., and Huang, C.-Y · 2024
Cited alongside, same era.
Fine-tuning and utilization methods of domain-specific llms
Jeong, C · 2024
Cited alongside, same era.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Chen, B., Sun, R., Wang, Y., and Yang, Y · 2024
Cited alongside, same era.
Publicly shareable clinical large language model built on synthetic clinical notes
Kweon, S., Kim, J., Kim, J., Im, S., Cho, E., Bae, S., Oh, J., Lee, G., Moon, J. H., You, S. C., Baek, S., Han, C. H., Jung, Y. B., Jo, Y., and Choi, E · 2024
Cited alongside, same era.
No two devils alike: Unveiling distinct mechanisms of fine-tuning attacks, 2024
Leong, C. T., Cheng, Y., Xu, K., Wang, J., Wang, H., and Li, W · 2024
Cited alongside, same era.
Xia, M., Malladi, S., Gururangan, S., Arora, S., and Chen, D · 2024
Later among the works it cites.
Emerging safety attack and defense in federated instruction tuning of large language models, 2024
Ye, R., Chai, J., Liu, X., Yang, Y., Wang, Y., and Chen, S · 2024
Later among the works it cites.
On the vulnerability of safety alignment in open-access LLMs
Yi, J., Ye, R., Chen, Q., Zhu, B., Chen, S., Lian, D., Sun, G., Xie, X., and Wu, F · 2024
Later among the works it cites.
On prompt-driven safeguarding for large language models
Zheng, C., Yin, F., Zhou, H., Meng, F., Zhou, J., Chang, K.-W., Huang, M., and Peng, N · 2024
Later among the works it cites.
Locking down the finetuned llms safety, 2024
Zhu, M., Yang, L., Wei, Y., Zhang, N., and Zhang, Y · 2024
Later among the works it cites.
Model tampering attacks enable more rigorous evaluations of llm capabilities, 2025
Che, Z., Casper, S., Kirk, R., Satheesh, A., Slocum, S., McKinney, L. E., Gandikota, R., Ewart, A., Rosati, D., Wu, Z., Cai, Z., Chughtai, B., Gal, Y., Huang, F., and Hadfield-Menell, D · 2025
Closest in time.
On weaponization-resistant large language models with prospect theoretic alignment
Cheng, Z., Zhang, M., Sun, J., and Dai, W · 2025
Closest in time.
Your task may vary: A systematic understanding of alignment and safety degradation when fine-tuning LLMs, 2025
Hsiung, L., Pang, T., Tang, Y.-C., Song, L., Ho, T.-Y., Chen, P.-Y., and Yang, Y · 2025
Closest in time.
Virus: Harmful fine-tuning attack for large language models bypassing guardrail moderation, 2025
Huang, T., Hu, S., Ilhan, F., Tekin, S. F., and Liu, L · 2025
Closest in time.
Safety alignment shouldn’t be complicated
Li, J. and Kim, J.-E · 2025
Closest in time.
Salora: Safety-alignment preserved low-rank adaptation
Li, M., Si, W. M., Backes, M., Zhang, Y., and Wang, Y · 2025
Closest in time.
On evaluating the durability of safeguards for open-weight LLMs
Qi, X., Wei, B., Carlini, N., Huang, Y., Xie, T., He, L., Jagielski, M., Nasr, M., Mittal, P., and Henderson, P · 2025
Closest in time.
Wang, Y., Huang, T., Shen, L., Yao, H., Luo, H., Liu, R., Tan, N., Huang, J., and Tao, D · 2025
Closest in time.
Probe before you talk: Towards black-box defense against backdoor unalignment for large language models
Yi, B., Huang, T., Chen, S., Li, T., Liu, Z., Chu, Z., and Li, Y · 2025
Closest in time.