Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) fine-tuning transforms large language models while creating a vulnerability we experimentally verify: Our experiment shows that malicious RL fine-tuning dismantles safety guardrails with remarkable efficiency, requiring only 50 steps and minimal adversarial prompts, with harmful escalating from 0-2 to 7-9.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Mistral AI · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
@ Meta · 2023
Earlier work this paper cites.
Unveiling the implicit toxicity in large language models
Jiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang, Chengfei Li, Jinfeng Bai, and Minlie Huang · 2023
Earlier work this paper cites.
Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt, 2023
Jiawen Shi, Yixin Liu, Pan Zhou, and Lichao Sun · 2023
Earlier work this paper cites.
Shadow alignment: The ease of subverting safely-aligned language models, 2023
Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin · 2023
Earlier work this paper cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Earlier work this paper cites.
DEPN: Detecting and editing privacy neurons in pretrained language models
Xinwei Wu, Junzhuo Li, Minghui Xu, Weilong Dong, Shuangzhi Wu, Chao Bian, and Deyi Xiong · 2023
Earlier work this paper cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo · 2024
Earlier work this paper cites.
@ OpenAI · 2024
Cited alongside, same era.
The llama 3 herd of models, 2024
@ Meta · 2024
Cited alongside, same era.
Vaccine: Perturbation-aware alignment for large language models against harmful fine-tuning attack
Tiansheng Huang, Sihao Hu, and Ling Liu · 2024
Cited alongside, same era.
Immunization against harmful fine-tuning attacks
Domenic Rosati, Jan Wehner, Kai Williams, Lukasz Bartoszcze, Hassan Sajjad, and Frank Rudzicz · 2024
Cited alongside, same era.
Lisa: Lazy safety alignment for large language models against harmful fine-tuning attack
Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, and Ling Liu · 2024
Cited alongside, same era.
Backdooralign: Mitigating fine-tuning based jailbreak attack with backdoor enhanced safety alignment
Jiongxiao Wang, Jiazhao Li, Yiquan Li, Xiangyu Qi, Junjie Hu, Yixuan Li, Patrick McDaniel, Muhao Chen, Bo Li, and Chaowei Xiao · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model, 2024
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
@ DeepSeek-AI · 2025
Closest in time.
Qwen2.5 technical report, 2025
Alibaba · 2025
Closest in time.
Booster: Tackling harmful fine-tuning for large language models via attenuating harmful perturbation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Piyush Jha, Arnav Arora, and Vijay Ganesh · 2024
Cited alongside, same era.
Enhancing llm safety via constrained direct preference optimization, 2024
Zixuan Liu, Xiaolin Sun, and Zizhan Zheng · 2024
Cited alongside, same era.
Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs
Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Olivier Pietquin, Ahmet Üstün, and Sara Hooker · 2024
Cited alongside, same era.
ODIN: Disentangled reward mitigates hacking in RLHF
Lichang Chen, Chen Zhu, Jiuhai Chen, Davit Soselia, Tianyi Zhou, Tom Goldstein, Heng Huang, Mohammad Shoeybi, and Bryan Catanzaro · 2024
Cited alongside, same era.
Removing RLHF protections in GPT-4 via fine-tuning
Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang · 2024
Cited alongside, same era.
Towards understanding jailbreak attacks in LLMs: A representation space analysis
Yuping Lin, Pengfei He, Han Xu, Yue Xing, Makoto Yamada, Hui Liu, and Jiliang Tang · 2024
Cited alongside, same era.
Dissecting learning and forgetting in language model finetuning
Xiao Zhang and Ji Wu · 2024
Cited alongside, same era.
Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, and Ling Liu · 2025
Closest in time.
Tamper-resistant safeguards for open-weight llms, 2025
Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li, Dan Hendrycks, and Mantas Mazeika · 2025
Closest in time.
Defining and characterizing reward hacking, 2025
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger · 2025
Closest in time.
The energy loss phenomenon in rlhf: A new perspective on mitigating reward hacking, 2025
Yuchun Miao, Sen Zhang, Liang Ding, Yuqi Zhang, Lefei Zhang, and Dacheng Tao · 2025
Closest in time.
Adversarial agents: Black-box evasion attacks with reinforcement learning, 2025
Kyle Domico, Jean-Charles Noirot Ferrand, Ryan Sheatsley, Eric Pauley, Josiah Hanna, and Patrick McDaniel · 2025
Closest in time.
Safety layers in aligned large language models: The key to llm security, 2025
Shen Li, Liuyi Yao, Lan Zhang, and Yaliang Li · 2025
Closest in time.
Safety alignment should be made more than just a few tokens deep
Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson · 2025
Closest in time.
Grpo-flat: Zero-shot grpo training framework with limited resources
Yijie Xu · 2025
Closest in time.