Fetching the paper…
Reading the bibliography…
Leading language model (LM) providers like OpenAI and Anthropic allow customers to fine-tune frontier LMs for specific use cases.
Antidote: Understanding and defending against poisoning of anomaly detectors
Rubinstein, B., Nelson, B., Huang, L., Joseph, A. D., Lau, S., Rao, S., Taft, N., and Tygar, J · 2009
Earlier work this paper cites.
How much spam can you take? an analysis of crowdsourcing results to increase accuracy
Vuurens, J., de Vries, A. P., and Eickhoff, C · 2011
Earlier work this paper cites.
Adversarial support vector machine learning
Zhou, Y., Kantarcioglu, M., Thuraisingham, B., and Xi, B · 2012
Earlier work this paper cites.
Is data clustering in adversarial settings secure?
Biggio, B., Pillai, I., Bulò, S. R., Ariu, D., Pelillo, M., and Roli, F · 2013
Earlier work this paper cites.
Poisoning complete-linkage hierarchical clustering
Biggio, B., Rota, B. S., Ignazio, P., Michele, M., Zemene, M. E., Marcello, P., and Fabio, R · 2014
Earlier work this paper cites.
The security of latent dirichlet allocation
Mei, S. and Zhu, X · 2015
Earlier work this paper cites.
Systematic poisoning attacks on and defenses for machine learning in healthcare
Mozaffari-Kermani, M., Sur-Kolay, S., Raghunathan, A., and Jha, N. K · 2015
Earlier work this paper cites.
Is feature selection secure against training data poisoning?
Xiao, H., Biggio, B., Brown, G., Fumera, G., Eckert, C., and Roli, F · 2015
Earlier work this paper cites.
Data poisoning attacks on factorization-based collaborative filtering
Li, B., Wang, Y., Singh, A., and Vorobeychik, Y · 2016
Earlier work this paper cites.
Combating Attacks and Abuse in Large Online Communities
Wang, G · 2016
Earlier work this paper cites.
Certified defenses for data poisoning attacks
Steinhardt, J., Koh, P. W., and Liang, P · 2017
Earlier work this paper cites.
A multi-dimensional machine learning approach to predict advanced malware
Bahtiyar, Ş., Yaman, M. B., and Altıniğne, C. Y · 2019
Earlier work this paper cites.
A survey of attacks against twitter spam detectors in an adversarial environment
Imam, N. H. and Vassilakis, V. G · 2019
Earlier work this paper cites.
A benchmark study of backdoor data poisoning defenses for deep neural network classifiers and a novel defense
Xiang, Z., Miller, D. J., and Kesidis, G · 2019
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Exploring data and model poisoning attacks to deep learning-based nlp systems
Marulli, F., Verde, L., and Campanile, L · 2021
Earlier work this paper cites.
Challenges in detoxifying language models, 2021
Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J · 2022
Earlier work this paper cites.
Truth serum: Poisoning machine learning models to reveal their secrets
Tramèr, F., Shokri, R., San Joaquin, A., Le, H., Jagielski, M., Hong, S., and Carlini, N · 2022
Earlier work this paper cites.
Not all poisons are created equal: Robust training against data poisoning
Yang, Y., Liu, T. Y., and Mirzasoleiman, B · 2022
Earlier work this paper cites.
Are aligned neural networks adversarially aligned?
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Koh, P. W., Ippolito, D., Tramèr, F., and Schmidt, L · 2023
Earlier work this paper cites.
Call for ai pause highlights potential dangers
Clarke, L · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations, 2023
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., and Khabsa, M · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Earlier work this paper cites.
Generative ai as a dangerous new form of media
Rosenberg, L · 2023
Earlier work this paper cites.
The dangers of generative artificial intelligence
Tredinnick, L. and Laybats, C · 2023
Earlier work this paper cites.
Poisoning language models during instruction tuning
Wan, A., Wallace, E., Shen, S., and Klein, D · 2023
Earlier work this paper cites.
Temporal robustness against data poisoning
Wang, W. and Feizi, S · 2023
Earlier work this paper cites.
Helpsteer: Multi-attribute helpfulness dataset for steerlm, 2023
Wang, Z., Dong, Y., Zeng, J., Adams, V., Sreedhar, M. N., Egert, D., Delalleau, O., Scowcroft, J. P., Kant, N., Swope, A., and Kuchaiev, O · 2023
Cited alongside, same era.
Jailbroken: How does LLM safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Cited alongside, same era.
Shadow alignment: The ease of subverting safely-aligned language models, 2023
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Cited alongside, same era.
Removing rlhf protections in gpt-4 via fine-tuning
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D · 2023
Cited alongside, same era.
The real dangers of generative ai
Allen, D. and Weyl, E. G · 2024
Cited alongside, same era.
Seal: Safety-enhanced aligned llm fine-tuning via bilevel data selection, 2024
Shen, H., Chen, P.-Y., Das, P., and Chen, T · 2024
Later among the works it cites.
On the exploitability of instruction tuning
Shu, M., Wang, J., Zhu, C., Geiping, J., Xiao, C., and Goldstein, T · 2024
Later among the works it cites.
Tamper-resistant safeguards for open-weight llms
Tamirisa, R., Bharathi, B., Phan, L., Zhou, A., Gatti, A., Suresh, T., Lin, M., Wang, J., Wang, R., Arel, R., et al · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology, 2024
Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., Tafti, P., Hussenot, L., Sessa, P. G., Chowdhery, A., Roberts, A., Barua, A., Botev, A., Castro-Ros, A., Slone, A., Héliou, A., Tacchetti, A., Bulanova, A., Paterson, A., Tsai, B., Shahriari, B., Lan, C. L., Choquette-Choo, C. A., Crepy, C., Cer, D., Ippolito, D., Reid, D., Buchatskaya, E., Ni, E., Noland, E., Yan, G., Tucker, G., Muraru, G.-C., Rozhdestvenskiy, G., Michalewski, H., Tenney, I., Grishchenko, I., Austin, J., Keeling, J., Labanowski, J., Lespiau, J.-B., Stanway, J., Brennan, J., Chen, J., Ferret, J., Chiu, J., Mao-Jones, J., Lee, K., Yu, K., Millican, K., Sjoesund, L. L., Lee, L., Dixon, L., Reid, M., Mikuła, M., Wirth, M., Sharman, M., Chinaev, N., Thain, N., Bachem, O., Chang, O., Wahltinez, O., Bailey, P., Michel, P., Yotov, P., Chaabouni, R., Comanescu, R., Jana, R., Anil, R., McIlroy, R., Liu, R., Mullins, R., Smith, S. L., Borgeaud, S., Girgin, S., Douglas, S., Pandya, S., Shakeri, S., De, S., Klimenko, T., Hennigan, T., Feinberg, V., Stokowiec, W., hui Chen, Y., Ahmed, Z., Gong, Z., Warkentin, T., Peran, L., Giang, M., Farabet, C., Vinyals, O., Dean, J., Kavukcuoglu, K., Hassabis, D., Ghahramani, Z., Eck, D., Barral, J., Pereira, F., Collins, E., Joulin, A., Fiedel, N., Senter, E., Andreev, A., and Kenealy, K · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Best-of-venom: Attacking RLHF by injecting poisoned preference data
Baumgärtner, T., Gao, Y., Alon, D., and Metzler, D · 2024
Cited alongside, same era.
Safety-tuned LLaMAs: Lessons from improving the safety of large language models that follow instructions
Bianchi, F., Suzgun, M., Attanasio, G., Rottger, P., Jurafsky, D., Hashimoto, T., and Zou, J · 2024
Cited alongside, same era.
Poisoning web-scale training datasets is practical
Carlini, N., Jagielski, M., Choquette-Choo, C. A., Paleka, D., Pearce, W., Anderson, H., Terzis, A., Thomas, K., and Tramèr, F · 2024
Cited alongside, same era.
Defending against unforeseen failure modes with latent adversarial training, 2024
Casper, S., Schulze, L., Patel, O., and Hadfield-Menell, D · 2024
Cited alongside, same era.
Safety-aware fine-tuning of large language models, 2024
Choi, H. K., Du, X., and Li, Y · 2024
Cited alongside, same era.
Building guardrails for large language models
Dong, Y., Mu, R., Jin, G., Qi, Y., Hu, J., Zhao, X., Meng, J., Ruan, W., and Huang, X · 2024
Cited alongside, same era.
Towards secure tuning: Mitigating security risks arising from benign instruction fine-tuning, 2024
Du, Y., Zhao, S., Cao, J., Ma, M., Zhao, D., Fan, F., Liu, T., and Qin, B · 2024
Cited alongside, same era.
Later among the works it cites.
Assessing the brittleness of safety alignment via pruning and low-rank modifications, 2024
Wei, B., Huang, K., Huang, Y., Xie, T., Qi, X., Xia, M., Mittal, P., Wang, M., and Henderson, P · 2024
Later among the works it cites.
Shadowcast: Stealthy data poisoning attacks against vision-language models, 2024
Xu, Y., Yao, J., Shu, M., Sun, Y., Wu, Z., Yu, N., Goldstein, T., and Huang, F · 2024
Later among the works it cites.
No free lunch for defending against prefilling attack by in-context learning, 2024
Xue, Z., Liu, G., Chen, B., Johnson, K. M., and Pedarsani, R · 2024
Later among the works it cites.
Social dangers of generative artificial intelligence: review and guidelines
Yang, A. and Yang, T. A · 2024
Later among the works it cites.
Shieldgemma: Generative ai content moderation based on gemma, 2024
Zeng, W., Liu, Y., Mullins, R., Peran, L., Fernandez, J., Harkous, H., Narasimhan, K., Proud, D., Kumar, P., Radharapu, B., Sturman, O., and Wahltinez, O · 2024
Later among the works it cites.
Locking down the finetuned llms safety, 2024
Zhu, M., Yang, L., Wei, Y., Zhang, N., and Zhang, Y · 2024
Later among the works it cites.
Openai bug bounty program, 2025
Bugcrowd · 2025
Closest in time.
Amazing ”jailbreak” bypasses chatgpt’s ethics safeguards, February 4 2023
Christian, J · 2025
Closest in time.
Fundamental limitations in defending llm finetuning apis, 2025
Davies, X., Winsor, E., Korbak, T., Souly, A., Kirk, R., de Witt, C. S., and Gal, Y · 2025
Closest in time.
Do as i do (safely): Mitigating task-specific fine-tuning risks in large language models, 2025
Eiras, F., Petrov, A., Torr, P. H. S., Kumar, M. P., and Bibi, A · 2025
Closest in time.
Covert malicious finetuning: challenges in safeguarding llm adaptation
Halawi, D., Wei, A., Wallace, E., Wang, T., Haghtalab, N., and Steinhardt, J · 2025
Closest in time.
Your task may vary: A systematic understanding of alignment and safety degradation when fine-tuning LLMs, 2025
Hsiung, L., Pang, T., Tang, Y.-C., Song, L., Ho, T.-Y., Chen, P.-Y., and Yang, Y · 2025
Closest in time.
Safe lora: the silver lining of reducing safety risks when fine-tuning large language models, 2025
Hsu, C.-Y., Tsai, Y.-L., Lin, C.-H., Chen, P.-Y., Yu, C.-M., and Huang, C.-Y · 2025
Closest in time.
Virus: Harmful fine-tuning attack for large language models bypassing guardrail moderation, 2025
Huang, T., Hu, S., Ilhan, F., Tekin, S. F., and Liu, L · 2025
Closest in time.
Safety alignment shouldn’t be complicated, 2025
Li, J. and Kim, J.-E · 2025
Closest in time.
Keeping llms aligned after fine-tuning: The crucial role of prompt templates, 2025
Lyu, K., Zhao, H., Gu, X., Yu, D., Goyal, A., and Arora, S · 2025
Closest in time.
Fine-tuning models, 2024
OpenAI · 2025
Closest in time.
Usage policies
OpenAI · 2025
Closest in time.
Towards understanding the fragility of multilingual llms against fine-tuning attacks, 2025
Poppi, S., Yong, Z.-X., He, Y., Chern, B., Zhao, H., Yang, A., and Chi, J · 2025
Closest in time.
Operationalizing a threat model for red-teaming large language models (LLMs)
Verma, A., Krishna, S., Gehrmann, S., Seshadri, M., Pradhan, A., Doucette, J. A., Rabinowitz, D., Barrett, L., Ault, T., and Phan, H · 2025
Closest in time.
Wu, D., Lu, X., Zhao, Y., and Qin, B · 2025
Closest in time.
Probe before you talk: Towards black-box defense against backdoor unalignment for large language models
Yi, B., Huang, T., Chen, S., Li, T., Liu, Z., Chu, Z., and Li, Y · 2025
Closest in time.
Safety fine-tuning at (almost) no cost: a baseline for vision large language models
Zong, Y., Bohdal, O., Yu, T., Yang, Y., and Hospedales, T · 2025
Closest in time.