Fetching the paper…
Reading the bibliography…
Black-box finetuning is an emerging interface for adapting state-of-the-art language models to user needs.
Targeted backdoor attacks on deep learning systems using data poisoning
Chen, X., Liu, C., Li, B., Lu, K., and Song, D · 2017
Earlier work this paper cites.
Think you have solved question answering? Try ARC, the AI2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Concealed data poisoning attacks on NLP models
Wallace, E., Zhao, T., Feng, S., and Singh, S · 2021
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Earlier work this paper cites.
Backdoor learning: A survey
Li, Y., Jiang, Y., Li, Z., and Xia, S.-T · 2022
Earlier work this paper cites.
Introducing ChatGPT
OpenAI · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Earlier work this paper cites.
Solving math word problems with process-and outcome-based feedback
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
ShareGPT_Vicuna_unfiltered
anon8231489123 · 2023
Earlier work this paper cites.
Introducing Claude 2.1, Nov 2023
Anthropic · 2023
Earlier work this paper cites.
Bianchi, F., Suzgun, M., Attanasio, G., Röttger, P., Jurafsky, D., Hashimoto, T., and Zou, J · 2023
Earlier work this paper cites.
BadLLaMA: cheaply removing safety fine-tuning from LLaMA 2-chat 13B, 2023
Gade, P., Lermen, S., Rogers-Smith, C., and Ladish, J · 2023
Cited alongside, same era.
Bard: A conversational AI tool by Google, 2023
Google · 2023
Cited alongside, same era.
Overthinking the truth: Understanding how language models process false demonstrations, 2023
Halawi, D., Denain, J.-S., and Steinhardt, J · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models, 2023
Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., Chiang, P., Goldblum, M., Saha, A., Geiping, J., and Goldstein, T · 2023
Cited alongside, same era.
LoRA fine-tuning efficiently undoes safety training in LLaMA 2-Chat 70B, 2023
Lermen, S., Rogers-Smith, C., and Ladish, J · 2023
Cited alongside, same era.
OpenAI debuts GPT-4 Turbo and fine-tuning program for GPT-4, Nov 2023
Wiggers, K · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Later among the works it cites.
GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher
Yuan, Y., Jiao, W., Wang, W., Huang, J.-t., He, P., Shi, S., and Tu, Z · 2023
Later among the works it cites.
Removing RLHF protections in GPT-4 via fine-tuning
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D · 2023
Later among the works it cites.
Learning and forgetting unsafe examples in large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Cited alongside, same era.
Pelrine, K., Taufeeque, M., Zając, M., McLean, E., and Gleave, A · 2023
Cited alongside, same era.
On the exploitability of instruction tuning
Shu, M., Wang, J., Zhu, C., Geiping, J., Xiao, C., and Goldstein, T · 2023
Cited alongside, same era.
Stanford Alpaca: An instruction-following LLaMA model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
LLaMA 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Poisoning language models during instruction tuning
Wan, A., Wallace, E., Shen, S., and Klein, D · 2023
Cited alongside, same era.
Jailbroken: How does LLM safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Cited alongside, same era.
Zhao, J., Deng, Z., Madras, D., Zou, J., and Ren, M · 2023
Later among the works it cites.
Making harmful behaviors unlearnable for large language models, 2023
Zhou, X., Lu, Y., Ma, R., Gui, T., Zhang, Q., and Huang, X · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D · 2024
Closest in time.
OpenAI API Documentation
OpenAI · 2024
Closest in time.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2024
Closest in time.
Robust prompt optimization for defending language models against jailbreaking attacks, 2024
Zhou, A., Li, B., and Wang, H · 2024
Closest in time.
Improving alignment and robustness with circuit breakers, 2024
Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., Wang, R., Kolter, Z., Fredrikson, M., and Hendrycks, D · 2024
Closest in time.