Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are vulnerable to jailbreaking attacks that lead to generation of inappropriate or harmful content.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P. F., and Irving, G · 1909
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing NLP
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 2019
Earlier work this paper cites.
AutoPrompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., IV, R. L. L., Wallace, E., and Singh, S · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
von Werra, L., Belkada, Y., Tunstall, L., Beeching, E., Thrush, T., Lambert, N., and Huang, S · 2020
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers
Guo, C., Sablayrolles, A., Jegou, H., and Kiela, D · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J · 2022
Earlier work this paper cites.
Learning to optimize: A primer and a benchmark
Chen, T., Chen, X., Chen, W., Heaton, H., Liu, J., Wang, Z., and Yin, W · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., Jones, A., Bowman, S., Chen, A., Conerly, T., DasSarma, N., Drain, D., Elhage, N., El-Showk, S., Fort, S., Hatfield-Dodds, Z., Henighan, T., Hernandez, D., Hume, T., Jacobson, J., Johnston, S., Kravec, S., Olsson, C., Ringer, S., Tran-Johnson, E., Amodei, D., Brown, T., Joseph, N., McCandlish, S., Olah, C., Kaplan, J., and Clark, J · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Red teaming language models with language models, 2022
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Perez, F. and Ribeiro, I · 2022
Earlier work this paper cites.
Large language models are human-level prompt engineers, 2022
Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J · 2022
Earlier work this paper cites.
Tutorial on amortized optimization
Amos, B · 2023
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling, 2023
Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., and van der Wal, O · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Earlier work this paper cites.
Instructzero: Efficient instruction optimization for black-box large language models
Chen, L., Chen, J., Goldstein, T., Huang, H., and Zhou, T · 2023
Earlier work this paper cites.
Black-box prompt learning for pre-trained language models
Diao, S., Huang, Z., Xu, R., Li, X., Lin, Y., Zhou, X., and Zhang, T · 2023
Earlier work this paper cites.
Connecting large language models with evolutionary algorithms yields powerful prompt optimizers, 2023
Guo, Q., Wang, R., Guo, J., Li, B., Song, K., Tan, X., Liu, G., Bian, J., and Yang, Y · 2023
Cited alongside, same era.
Catastrophic jailbreak of open-source llms via exploiting generation
Huang, Y., Gupta, S., Xia, M., Li, K., and Chen, D · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., Chiang, P.-y., Goldblum, M., Saha, A., Geiping, J., and Goldstein, T · 2023
Cited alongside, same era.
Mistral 7b, 2023
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Cited alongside, same era.
Automatically auditing large language models via discrete optimization, 2023
Jailbreak and guard aligned language models with only few in-context demonstrations, 2023
Wei, Z., Wang, Y., and Wang, Y · 2023
Later among the works it cites.
Large language models as optimizers, 2023
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts, 2023
Yu, J., Lin, X., Yu, Z., and Xing, X · 2023
Later among the works it cites.
Removing rlhf protections in gpt-4 via fine-tuning, 2023
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D · 2023
Later among the works it cites.
Autodan: Automatic and interpretable adversarial attacks on large language models
Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jones, E., Dragan, A., Raghunathan, A., and Steinhardt, J · 2023
Cited alongside, same era.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Kang, D., Li, X., Stoica, I., Guestrin, C., Zaharia, M., and Hashimoto, T · 2023
Cited alongside, same era.
Llm+p: Empowering large language models with optimal planning proficiency
Liu, B., Jiang, Y., Zhang, X., Liu, Q., Zhang, S., Biswas, J., and Stone, P · 2023
Cited alongside, same era.
Black box adversarial prompting for foundation models, 2023
Maus, N., Chao, P., Wong, E., and Gardner, J · 2023
Cited alongside, same era.
Tree of attacks: Jailbreaking black-box llms automatically
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H. S., Singer, Y., and Karbasi, A · 2023
Cited alongside, same era.
Do anything now. https://gist.github.com/coolaj86/6f4f7b30129b0251f61fa7baaa881516 , 2023
ONeal, A · 2023
Cited alongside, same era.
Autoplan: Automatic planning of interactive decision-making tasks with large language models
Ouyang, S. and Li, L · 2023
Cited alongside, same era.
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J · 2023
Cited alongside, same era.
Zou, A., Wang, Z., Kolter, Z. J., and Fredrikson, M · 2023
Later among the works it cites.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024
Ahmadian, A., Cremer, C., Gallé, M., Fadaee, M., Kreutzer, J., Pietquin, O., Üstün, A., and Hooker, S · 2024
Closest in time.
The llama 3 herd of models
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., Rao, A., Zhang, A., Rodriguez, A., Gregerson, A., Spataru, A., Rozière, B., Biron, B., Tang, B., Chern, B., Caucheteux, C., Nayak, C., Bi, C., Marra, C., McConnell, C., Keller, C., Touret, C., Wu, C., Wong, C., Ferrer, C. C., Nikolaidis, C., Allonsius, D., Song, D., Pintz, D., Livshits, D., Esiobu, D., Choudhary, D., Mahajan, D., Garcia-Olano, D., Perino, D., Hupkes, D., Lakomkin, E., AlBadawy, E., Lobanova, E., Dinan, E., Smith, E. M., Radenovic, F., Zhang, F., Synnaeve, G., Lee, G., Anderson, G. L., Nail, G., Mialon, G., Pang, G., Cucurell, G., Nguyen, H., Korevaar, H., Xu, H., Touvron, H., Zarov, I., Ibarra, I. A., Kloumann, I. M., Misra, I., Evtimov, I., Copet, J., Lee, J., Geffert, J., Vranes, J., Park, J., Mahadeokar, J., Shah, J., van der Linde, J., Billock, J., Hong, J., Lee, J., Fu, J., Chi, J., Huang, J., Liu, J., Wang, J., Yu, J., Bitton, J., Spisak, J., Park, J., Rocca, J., Johnstun, J., Saxe, J., Jia, J., Alwala, K. V., Upasani, K., Plawiak, K., Li, K., Heafield, K., Stone, K., and et al · 2024
Closest in time.
Liao, Z. and Sun, H · 2024
Closest in time.
Mistral-large api, mistral ai team, 2024
Mistral-Large, 2024 · 2024
Closest in time.
Gpt-4 technical report
OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett-Shapiro, G., Berner, C., Bogdonoff, L., Boiko, O., Boyd, M., Brakman, A.-L., Brockman, G., Brooks, T., Brundage, M., Button, K., Cai, T., Campbell, R., Cann, A., Carey, B., Carlson, C., Carmichael, R., Chan, B., Chang, C., Chantzis, F., Chen, D., Chen, S., Chen, R., Chen, J., Chen, M., Chess, B., Cho, C., Chu, C., Chung, H. W., Cummings, D., Currier, J., Dai, Y., Decareaux, C., Degry, T., Deutsch, N., Deville, D., Dhar, A., Dohan, D., Dowling, S., Dunning, S., Ecoffet, A., Eleti, A., Eloundou, T., Farhi, D., Fedus, L., Felix, N., Fishman, S. P., Forte, J., Fulford, I., Gao, L., Georges, E., Gibson, C., Goel, V., Gogineni, T., Goh, G., Gontijo-Lopes, R., Gordon, J., Grafstein, M., Gray, S., Greene, R., Gross, J., Gu, S. S., Guo, Y., Hallacy, C., Han, J., Harris, J., He, Y., Heaton, M., Heidecke, J., Hesse, C., Hickey, A., Hickey, W., Hoeschele, P., Houghton, B., Hsu, K., Hu, S., Hu, X., Huizinga, J., Jain, S., Jain, S., Jang, J., Jiang, A., Jiang, R., Jin, H., Jin, D., Jomoto, S., Jonn, B., Jun, H., Kaftan, T., Łukasz Kaiser, Kamali, A., Kanitscheider, I., Keskar, N. S., Khan, T., Kilpatrick, L., Kim, J. W., Kim, C., Kim, Y., Kirchner, J. H., Kiros, J., Knight, M., Kokotajlo, D., Łukasz Kondraciuk, Kondrich, A., Konstantinidis, A., Kosic, K., Krueger, G., Kuo, V., Lampe, M., Lan, I., Lee, T., Leike, J., Leung, J., Levy, D., Li, C. M., Lim, R., Lin, M., Lin, S., Litwin, M., Lopez, T., Lowe, R., Lue, P., Makanju, A., Malfacini, K., Manning, S., Markov, T., Markovski, Y., Martin, B., Mayer, K., Mayne, A., McGrew, B., McKinney, S. M., McLeavey, C., McMillan, P., McNeil, J., Medina, D., Mehta, A., Menick, J., Metz, L., Mishchenko, A., Mishkin, P., Monaco, V., Morikawa, E., Mossing, D., Mu, T., Murati, M., Murk, O., Mély, D., Nair, A., Nakano, R., Nayak, R., Neelakantan, A., Ngo, R., Noh, H., Ouyang, L., O’Keefe, C., Pachocki, J., Paino, A., Palermo, J., Pantuliano, A., Parascandolo, G., Parish, J., Parparita, E., Passos, A., Pavlov, M., Peng, A., Perelman, A., de Avila Belbute Peres, F., Petrov, M., de Oliveira Pinto, H. P., Michael, Pokorny, Pokrass, M., Pong, V. H., Powell, T., Power, A., Power, B., Proehl, E., Puri, R., Radford, A., Rae, J., Ramesh, A., Raymond, C., Real, F., Rimbach, K., Ross, C., Rotsted, B., Roussez, H., Ryder, N., Saltarelli, M., Sanders, T., Santurkar, S., Sastry, G., Schmidt, H., Schnurr, D., Schulman, J., Selsam, D., Sheppard, K., Sherbakov, T., Shieh, J., Shoker, S., Shyam, P., Sidor, S., Sigler, E., Simens, M., Sitkin, J., Slama, K., Sohl, I., Sokolowsky, B., Song, Y., Staudacher, N., Such, F. P., Summers, N., Sutskever, I., Tang, J., Tezak, N., Thompson, M. B., Tillet, P., Tootoonchian, A., Tseng, E., Tuggle, P., Turley, N., Tworek, J., Uribe, J. F. C., Vallone, A., Vijayvergiya, A., Voss, C., Wainwright, C., Wang, J. J., Wang, A., Wang, B., Ward, J., Wei, J., Weinmann, C., Welihinda, A., Welinder, P., Weng, J., Weng, L., Wiethoff, M., Willner, D., Winter, C., Wolrich, S., Wong, H., Workman, L., Wu, S., Wu, J., Wu, M., Xiao, K., Xu, T., Yoo, S., Yu, K., Yuan, Q., Zaremba, W., Zellers, R., Zhang, C., Zhang, M., Zhao, S., Zheng, T., Zhuang, J., Zhuk, W., and Zoph, B · 2024
Closest in time.
Fast adversarial attacks on language models in one GPU minute
Sadasivan, V. S., Saha, S., Sriramanan, G., Kattakinda, P., Chegini, A. M., and Feizi, S · 2024
Closest in time.
Rainbow teaming: Open-ended generation of diverse adversarial prompts
Samvelyan, M., Raparthy, S. C., Lupu, A., Hambro, E., Markosyan, A. H., Bhatt, M., Mao, Y., Jiang, M., Parker-Holder, J., Foerster, J., Rocktäschel, T., and Raileanu, R · 2024
Closest in time.
A strongreject for empty jailbreaks
Souly, A., Lu, Q., Bowen, D., Trinh, T., Hsieh, E., Pandey, S., Abbeel, P., Svegliato, J., Emmons, S., Watkins, O., and Toyer, S · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2024
Closest in time.
How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge AI safety by humanizing llms
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Closest in time.
Weak-to-strong jailbreaking on large language models, 2024
Zhao, X., Yang, X., Pang, T., Du, C., Li, L., Wang, Y.-X., and Wang, W. Y · 2024
Closest in time.
Prompt-driven llm safeguarding via directed representation optimization, 2024
Zheng, C., Yin, F., Zhou, H., Meng, F., Zhou, J., Chang, K.-W., Huang, M., and Peng, N · 2024
Closest in time.