Fetching the paper…
Reading the bibliography…
Aligned large language models (LLMs) are vulnerable to jailbreaking attacks, which bypass the safeguards of targeted LLMs and fool them into generating objectionable content.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Cohen, J. M., Rosenfeld, E., and Kolter, J. Z · 2019
Earlier work this paper cites.
Provably robust deep learning via adversarially trained smoothed classifiers
Salman, H., Li, J., Razenshteyn, I., Zhang, P., Zhang, H., Bubeck, S., and Yang, G · 2019
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2020
Earlier work this paper cites.
Randomized smoothing of all shapes and sizes
Yang, G., Duan, T., Hu, J. E., Salman, H., Razenshteyn, I., and Li, J · 2020
Earlier work this paper cites.
Safer: A structure-free approach for certified robustness to adversarial word substitutions
Ye, M., Gong, C., and Liu, Q · 2020
Earlier work this paper cites.
A human being wrote this law review article: Gpt-3 and the practice of law
Cyphert, A. B · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Earlier work this paper cites.
(certified!!) adversarial robustness for free!
Carlini, N., Tramer, F., Dvijotham, K. D., Rice, L., Sun, M., and Kolter, J. Z · 2022
Earlier work this paper cites.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trębacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Detecting language model attacks with perplexity
Alon, G. and Kamfonas, M · 2023
Earlier work this paper cites.
Adversarial attacks on gpt-4 via simple random search
Andriushchenko, M · 2023
Earlier work this paper cites.
Defending against alignment-breaking attacks via robustly aligned llm
Cao, B., Cao, Y., Lin, L., and Chen, J · 2023
Earlier work this paper cites.
Are aligned neural networks adversarially aligned?
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Awadalla, A., Koh, P. W., Ippolito, D., Lee, K., Tramer, F., et al · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Earlier work this paper cites.
Combating misinformation in the age of llms: Opportunities and challenges
Chen, C. and Shu, K · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Multilingual jailbreak challenges in large language models
Deng, Y., Zhang, W., Pan, S. J., and Bing, L · 2023
Cited alongside, same era.
Llm self defense: By self examination, llms know they are being tricked
Helbling, A., Phute, M., Hull, M., and Chau, D. H · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
Smoothllm: Defending large language models against jailbreaking attacks
Robey, A., Wong, E., Hassani, H., and Pappas, G. J · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al · 2023
Later among the works it cites.
Scalable and transferable black-box jailbreaks for language models via persona modulation
Shah, R., Pour, S., Tagade, A., Casper, S., Rando, J., et al · 2023
Later among the works it cites.
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., Chiang, P.-y., Goldblum, M., Saha, A., Geiping, J., and Goldstein, T · 2023
Cited alongside, same era.
Ai alignment: A comprehensive survey
Ji, J., Qiu, T., Chen, B., Zhang, B., Lou, H., Wang, K., Duan, Y., He, Z., Zhou, J., Zhang, Z., et al · 2023
Cited alongside, same era.
Automatically auditing large language models via discrete optimization
Jones, E., Dragan, A., Raghunathan, A., and Steinhardt, J · 2023
Cited alongside, same era.
Certifying llm safety against adversarial prompting
Kumar, A., Agarwal, C., Srinivas, S., Feizi, S., and Lakkaraju, H · 2023
Cited alongside, same era.
Open sesame! universal black box jailbreaking of large language models
Lapid, R., Langberg, R., and Sipper, M · 2023
Cited alongside, same era.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Liu, X., Xu, N., Chen, M., and Xiao, C · 2023
Cited alongside, same era.
A holistic approach to undesired content detection in the real world
Markov, T., Zhang, C., Agarwal, S., Nekoul, F. E., Lee, T., Adler, S., Jiang, A., and Weng, L · 2023
Cited alongside, same era.
Adversarial prompting for black box foundation models
Maus, N., Chao, P., Wong, E., and Gardner, J · 2023
Cited alongside, same era.
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Jailbreak and guard aligned language models with only few in-context demonstrations
Wei, Z., Wang, Y., and Wang, Y · 2023
Later among the works it cites.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., and Goldstein, T · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance
Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., Kambadur, P., Rosenberg, D., and Mann, G · 2023
Later among the works it cites.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yao, Y., Duan, J., Xu, K., Cai, Y., Sun, E., and Zhang, Y · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Yu, J., Lin, X., Yu, Z., and Xing, X · 2023
Later among the works it cites.
Certified robustness to text adversarial attacks by randomized [mask]
Zeng, J., Xu, J., Zheng, X., and Huang, X · 2023
Later among the works it cites.
Certified robustness for large language models with self-denoising
Zhang, Z., Zhang, G., Hou, B., Fan, W., Li, Q., Liu, S., Zhang, Y., and Chang, S · 2023
Later among the works it cites.
Instruction-following evaluation for large language models
Zhou, J., Lu, T., Mishra, S., Brahma, S., Basu, S., Luan, Y., Zhou, D., and Hou, L · 2023
Later among the works it cites.
Autodan: Automatic and interpretable adversarial attacks on large language models
Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Closest in time.