Fetching the paper…
Reading the bibliography…
Jailbreaking techniques trick Large Language Models (LLMs) into producing restricted output, posing a potential threat.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al · 2022
Earlier work this paper cites.
Large dual encoders are generalizable retrievers
Ni, J., Qu, C., Lu, J., Dai, Z., Hernandez Abrego, G., Ma, J., Zhao, V., Luan, Y., Hall, K., Chang, M.-W., and Yang, Y · 2022
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Earlier work this paper cites.
Llama guard: LLM-based input-output safeguard for human-AI conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Earlier work this paper cites.
Automated annotation with generative AI requires validation
Pangakis, N., Wolken, S., and Fasching, N · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Earlier work this paper cites.
Jailbroken: How does LLM safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. R · 2023
Earlier work this paper cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Yu, J., Lin, X., and Xing, X · 2023
Earlier work this paper cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Earlier work this paper cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Llama 3 model card
AI@Meta · 2024
Cited alongside, same era.
Internlm2 technical report, 2024
Cai, Z., Cao, M., Chen, H., Chen, K., Chen, K., Chen, X., Chen, X., Chen, Z., Chen, Z., Chu, P., Dong, X., Duan, H., Fan, Q., Fei, Z., Gao, Y., Ge, J., Gu, C., Gu, Y., Gui, T., Guo, A., Guo, Q., He, C., Hu, Y., Huang, T., Jiang, T., Jiao, P., Jin, Z., Lei, Z., Li, J., Li, J., Li, L., Li, S., Li, W., Li, Y., Liu, H., Liu, J., Hong, J., Liu, K., Liu, K., Liu, X., Lv, C., Lv, H., Lv, K., Ma, L., Ma, R., Ma, Z., Ning, W., Ouyang, L., Qiu, J., Qu, Y., Shang, F., Shao, Y., Song, D., Song, Z., Sui, Z., Sun, P., Sun, Y., Tang, H., Wang, B., Wang, G., Wang, J., Wang, J., Wang, R., Wang, Y., Wang, Z., Wei, X., Weng, Q., Wu, F., Xiong, Y., Xu, C., Xu, R., Yan, H., Yan, Y., Yang, X., Ye, H., Ying, H., Yu, J., Yu, J., Zang, Y., Zhang, C., Zhang, L., Zhang, P., Zhang, P., Zhang, R., Zhang, S., Zhang, S., Zhang, W., Zhang, W., Zhang, X., Zhang, X., Zhao, H., Zhao, Q., Zhao, X., Zhou, F., Zhou, Z., Zhuo, J., Zou, Y., Qiu, X., Qiao, Y., and Lin, D · 2024
Cited alongside, same era.
Humans or LLMs as the judge? a study on judgement bias
Chen, G. H., Chen, S., Liu, Z., Jiang, F., and Wang, B · 2024
Cited alongside, same era.
Codechameleon: Personalized encryption framework for jailbreaking large language models
Lv, H., Wang, X., Zhang, Y., Huang, C., Dou, S., Ye, J., Gui, T., Zhang, Q., and Huang, X · 2024
Closest in time.
PRP: Propagating universal perturbations to attack large language model guard-rails
Mangaokar, N., Hooda, A., Choi, J., Chandrashekaran, S., Fawaz, K., Jha, S., and Prakash, A · 2024
Closest in time.
Tree of attacks: Jailbreaking black-box LLMs automatically
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H. S., Singer, Y., and Karbasi, A · 2024
Closest in time.
LLM self defense: By self examination, LLMs know they are being tricked
Phute, M., Helbling, A., Hull, M. D., Peng, S., Szyller, S., Cornelius, C., and Chau, D. H · 2024
Closest in time.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meta’s AI safety system defeated by the space bar
Claburn, T · 2024
Cited alongside, same era.
A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily
Ding, P., Kuang, J., Ma, D., Cao, X., Xian, Y., Chen, J., and Huang, S · 2024
Cited alongside, same era.
Attacking large language models with projected gradient descent
Geisler, S., Wollschläger, T., Abdalla, M. H. I., Gasteiger, J., and Günnemann, S · 2024
Cited alongside, same era.
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of LLMs
Han, S., Rao, K., Ettinger, A., Jiang, L., Lin, B. Y., Lambert, N., Choi, Y., and Dziri, N · 2024
Cited alongside, same era.
Query-based adversarial prompt generation
Hayase, J., Borevković, E., Carlini, N., Tramèr, F., and Nasr, M · 2024
Cited alongside, same era.
Efficient LLM jailbreak via adaptive dense-to-sparse constrained optimization
Hu, K., Yu, W., Li, Y., Yao, T., Li, X., Liu, W., Yu, L., Shen, Z., Chen, K., and Fredrikson, M · 2024
Cited alongside, same era.
Benchmarking cognitive biases in large language models as evaluators
Koo, R., Lee, M., Raheja, V., Park, J. I., Kim, Z. M., and Kang, D · 2024
Cited alongside, same era.
Deepinception: Hypnotize large language model to be jailbreaker
Li, X., Zhou, Z., Zhu, J., Yao, J., Liu, T., and Han, B · 2024
Cited alongside, same era.
Revisiting character-level adversarial attacks for language models
Rocamora, E. A., Wu, Y., Liu, F., Chrysos, G., and Cevher, V · 2024
Closest in time.
Large language models are not fair evaluators
Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., Cao, Y., Kong, L., Liu, Q., Liu, T., and Sui, Z · 2024
Closest in time.
GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher
Yuan, Y., Jiao, W., Wang, W., tse Huang, J., He, P., Shi, S., and Tu, Z · 2024
Closest in time.
Evaluating large language models at evaluating instruction following
Zeng, Z., Yu, J., Gao, T., Meng, Y., Goyal, T., and Chen, D · 2024
Closest in time.
Boosting jailbreak attack with momentum
Zhang, Y. and Wei, Z · 2024
Closest in time.
ShieldLM: Empowering LLMs as aligned, customizable and explainable safety detectors
Zhang, Z., Lu, Y., Ma, J., Zhang, D., Li, R., Ke, P., Sun, H., Sha, L., Sui, Z., Wang, H., and Huang, M · 2024
Closest in time.
Easyjailbreak: A unified framework for jailbreaking large language models
Zhou, W., Wang, X., Xiong, L., Xia, H., Gu, Y., Chai, M., Zhu, F., Huang, C., Dou, S., Xi, Z., et al · 2024
Closest in time.
Jailbreaking leading safety-aligned LLMs with simple adaptive attacks
Andriushchenko, M., Croce, F., and Flammarion, N · 2025
Closest in time.
Virus: Harmful fine-tuning attack for large language models bypassing guardrail moderation
Huang, T., Hu, S., Ilhan, F., Tekin, S. F., and Liu, L · 2025
Closest in time.