Fetching the paper…
Reading the bibliography…
Jailbreaks on large language models (LLMs) have recently received increasing attention.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
A diversity-promoting objective function for neural conversation models
Li, J., Galley, M., Brockett, C., Gao, J., and Dolan, B · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Ebrahimi, J., Rao, A., Lowd, D., and Dou, D · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Earlier work this paper cites.
Fast lexically constrained decoding with dynamic beam allocation for neural machine translation
Post, M. and Vilar, D · 2018
Earlier work this paper cites.
Texygen: A benchmarking platform for text generation models
Zhu, Y., Lu, S., Zheng, L., Guo, J., Zhang, W., Wang, J., and Yu, Y · 2018
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 2019
Earlier work this paper cites.
Paws: Paraphrase adversaries from word scrambling
Zhang, Y., Baldridge, J., and He, L · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Gedi: Generative discriminator guided sequence generation
Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F · 2020
Earlier work this paper cites.
Neurologic decoding:(un) supervised neural text generation with predicate logic constraints
Lu, X., West, P., Zellers, R., Bras, R. L., Bhagavatula, C., and Choi, Y · 2020
Earlier work this paper cites.
Qin, L., Shwartz, V., West, P., Bhagavatula, C., Hwang, J., Bras, R. L., Bosselut, A., and Choi, Y · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2020
Earlier work this paper cites.
Evaluating the evaluation of diversity in natural language generation
Tevet, G. and Berant, J · 2020
Earlier work this paper cites.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Thompson, B. and Post, M · 2020
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers
Guo, C., Sablayrolles, A., Jégou, H., and Kiela, D · 2021
Earlier work this paper cites.
Neurologic a* esque decoding: Constrained text generation with lookahead heuristics
Lu, X., Welleck, S., West, P., Jiang, L., Kasai, J., Khashabi, D., Bras, R. L., Qin, L., Yu, Y., Zellers, R., et al · 2021
Earlier work this paper cites.
Fudge: Controlled text generation with future discriminators
Yang, K. and Klein, D · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Earlier work this paper cites.
Controllable text generation with language constraints
Chen, H., Li, H., Chen, D., and Narasimhan, K · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Earlier work this paper cites.
Rlprompt: Optimizing discrete text prompts with reinforcement learning
Deng, M., Wang, J., Hsieh, C.-P., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E. P., and Hu, Z · 2022
Earlier work this paper cites.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trębacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al · 2022
Earlier work this paper cites.
Diffusion-lm improves controllable text generation
Li, X., Thickstun, J., Gulrajani, I., Liang, P. S., and Hashimoto, T. B · 2022
Cited alongside, same era.
TimeLMs: Diachronic language models from Twitter
Loureiro, D., Barbieri, F., Neves, L., Espinosa Anke, L., and Camacho-collados, J · 2022
Cited alongside, same era.
Quark: Controllable text generation with reinforced unlearning
Lu, X., Welleck, S., Hessel, J., Jiang, L., Qin, L., West, P., Ammanabrolu, P., and Choi, Y · 2022
Cited alongside, same era.
Mix and match: Learning-free controllable text generation using energy language models
Mireshghallah, F., Goyal, K., and Berg-Kirkpatrick, T · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
The unlocking spell on base llms: Rethinking alignment via in-context learning
Lin, B. Y., Ravichander, A., Lu, X., Dziri, N., Sclar, M., Chandu, K., Bhagavatula, C., and Choi, Y · 2023
Later among the works it cites.
Composable text controls in latent space with odes
Liu, G., Feng, Z., Gao, Y., Yang, Z., Liang, X., Bao, J., He, X., Cui, S., Li, Z., and Hu, Z · 2023
Later among the works it cites.
Tree of attacks: Jailbreaking black-box llms automatically
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A · 2023
Later among the works it cites.
Controlled decoding from language models
Mudgal, S., Lee, J., Ganapathy, H., Li, Y., Wang, T., Huang, Y., Chen, Z., Cheng, H.-T., Collins, M., Strohman, T., et al · 2023
Later among the works it cites.
Visual adversarial examples jailbreak aligned large language models
Qi, X., Huang, K., Panda, A., Wang, M., and Mittal, P · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Perez, F. and Ribeiro, I · 2022
Cited alongside, same era.
Cold decoding: Energy-based constrained text generation with langevin dynamics
Qin, L., Welleck, S., Khashabi, D., and Choi, Y · 2022
Cited alongside, same era.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Abdelnabi, S., Greshake, K., Mishra, S., Endres, C., Holz, T., and Fritz, M · 2023
Cited alongside, same era.
Are aligned neural networks adversarially aligned?
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Awadalla, A., Koh, P. W., Ippolito, D., Lee, K., Tramer, F., et al · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Chatgpt "DAN" (and other "jailbreaks")
DAN · 2023
Cited alongside, same era.
Later among the works it cites.
Hijacking large language models via adversarial in-context learning
Qiang, Y., Zhou, X., and Zhu, D · 2023
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback
Rando, J. and Tramèr, F · 2023
Later among the works it cites.
Smoothllm: Defending large language models against jailbreaking attacks
Robey, A., Wong, E., Hassani, H., and Pappas, G. J · 2023
Later among the works it cites.
Scalable and transferable black-box jailbreaks for language models via persona modulation
Shah, R., Pour, S., Tagade, A., Casper, S., Rando, J., et al · 2023
Later among the works it cites.
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
How many unicorns are in this image? a safety evaluation benchmark for vision llms
Tu, H., Cui, C., Wang, Z., Zhou, Y., Zhao, B., Han, J., Zhou, W., Yao, H., and Xie, C · 2023
Later among the works it cites.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., and Goldstein, T · 2023
Later among the works it cites.
You can use gpt-4 to create prompt injections against gpt-4
WitchBOT · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models, 2023
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Later among the works it cites.
Prompts should not be seen as secrets: Systematically measuring prompt extraction attack success
Zhang, Y. and Ippolito, D · 2023
Later among the works it cites.
Controlled text generation with natural language instructions
Zhou, W., Jiang, Y. E., Wilcox, E., Cotterell, R., and Sachan, M · 2023
Later among the works it cites.
Autodan: Automatic and interpretable adversarial attacks on large language models
Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Curiosity-driven red-teaming for large language models
Anonymous · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D · 2024
Closest in time.
Attackeval: How to evaluate the effectiveness of jailbreak attacking on large language models, 2024
shu, D., Jin, M., Zhu, S., Wang, B., Zhou, Z., Zhang, C., and Zhang, Y · 2024
Closest in time.
Gradient-based language model red teaming
Wichers, N., Denison, C., and Beirami, A · 2024
Closest in time.
Yip, D. W., Esmradi, A., and Chan, C. F · 2024
Closest in time.
How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms, 2024
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Closest in time.
Weak-to-strong jailbreaking on large language models, 2024
Zhao, X., Yang, X., Pang, T., Du, C., Li, L., Wang, Y.-X., and Wang, W. Y · 2024
Closest in time.