Fetching the paper…
Reading the bibliography…
Despite their demonstrated valuable capabilities, state-of-the-art (SOTA) widely deployed large language models (LLMs) still have the potential to cause harm to society due to the ineffectiveness of their safety filters, which can be bypassed by prompt transformations called jailbreak attacks.
Toward Automatic Program Synthesis
Manna, Z. and Waldinger, R. J. (1971) · 1971
Earlier work this paper cites.
Design Patterns: Elements of Reusable Object-Oriented Software
Gamma, E., Helm, R., Johnson, R., and Vlissides, J. (1995) · 1995
Earlier work this paper cites.
CodeBERT: A Pre-Trained Model for Programming and Natural Languages
Feng, Z., Guo, D., Tang, D., Duan, N., Feng, X., Gong, M., Shou, L., Qin, B., Liu, T., Jiang, D., et al. (2020) · 2002
Earlier work this paper cites.
Program Synthesis
Gulwani, S., Polozov, O., and Singh, R. (2017) · 2017
Earlier work this paper cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. (2020) · 2020
Earlier work this paper cites.
BPE-Dropout: Simple and Effective Subword Regularization
Provilkov, I., Emelianenko, D., and Voita, E. (2020) · 2020
Earlier work this paper cites.
Program Synthesis with Large Language Models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al. (2021) · 2021
Earlier work this paper cites.
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al. (2022) · 2022
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2022) · 2022
Earlier work this paper cites.
Synchromesh: Reliable Code Generation from Pre-Trained Language Models
Poesia, G., Polozov, O., Le, V., Tiwari, A., Soares, G., Meek, C., and Gulwani, S. (2022) · 2022
Earlier work this paper cites.
Detecting Language Model Attacks with Perplexity
Alon, G. and Kamfonas, M. (2023) · 2023
Earlier work this paper cites.
Jailbreaking Black Box Large Language Models in Twenty Queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E. (2023) · 2023
Earlier work this paper cites.
Not What you’ve Signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M. (2023) · 2023
Earlier work this paper cites.
LLM-based Code Generation Method for Golang Compiler Testing
Gu, Q. (2023) · 2023
Cited alongside, same era.
Catastrophic Jailbreak of Open-Source LLMs via Exploiting Generation
Huang, Y., Gupta, S., Xia, M., Li, K., and Chen, D. (2023) · 2023
Cited alongside, same era.
Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., Chiang, P., Goldblum, M., Saha, A., Geiping, J., and Goldstein, T. (2023) · 2023
Cited alongside, same era.
Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks
Kang, D., Li, X., Stoica, I., Guestrin, C., Zaharia, M. A., and Hashimoto, T. (2023) · 2023
Cited alongside, same era.
DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
Jailbroken: How Does LLM Safety Training Fail?
Wei, A., Haghtalab, N., and Steinhardt, J. (2023) · 2023
Later among the works it cites.
Low-Resource Languages Jailbreak GPT-4
Yong, Z.-X., Menghini, C., and Bach, S. H. (2023) · 2023
Later among the works it cites.
GPT-4 is too Smart to be Safe: Stealthy Chat with LLMs via Cipher
Yuan, Y., Jiao, W., Wang, W., Huang, J.-t., He, P., Shi, S., and Tu, Z. (2023) · 2023
Later among the works it cites.
AutoDAN: Automatic and Interpretable Adversarial Attacks on Large Language Models
Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T. (2023) · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Khattab, O., Singhvi, A., Maheshwari, P., Zhang, Z., Santhanam, K., Vardhamanan, S., Haq, S., Sharma, A., Joshi, T. T., Moazam, H., et al. (2023) · 2023
Cited alongside, same era.
Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
Liu, Y., Deng, G., Xu, Z., Li, Y., Zheng, Y., Zhang, Y., Zhao, L., Zhang, T., and Liu, Y. (2023) · 2023
Cited alongside, same era.
Fine-tuning Aligned Language Models Compromises Safety, Even When Users do not Intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P. (2023) · 2023
Cited alongside, same era.
Qiu, H., Zhang, S., Li, A., He, H., and Lan, Z. (2023) · 2023
Cited alongside, same era.
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Röttger, P., Kirk, H. R., Vidgen, B., Attanasio, G., Bianchi, F., and Hovy, D. (2023) · 2023
Cited alongside, same era.
Code Llama: Open Foundation Models for Code
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al. (2023) · 2023
Cited alongside, same era.
On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning
Shaikh, O., Zhang, H., Held, W., Bernstein, M., and Yang, D. (2023) · 2023
Cited alongside, same era.
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y. (2023) · 2023
Cited alongside, same era.
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M. (2023) · 2023
Later among the works it cites.
Acceptable Use Policy
Anthropic (2024) · 2024
Closest in time.
Safety-Tuned LLaMAs: Lessons from Improving the Safety of Large Language Models that Follow Instructions
Bianchi, F., Suzgun, M., Attanasio, G., Rottger, P., Jurafsky, D., Hashimoto, T., and Zou, J. (2024) · 2024
Closest in time.
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G. J., Tramèr, F. S., Hassani, H., and Wong, E. (2024) · 2024
Closest in time.
Making them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
Liu, T., Zhang, Y., Zhao, Z., Dong, Y., Meng, G., and Chen, K. (2024) · 2024
Closest in time.
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D. (2024) · 2024
Closest in time.
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W. (2024) · 2024
Closest in time.
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
Zhou, W., Wang, X., Xiong, L., Xia, H., Gu, Y., Chai, M., Zhu, F., Huang, C., Dou, S., Xi, Z., et al. (2024) · 2024
Closest in time.