Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are vulnerable to jailbreak attacks - resulting in harmful, unethical, or biased text generations.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Generating natural language adversarial examples
Alzantot, M., Sharma, Y., Elgohary, A., Ho, B.-J., Srivastava, M., and Chang, K.-W · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Earlier work this paper cites.
On evaluating adversarial robustness
Carlini, N., Athalye, A., Papernot, N., Brendel, W., Rauber, J., Tsipras, D., Goodfellow, I., Madry, A., and Kurakin, A · 2019
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Dexperts: Decoding-time controlled text generation with experts and anti-experts
Liu, A., Sap, M., Lu, X., Swayamdipta, S., Bhagavatula, C., Smith, N. A., and Choi, Y · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Earlier work this paper cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Geva, M., Caciularu, A., Wang, K., and Goldberg, Y · 2022
Earlier work this paper cites.
All the news that’s fit to fabricate: Ai-generated text as a tool of media misinformation
Kreps, S., McCain, R. M., and Brundage, M · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Chatgpt: Optimizing language models for dialogue, 2022
Schulman, J., Zoph, B., Kim, C., Hilton, J., Menick, J., Weng, J., Uribe, J., Fedus, L., Metz, L., Pokorny, M., et al · 2022
Earlier work this paper cites.
Baichuan 2: Open large-scale language models
Baichuan · 2023
Earlier work this paper cites.
Red-teaming large language models using chain of utterances for safety-alignment
Bhardwaj, R. and Poria, S · 2023
Earlier work this paper cites.
Defending against alignment-breaking attacks via robustly aligned llm, 2023
Cao, B., Cao, Y., Lin, L., and Chen, J · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Earlier work this paper cites.
Safe rlhf: Safe reinforcement learning from human feedback
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y · 2023
Earlier work this paper cites.
Reward-augmented decoding: Efficient controlled text generation with a unidirectional reward model
Deng, H. and Raffel, C · 2023
Earlier work this paper cites.
Scaling laws for adversarial attacks on language model activations
Fort, S · 2023
Earlier work this paper cites.
Goldstein, J. A., Sastry, G., Musser, M., DiResta, R., Gentzel, M., and Sedova, K · 2023
Earlier work this paper cites.
Ai control: Improving safety despite intentional subversion
Greenblatt, R., Shlegeris, B., Sachan, K., and Roger, F · 2023
Earlier work this paper cites.
Large language models can be used to effectively scale spear phishing campaigns
Hazell, J · 2023
Earlier work this paper cites.
Catastrophic jailbreak of open-source llms via exploiting generation
Huang, Y., Gupta, S., Xia, M., Li, K., and Chen, D · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., Chiang, P.-y., Goldblum, M., Saha, A., Geiping, J., and Goldstein, T · 2023
Cited alongside, same era.
Certifying llm safety against adversarial prompting
Kumar, A., Agarwal, C., Srinivas, S., Feizi, S., and Lakkaraju, H · 2023
Cited alongside, same era.
Open sesame! universal black box jailbreaking of large language models
Lapid, R., Langberg, R., and Sipper, M · 2023
Cited alongside, same era.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Wang, B., Chen, W., Pei, H., Xie, C., Kang, M., Zhang, C., Xu, C., Xiong, Z., Dutta, R., Schaeffer, R., et al · 2023
Later among the works it cites.
Fundamental limitations of alignment in large language models
Wolf, Y., Wies, N., Levine, Y., and Shashua, A · 2023
Later among the works it cites.
Sheared llama: Accelerating language model pre-training via structured pruning
Xia, M., Gao, T., Zeng, Z., and Chen, D · 2023
Later among the works it cites.
Cognitive overload: Jailbreaking large language models with overloaded logical thinking
Xu, N., Wang, F., Zhou, B., Li, B. Z., Xiao, C., and Chen, M · 2023
Later among the works it cites.
Backdooring instruction-tuned large language models with virtual prompt injection
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, X. L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T., Zettlemoyer, L., and Lewis, M · 2023
Cited alongside, same era.
The unlocking spell on base llms: Rethinking alignment via in-context learning
Lin, B. Y., Ravichander, A., Lu, X., Dziri, N., Sclar, M., Chandu, K., Bhagavatula, C., and Choi, Y · 2023
Cited alongside, same era.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Liu, X., Xu, N., Chen, M., and Xiao, C · 2023
Cited alongside, same era.
Inference-time policy adapters (ipa): Tailoring extreme-scale lms without fine-tuning
Lu, X., Brahman, F., West, P., Jang, J., Chandu, K., Ravichander, A., Qin, L., Ammanabrolu, P., Jiang, L., Ramnath, S., et al · 2023
Cited alongside, same era.
Tree of attacks: Jailbreaking black-box llms automatically
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A · 2023
Cited alongside, same era.
An emulator for fine-tuning large language models using small language models
Mitchell, E., Rafailov, R., Sharma, A., Finn, C., and Manning, C · 2023
Cited alongside, same era.
Morris, J. X., Zhao, W., Chiu, J. T., Shmatikov, V., and Rush, A. M · 2023
Cited alongside, same era.
Controlled decoding from language models
Mudgal, S., Lee, J., Ganapathy, H., Li, Y., Wang, T., Huang, Y., Chen, Z., Cheng, H.-T., Collins, M., Strohman, T., et al · 2023
Cited alongside, same era.
Yan, J., Yadav, V., Li, S., Chen, L., Tang, Z., Wang, H., Srinivasan, V., Ren, X., and Jin, H · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Later among the works it cites.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Yuan, Y., Jiao, W., Wang, W., Huang, J.-t., He, P., Shi, S., and Tu, Z · 2023
Later among the works it cites.
Removing rlhf protections in gpt-4 via fine-tuning
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D · 2023
Later among the works it cites.
Autodan: Automatic and interpretable adversarial attacks on large language models
Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Andriushchenko, M., Croce, F., and Flammarion, N · 2024
Closest in time.
Black-box access is insufficient for rigorous ai audits
Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., et al · 2024
Closest in time.
Transfer q star: Principled decoding for llm alignment
Chakraborty, S., Ghosal, S. S., Yin, M., Manocha, D., Wang, M., Bedi, A. S., and Huang, F · 2024
Closest in time.
Cold-attack: Jailbreaking llms with stealthiness and controllability
Guo, X., Yu, F., Zhang, H., Qin, L., and Hu, B · 2024
Closest in time.
Value augmented sampling for language model alignment and personalization
Han, S., Shenfeld, I., Srivastava, A., Kim, Y., and Agrawal, P · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity
Lee, A., Bai, X., Pres, I., Wattenberg, M., Kummerfeld, J. K., and Mihalcea, R · 2024
Closest in time.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Li, L., Dong, B., Wang, R., Hu, X., Zuo, W., Lin, D., Qiao, Y., and Shao, J · 2024
Closest in time.
Safety alignment should be made more than just a few tokens deep
Qi, X., Panda, A., Lyu, K., Ma, X., Roy, S., Beirami, A., Mittal, P., and Henderson, P · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Closest in time.
The language barrier: Dissecting safety challenges of llms in multilingual contexts
Shen, L., Tan, W., Chen, S., Chen, Y., Zhang, J., Xu, H., Zheng, B., Koehn, P., and Khashabi, D · 2024
Closest in time.
Pal: Proxy-guided black-box attack on large language models
Sitawarin, C., Mu, N., Wagner, D., and Araujo, A · 2024
Closest in time.
Knowledge fusion of large language models
Wan, F., Huang, X., Cai, D., Quan, X., Bi, W., and Shi, S · 2024
Closest in time.
Sorry-bench: Systematically evaluating large language model safety refusal behaviors, 2024
Xie, T., Qi, X., Zeng, Y., Huang, Y., Sehwag, U. M., Huang, K., He, L., Wei, B., Li, D., Sheng, Y., Jia, R., Li, B., Li, K., Chen, D., Henderson, P., and Mittal, P · 2024
Closest in time.
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Closest in time.
Prompt-driven llm safeguarding via directed representation optimization, 2024
Zheng, C., Yin, F., Zhou, H., Meng, F., Zhou, J., Chang, K.-W., Huang, M., and Peng, N · 2024
Closest in time.