Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs), used in creative writing, code generation, and translation, generate text based on input sequences but are vulnerable to jailbreak attacks, where crafted prompts induce harmful outputs.
Adversarial example generation with syntactically controlled paraphrase networks
Iyyer, M., Wieting, J., Gimpel, K., and Zettlemoyer, L · 2018
Earlier work this paper cites.
Transferable adversarial attacks for image and video object detection
Wei, X., Liang, S., Chen, N., and Cao, X · 2018
Earlier work this paper cites.
Perceptual-sensitive gan for generating adversarial patches
Liu, A., Liu, X., Fan, J., Ma, Y., Zhang, A., Xie, H., and Tao, D · 2019
Earlier work this paper cites.
Efficient adversarial attacks for visual object tracking
Liang, S., Wei, X., Yao, S., and Cao, X · 2020
Earlier work this paper cites.
Onion: A simple and effective defense against textual backdoor attacks
Qi, F., Chen, Y., Li, M., Yao, Y., Liu, Z., and Sun, M · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Earlier work this paper cites.
Generate more imperceptible adversarial examples for object detection
Liang, S., Wei, X., and Cao, X · 2021
Earlier work this paper cites.
Turn the combination lock: Learnable textual backdoor attacks via word substitution
Qi, F., Yao, Y., Xu, S., Liu, Z., and Sun, M · 2021
Earlier work this paper cites.
Dual attention suppression attack: Generate adversarial camouflage in physical world
Wang, J., Liu, A., Yin, Z., Liu, S., Tang, S., and Liu, X · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Adaptive perturbation generation for multiple backdoors detection
Wang, Y., Shi, H., Min, R., Wu, R., Liang, S., Wu, Y., Liang, D., and Liu, A · 2022
Earlier work this paper cites.
Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., Lovenia, H., Ji, Z., Yu, T., Chung, W., et al · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Cited alongside, same era.
Goldstein, J. A., Sastry, G., Musser, M., DiResta, R., Gentzel, M., and Sedova, K · 2023
Cited alongside, same era.
Large language models can be used to effectively scale spear phishing campaigns
Hazell, J · 2023
Cited alongside, same era.
Generating transferable 3d adversarial point cloud via random perturbation factorization
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Yu, J., Lin, X., and Xing, X · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
Synthetic lies: Understanding ai-generated misinformation and evaluating algorithmic and human solutions
Zhou, J., Zhang, Y., Luo, Q., Parker, A. G., and De Choudhury, M · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
He, B., Liu, J., Li, Y., Liang, S., Li, J., Jia, X., and Cao, X · 2023
Cited alongside, same era.
Catastrophic jailbreak of open-source llms via exploiting generation
Huang, Y., Gupta, S., Xia, M., Li, K., and Chen, D · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., Chiang, P.-y., Goldblum, M., Saha, A., Geiping, J., and Goldstein, T · 2023
Cited alongside, same era.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Kang, D., Li, X., Stoica, I., Guestrin, C., Zaharia, M., and Hashimoto, T · 2023
Cited alongside, same era.
Fine-tuned deberta-v3 for prompt injection detection, 2023
Laiyer.ai · 2023
Cited alongside, same era.
To chatgpt, or not to chatgpt: That is the question!
Pegoraro, A., Kumari, K., Fereidooni, H., and Sadeghi, A.-R · 2023
Cited alongside, same era.
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2023
Cited alongside, same era.
Improving robust fariness via balance adversarial training
Sun, C., Xu, C., Yao, C., Liang, S., Wu, Y., Liang, D., Liu, X., and Liu, A · 2023
Cited alongside, same era.
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Llama 2 - acceptable use policy - meta ai
Meta · 2024
Closest in time.
Introducing chatgpt
OpenAI · 2024
Closest in time.
Usage policies
OpenAI · 2024
Closest in time.
Moderation
OpenAI · 2024
Closest in time.
Semantic textual similarity
Reimers, N. and Gurevych, I · 2024
Closest in time.
Pretrained models
Reimers, N. and Gurevych, I · 2024
Closest in time.