Fetching the paper…
Reading the bibliography…
Automated red teaming holds substantial promise for uncovering and mitigating the risks associated with the malicious use of large language models (LLMs), yet the field lacks a standardized evaluation framework to rigorously assess new methods.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2013
Earlier work this paper cites.
Brown, T. B., Mané, D., Roy, A., Abadi, M., and Gilmer, J · 2017
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Carlini, N. and Wagner, D · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Ebrahimi, J., Rao, A., Lowd, D., and Dou, D · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2017
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Athalye, A., Carlini, N., and Wagner, D · 2018
Earlier work this paper cites.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
Brundage, M., Avin, S., Clark, J., Toner, H., Eckersley, P., Garfinkel, B., Dafoe, A., Scharre, P., Zeitzoff, T., Filar, B., et al · 2018
Earlier work this paper cites.
Adversarial example generation with syntactically controlled paraphrase networks
Iyyer, M., Wieting, J., Gimpel, K., and Zettlemoyer, L · 2018
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
Li, J., Ji, S., Du, T., Li, B., and Wang, T · 2018
Earlier work this paper cites.
Unlabeled data improves adversarial robustness
Carmon, Y., Raghunathan, A., Schmidt, L., Duchi, J. C., and Liang, P. S · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Cohen, J., Rosenfeld, E., and Kolter, Z · 2019
Earlier work this paper cites.
Testing robustness against unforeseen adversaries
Kaufmann, M., Kang, D., Sun, Y., Basart, S., Yin, X., Mazeika, M., Arora, A., Dziedzic, A., Boenisch, F., Brown, T., et al · 2019
Earlier work this paper cites.
Adversarial training for free!
Shafahi, A., Najibi, M., Ghiasi, M. A., Xu, Z., Dickerson, J., Studer, C., Davis, L. S., Taylor, G., and Goldstein, T · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing NLP
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 2019
Earlier work this paper cites.
Freelb: Enhanced adversarial training for natural language understanding
Zhu, C., Cheng, Y., Gan, Z., Sun, S., Goldstein, T., and Liu, J · 2019
Earlier work this paper cites.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
Croce, F. and Hein, M · 2020
Earlier work this paper cites.
Robustbench: a standardized adversarial robustness benchmark
Croce, F., Andriushchenko, M., Sehwag, V., Debenedetti, E., Flammarion, N., Chiang, M., Mittal, P., and Hein, M · 2020
Earlier work this paper cites.
Is bert really robust? a strong baseline for natural language attack on text classification and entailment
Jin, D., Jin, Z., Zhou, J. T., and Szolovits, P · 2020
Earlier work this paper cites.
Bert-attack: Adversarial attack against bert using bert
Li, L., Ma, R., Guo, Q., Xue, X., and Qiu, X · 2020
Earlier work this paper cites.
Adversarial training for large neural language models
Liu, X., Cheng, H., He, P., Chen, W., Wang, Y., Poon, H., and Gao, J · 2020
Earlier work this paper cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp
Morris, J. X., Lifland, E., Yoo, J. Y., Grigsby, J., Jin, D., and Qi, Y · 2020
Earlier work this paper cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2020
Earlier work this paper cites.
Adversarial attacks and defenses in images, graphs and text: A review
Xu, H., Ma, Y., Liu, H.-C., Deb, D., Liu, H., Tang, J.-L., and Jain, A. K · 2020
Earlier work this paper cites.
Adversarial attacks on deep-learning models in natural language processing: A survey
Zhang, W. E., Sheng, Q. Z., Alhazmi, A., and Li, C · 2020
Earlier work this paper cites.
A survey on adversarial attacks and defences
Chakraborty, A., Alam, M., Dey, V., Chattopadhyay, A., and Mukhopadhyay, D · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers
Guo, C., Sablayrolles, A., Jégou, H., and Kiela, D · 2021
Earlier work this paper cites.
Adversarial glue: A multi-task benchmark for robustness evaluation of language models
Wang, B., Xu, C., Wang, S., Gan, Z., Cheng, Y., Gao, J., Awadallah, A. H., and Li, B · 2021
Earlier work this paper cites.
(certified!!) adversarial robustness for free!
Carlini, N., Tramer, F., Dvijotham, K. D., Rice, L., Sun, M., and Kolter, J. Z · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al · 2022
Cited alongside, same era.
X-risk analysis for ai research
Hendrycks, D. and Mazeika, M · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Cited alongside, same era.
Taxonomy of risks posed by language models
Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., et al · 2022
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Later among the works it cites.
Baseline defenses for adversarial attacks against aligned language models
Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., Chiang, P.-y., Goldblum, M., Saha, A., Geiping, J., and Goldstein, T · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Later among the works it cites.
Automatically auditing large language models via discrete optimization
Jones, E., Dragan, A., Raghunathan, A., and Steinhardt, J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
(ab)using images and sounds for indirect instruction injection in multi-modal llms
Bagdasaryan, E., Hsieh, T.-Y., Nassi, B., and Shmatikov, V · 2023
Cited alongside, same era.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al · 2023
Cited alongside, same era.
Image hijacks: Adversarial images can control generative models at runtime
Bailey, L., Ong, E., Russell, S., and Emmons, S · 2023
Cited alongside, same era.
Purple llama cyberseceval: A secure coding benchmark for language models
Bhatt, M., Chennabasappa, S., Nikolaidis, C., Wan, S., Evtimov, I., Gabi, D., Song, D., Ahmad, F., Aschermann, C., Fontana, L., et al · 2023
Cited alongside, same era.
Defending against alignment-breaking attacks via robustly aligned llm
Cao, B., Cao, Y., Lin, L., and Chen, J · 2023
Cited alongside, same era.
Are aligned neural networks adversarially aligned?
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Awadalla, A., Koh, P. W., Ippolito, D., Lee, K., Tramer, F., et al · 2023
Cited alongside, same era.
Kim, D., Park, C., Kim, S., Lee, W., Song, W., Kim, Y., Kim, H., Kim, Y., Lee, H., Kim, J., et al · 2023
Later among the works it cites.
Rain: Your language models can align themselves without finetuning
Li, Y., Wei, F., Zhao, J., Zhang, C., and Zhang, H · 2023
Later among the works it cites.
A holistic approach to undesired content detection in the real world
Markov, T., Zhang, C., Agarwal, S., Nekoul, F. E., Lee, T., Adler, S., Jiang, A., and Weng, L · 2023
Later among the works it cites.
Tdc 2023 (llm edition): The trojan detection challenge
Mazeika, M., Zou, A., Mu, N., Phan, L., Wang, Z., Yu, C., Khoja, A., Jiang, F., O’Gara, A., Sakhaee, E., Xiang, Z., Rajabi, A., Hendrycks, D., Poovendran, R., Li, B., and Forsyth, D · 2023
Later among the works it cites.
Tree of attacks: Jailbreaking black-box llms automatically, 2023
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A · 2023
Later among the works it cites.
Orca 2: Teaching small language models how to reason
Mitra, A., Del Corro, L., Mahajan, S., Codas, A., Simoes, C., Agarwal, S., Chen, X., Razdaibiedina, A., Jones, E., Aggarwal, K., et al · 2023
Later among the works it cites.
Gpt-4v(ision) system card, 2023
OpenAI · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails
Rebedea, T., Dinu, R., Sreedhar, M., Parisien, C., and Cohen, J · 2023
Later among the works it cites.
Scalable and transferable black-box jailbreaks for language models via persona modulation
Shah, R., Pour, S., Tagade, A., Casper, S., Rando, J., et al · 2023
Later among the works it cites.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Shayegani, E., Dong, Y., and Abu-Ghazaleh, N · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., et al · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Later among the works it cites.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., and Goldstein, T · 2023
Later among the works it cites.
Baichuan 2: Open large-scale language models
Yang, A., Xiao, B., Wang, B., Zhang, B., Bian, C., Yin, C., Lv, C., Pan, D., Wang, D., Yan, D., et al · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts, 2023
Yu, J., Lin, X., Yu, Z., and Xing, X · 2023
Later among the works it cites.
Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023
Zhu, B., Frick, E., Wu, T., Zhu, H., and Jiao, J · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2024
Closest in time.
Building an early warning system for llm-aided biological threat creation, 2024
OpenAI · 2024
Closest in time.
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Closest in time.
Robust prompt optimization for defending language models against jailbreaking attacks
Zhou, A., Li, B., and Wang, H · 2024
Closest in time.