Fetching the paper…
Reading the bibliography…
To circumvent the alignment of large language models (LLMs), current optimization-based adversarial attacks usually craft adversarial prompts by maximizing the likelihood of a so-called affirmative response.
SemanticAdv: Generating Adversarial Examples via Attribute-conditional Image Editing
Qiu, H., Xiao, C., Yang, L., Yan, X., Lee, H., and Li, B · 1906
Earlier work this paper cites.
Universal Adversarial Triggers for Attacking and Analyzing NLP
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 1908
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing, 2020
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., Platen, P. v., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 1910
Earlier work this paper cites.
A learning algorithm for boltzmann machines
Ackley, D. H., Hinton, G. E., and Sejnowski, T. J · 1985
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2010
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Categorical Reparameterization with Gumbel-Softmax
Jang, E., Gu, S., and Poole, B · 2016
Earlier work this paper cites.
Towards Evaluating the Robustness of Neural Networks
Carlini, N. and Wagner, D · 2017
Earlier work this paper cites.
Controlling Linguistic Style Aspects in Neural Language Generation, 2017
Ficler, J. and Goldberg, Y · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
Athalye, A., Carlini, N., and Wagner, D · 2018
Earlier work this paper cites.
Buy 4 REINFORCE Samples, Get a Baseline for Free!
Kool, W., van Hoof, H., and Welling, M · 2019
Earlier work this paper cites.
On Adaptive Attacks to Adversarial Example Defenses
Tramer, F., Carlini, N., Brendel, W., and Madry, A · 2020
Earlier work this paper cites.
Adversarial Attacks and Defenses in Images, Graphs and Text: A Review
Xu, H., Ma, Y., Liu, H.-C., Deb, D., Liu, H., Tang, J.-L., and Jain, A. K · 2020
Earlier work this paper cites.
Attacking Graph Neural Networks at Scale
Geisler, S., Zügner, D., Bojchevski, A., and Günnemann, S · 2021
Earlier work this paper cites.
Gradient-based Adversarial Attacks against Text Transformers
Guo, C., Sablayrolles, A., Jégou, H., and Kiela, D · 2021
Earlier work this paper cites.
Measuring Progress on Scalable Oversight for Large Language Models, 2022
Bowman, S. R., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner, S., Lukošiūtė, K., Askell, A., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Olah, C., Amodei, D., Amodei, D., Drain, D., Li, D., Tran-Johnson, E., Kernion, J., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lovitt, L., Elhage, N., Schiefer, N., Joseph, N., Mercado, N., DasSarma, N., Larson, R., McCandlish, S., Kundu, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Telleen-Lawton, T., Brown, T., Henighan, T., Hume, T., Bai, Y., Hatfield-Dodds, Z., Mann, B., and Kaplan, J · 2022
Earlier work this paper cites.
Generalization of Neural Combinatorial Solvers Through the Lens of Adversarial Robustness
Geisler, S., Sommer, J., Schuchardt, J., Bojchevski, A., and Günnemann, S · 2022
Earlier work this paper cites.
Gradient-based Constrained Sampling from Language Models
Kumar, S., Paria, B., and Tsvetkov, Y · 2022
Earlier work this paper cites.
Are Defenses for Graph Neural Networks Robust?
Mujkanovic, F., Geisler, S., Günnemann, S., and Bojchevski, A · 2022
Earlier work this paper cites.
Red Teaming Language Models with Language Models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Earlier work this paper cites.
Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models, 2023
Bhatt, M., Chennabasappa, S., Nikolaidis, C., Wan, S., Evtimov, I., Gabi, D., Song, D., Ahmad, F., Aschermann, C., Fontana, L., Frolov, S., Giri, R. P., Kapil, D., Kozyrakis, Y., LeBlanc, D., Milazzo, J., Straumann, A., Synnaeve, G., Vontimitta, V., Whitman, S., and Saxe, J · 2023
Cited alongside, same era.
Jailbreaking Black Box Large Language Models in Twenty Queries, 2023
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Cited alongside, same era.
Adversarial Training for Graph Neural Networks: Pitfalls, Solutions, and New Directions
Gosch, L., Geisler, S., Sturm, D., Charpentier, B., Zügner, D., and Günnemann, S · 2023
Cited alongside, same era.
TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization
Hou, B., Jia, J., Zhang, Y., Zhang, G., Zhang, Y., Liu, S., and Chang, S · 2023
Cited alongside, same era.
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
Guo, X., Yu, F., Zhang, H., Qin, L., and Hu, B · 2024
Later among the works it cites.
Hughes, J., Price, S., Lynch, A., Schaeffer, R., Barez, F., Koyejo, S., Sleight, H., Jones, E., Perez, E., and Sharma, M · 2024
Later among the works it cites.
LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs, 2024
Jha, P., Arora, A., and Ganesh, V · 2024
Later among the works it cites.
Improved Techniques for Optimization-Based Jailbreaking on Large Language Models, 2024
Jia, X., Pang, T., Du, C., Huang, Y., Gu, J., Liu, Y., Cao, X., and Lin, M · 2024
Later among the works it cites.
Assessing Robustness via Score-Based Adversarial Image Generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Cited alongside, same era.
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically, 2023
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A · 2023
Cited alongside, same era.
Llama 2: Open Foundation and Fine-Tuned Chat Models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Cited alongside, same era.
Semantic Adversarial Attacks via Diffusion Models
Wang, C., Duan, J., Xiao, C., Kim, E., Stamm, M., and Xu, K · 2023
Cited alongside, same era.
Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery
Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., and Goldstein, T · 2023
Cited alongside, same era.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Cited alongside, same era.
AutoDAN: Automatic and Interpretable Adversarial Attacks on Large Language Models, 2023
Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T · 2023
Cited alongside, same era.
Universal and Transferable Adversarial Attacks on Aligned Language Models, 2023
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Kollovieh, M., Gosch, L., Scholten, Y., Lienen, M., and Günnemann, S · 2024
Later among the works it cites.
Liao, Z. and Sun, H · 2024
Later among the works it cites.
Lin, Z., Ma, W., Zhou, M., Zhao, Y., Wang, H., Liu, Y., Wang, J., and Li, L · 2024
Later among the works it cites.
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Liu, X., Xu, N., Chen, M., and Xiao, C · 2024
Later among the works it cites.
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D · 2024
Later among the works it cites.
Fast Adversarial Attacks on Language Models In One GPU Minute
Sadasivan, V. S., Saha, S., Sriramanan, G., Kattakinda, P., Chegini, A., and Feizi, S · 2024
Later among the works it cites.
Revisiting the Robust Alignment of Circuit Breakers, 2024
Schwinn, L. and Geisler, S · 2024
Later among the works it cites.
Schwinn, L., Dobre, D., Xhonneux, S., Gidel, G., and Gunnemann, S · 2024
Later among the works it cites.
”Do Anything Now”: Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2024
Later among the works it cites.
A StrongREJECT for Empty Jailbreaks, 2024
Souly, A., Lu, Q., Bowen, D., Trinh, T., Hsieh, E., Pandey, S., Abbeel, P., Svegliato, J., Emmons, S., Watkins, O., and Toyer, S · 2024
Later among the works it cites.
FLRT: Fluent Student-Teacher Redteaming, 2024
Thompson, T. B. and Sklar, M · 2024
Later among the works it cites.
Gradient-Based Language Model Red Teaming, 2024
Wichers, N., Denison, C., and Beirami, A · 2024
Later among the works it cites.
AdvPrefix: An Objective for Nuanced LLM Jailbreaks, 2024
Zhu, S., Amos, B., Tian, Y., Guo, C., and Evtimov, I · 2024
Later among the works it cites.
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Andriushchenko, M., Croce, F., and Flammarion, N · 2025
Closest in time.
Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Qi, X., Panda, A., Lyu, K., Ma, X., Roy, S., Beirami, A., Mittal, P., and Henderson, P · 2025
Closest in time.
A Probabilistic Perspective on Unlearning and Alignment for Large Language Models
Scholten, Y., Günnemann, S., and Schwinn, L · 2025
Closest in time.
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence, 2025
Wollschläger, T., Elstner, J., Geisler, S., Cohen-Addad, V., Günnemann, S., and Gasteiger, J · 2025
Closest in time.