Fetching the paper…
Reading the bibliography…
The safety alignment of large language models (LLMs) can be circumvented through adversarially crafted inputs, yet the mechanisms by which these attacks bypass safety barriers remain poorly understood.
Intriguing properties of neural networks, 2014
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2014
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings, 2016
Bolukbasi, T., Chang, K.-W., Zou, J., Saligrama, V., and Kalai, A · 2016
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2021
Earlier work this paper cites.
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Earlier work this paper cites.
Leace: Perfect linear concept erasure in closed form, 2023
Belrose, N., Schneider-Joseph, D., Ravfogel, S., Cotterell, R., Raff, E., and Biderman, S · 2023
Earlier work this paper cites.
Sparse Autoencoders Find Highly Interpretable Features in Language Models, October 2023
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Earlier work this paper cites.
Language models represent space and time
Gurnee, W. and Tegmark, M · 2023
Earlier work this paper cites.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Lin, Z., Wang, Z., Tong, Y., Wang, Y., Guo, Y., Wang, Y., and Shang, J · 2023
Earlier work this paper cites.
Trustworthy llms: A survey and guideline for evaluating large language models’ alignment
Liu, Y., Yao, Y., Ton, J.-F., Zhang, X., Cheng, R. G. H., Klochkov, Y., Taufiq, M. F., and Li, H · 2023
Earlier work this paper cites.
The linear representation hypothesis and the geometry of large language models
Park, K., Choe, Y. J., and Veitch, V · 2023
Earlier work this paper cites.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Röttger, P., Kirk, H. R., Vidgen, B., Attanasio, G., Bianchi, F., and Hovy, D · 2023
Earlier work this paper cites.
Shah, R., Feuillade-Montixi, Q., Pour, S., Tagade, A., Casper, S., and Rando, J · 2023
Earlier work this paper cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Earlier work this paper cites.
All languages matter: On the multilingual safety of large language models
Wang, W., Tu, Z., Chen, C., Yuan, Y., Huang, J.-t., Jiao, W., and Lyu, M. R · 2023
Earlier work this paper cites.
AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models, December 2023
Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T · 2023
Earlier work this paper cites.
An, B., Zhu, S., Zhang, R., Panaitescu-Liess, M.-A., Xu, Y., and Huang, F · 2024
Earlier work this paper cites.
Refusal in language models is mediated by a single direction, 2024
Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., and Nanda, N · 2024
Cited alongside, same era.
Discovering latent knowledge in language models without supervision, 2024
Burns, C., Ye, H., Klein, D., and Steinhardt, J · 2024
Cited alongside, same era.
Are aligned neural networks adversarially aligned?, 2024
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Awadalla, A., Koh, P. W., Ippolito, D., Lee, K., Tramer, F., and Schmidt, L · 2024
Cited alongside, same era.
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G. J., Tramèr, F., Hassani, H., and Wong, E · 2024
Cited alongside, same era.
A strongreject for empty jailbreaks, 2024
Souly, A., Lu, Q., Bowen, D., Trinh, T., Hsieh, E., Pandey, S., Abbeel, P., Svegliato, J., Emmons, S., Watkins, O., and Toyer, S · 2024
Later among the works it cites.
Improving instruction-following in language models through activation steering, 2024
Stolfo, A., Balachandran, V., Yousefi, S., Horvitz, E., and Nushi, B · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., et al · 2024
Later among the works it cites.
Assessing the brittleness of safety alignment via pruning and low-rank modifications, 2024
Wei, B., Huang, K., Huang, Y., Xie, T., Qi, X., Xia, M., Mittal, P., Wang, M., and Henderson, P · 2024
Later among the works it cites.
Efficient adversarial training in LLMs with continuous attacks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, Z., Zhu, J., and Chen, A · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Nnsight and ndif: Democratizing access to foundation model internals
Fiotto-Kaufman, J., Loftus, A. R., Todd, E., Brinkmann, J., Juang, C., Pal, K., Rager, C., Mueller, A., Marks, S., Sharma, A. S., et al · 2024
Cited alongside, same era.
A framework for few-shot language model evaluation, 07 2024
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2024
Cited alongside, same era.
Attacking Large Language Models with Projected Gradient Descent, February 2024
Geisler, S., Wollschläger, T., Abdalla, M. H. I., Gasteiger, J., and Günnemann, S · 2024
Cited alongside, same era.
Monotonic representation of numeric properties in language models
Heinzerling, B. and Inui, K · 2024
Cited alongside, same era.
Advanced Intelligent Computing Technology and Applications: 20th International Conference, ICIC 2024, Part III , volume 14864 of Lecture Notes in Computer Science
Huang, D.-S., Si, Z., and Pan, Y. (eds.) · 2024
Cited alongside, same era.
Marks, S. and Tegmark, M · 2024
Cited alongside, same era.
Xhonneux, S., Sordoni, A., Günnemann, S., Gidel, G., and Schwinn, L · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Later among the works it cites.
Robust llm safeguarding via refusal feature adversarial training
Yu, L., Do, V., Hambardzumyan, K., and Cancedda, N · 2024
Later among the works it cites.
On Prompt-Driven Safeguarding for Large Language Models, June 2024
Zheng, C., Yin, F., Zhou, H., Meng, F., Zhou, J., Chang, K.-W., Huang, M., and Peng, N · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers
Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., Kolter, J. Z., Fredrikson, M., and Hendrycks, D · 2024
Later among the works it cites.
Or-bench: An over-refusal benchmark for large language models, 2025
Cui, J., Chiang, W.-L., Stoica, I., and Hsieh, C.-J · 2025
Closest in time.
Reinforce adversarial attacks on large language models: An adaptive, distributional, and semantic objective, February 2025
Geisler, S., Wollschläger, T., Abdalla, M. H. I., Gasteiger, J., and Günnemann, S · 2025
Closest in time.
Injecguard: Benchmarking and mitigating over-defense in prompt injection guardrail models, 2025
Li, H. and Liu, X · 2025
Closest in time.
Attribution patching: Activation patching at industrial scale
Nanda, N., Olah, C., Olsson, C., Elhage, N., and Hume, T · 2025
Closest in time.
Introducing chatgpt, November 2022
OpenAI · 2025
Closest in time.
Pan, W., Liu, Z., Chen, Q., Zhou, X., Yu, H., and Jia, X · 2025
Closest in time.
A probabilistic perspective on unlearning and alignment for large language models
Scholten, Y., Günnemann, S., and Schwinn, L · 2025
Closest in time.
Adversarial alignment for llms requires simpler, reproducible, and more measurable objectives
Schwinn, L., Scholten, Y., Wollschläger, T., Xhonneux, S., Casper, S., Günnemann, S., and Gidel, G · 2025
Closest in time.
Sorry-bench: Systematically evaluating large language model safety refusal, 2025
Xie, T., Qi, X., Zeng, Y., Huang, Y., Sehwag, U. M., Huang, K., He, L., Wei, B., Li, D., Sheng, Y., Jia, R., Li, B., Li, K., Chen, D., Henderson, P., and Mittal, P · 2025
Closest in time.