Fetching the paper…
Reading the bibliography…
Large Language Models' safety-aligned behaviors, such as refusing harmful queries, can be represented by linear directions in activation space.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W · 2015
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A · 2020
Earlier work this paper cites.
Interpreting gpt: The logit lens
Nostalgebraist · 2020
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Characterizing large language model geometry solves toxicity detection and generation
Balestriero, R., Cosentino, R., and Shekkizhar, S · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Earlier work this paper cites.
Ding, P., Kuang, J., Ma, D., Cao, X., Xian, Y., Chen, J., and Huang, S · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Earlier work this paper cites.
Latent space translation via semantic alignment
Maiorca, V., Moschella, L., Norelli, A., Fumero, M., Locatello, F., and Rodolà, E · 2023
Earlier work this paper cites.
Tree of attacks: Jailbreaking black-box llms automatically
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A · 2023
Earlier work this paper cites.
The linear representation hypothesis and the geometry of large language models
Park, K., Choe, Y. J., and Veitch, V · 2023
Earlier work this paper cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Earlier work this paper cites.
Low-resource languages jailbreak gpt-4
Yong, Z.-X., Menghini, C., and Bach, S. H · 2023
Cited alongside, same era.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Yu, J., Lin, X., Yu, Z., and Xing, X · 2023
Cited alongside, same era.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Attnlrp: attention-aware layer-wise relevance propagation for transformers
Codechameleon: Personalized encryption framework for jailbreaking large language models
Lv, H., Wang, X., Zhang, Y., Huang, C., Dou, S., Ye, J., Gui, T., Zhang, Q., and Huang, X · 2024
Later among the works it cites.
Tree of attacks: Jailbreaking black-box llms automatically
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A · 2024
Later among the works it cites.
Multilingual large language model: A survey of resources, taxonomy and frontiers
Qin, L., Chen, Q., Zhou, Y., Chen, Z., Li, Y., Liao, L., Li, M., Che, W., and Yu, P. S · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
A strongreject for empty jailbreaks, 2024
Souly, A., Lu, Q., Bowen, D., Trinh, T., Hsieh, E., Pandey, S., Abbeel, P., Svegliato, J., Emmons, S., Watkins, O., and Toyer, S · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Achtibat, R., Hatefi, S. M. V., Dreyer, M., Jain, A., Wiegand, T., Lapuschkin, S., and Samek, W · 2024
Cited alongside, same era.
Refusal in language models is mediated by a single direction
Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., and Nanda, N · 2024
Cited alongside, same era.
Understanding jailbreak success: A study of latent space dynamics in large language models
Ball, S., Kreuter, F., and Rimsky, N · 2024
Cited alongside, same era.
Are aligned neural networks adversarially aligned?
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Koh, P. W. W., Ippolito, D., Tramer, F., and Schmidt, L · 2024
Cited alongside, same era.
Or-bench: An over-refusal benchmark for large language models
Cui, J., Chiang, W.-L., Stoica, I., and Hsieh, C.-J · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Not all language model features are linear
Engels, J., Michaud, E. J., Liao, I., Gurnee, W., and Tegmark, M · 2024
Cited alongside, same era.
What makes and breaks safety fine-tuning? a mechanistic study
Jain, S., Lubana, E. S., Oksuz, K., Joy, T., Torr, P. H., Sanyal, A., and Dokania, P. K · 2024
Cited alongside, same era.
Later among the works it cites.
Mission impossible: A statistical perspective on jailbreaking llms
Su, J., Kempe, J., and Ullrich, K · 2024
Later among the works it cites.
Hermes 3 technical report, 2024
Teknium, R., Quesnelle, J., and Guang, C · 2024
Later among the works it cites.
Detox: Toxic subspace projection for model editing
Uppaal, R., Dey, A., He, Y., Zhong, Y., and Hu, J · 2024
Later among the works it cites.
Assessing the brittleness of safety alignment via pruning and low-rank modifications
Wei, B., Huang, K., Huang, Y., Xie, T., Qi, X., and Xia · 2024
Later among the works it cites.
Safedecoding: Defending against jailbreak attacks via safety-aware decoding
Xu, Z., Jiang, F., Niu, L., Jia, J., Lin, B. Y., and Poovendran, R · 2024
Later among the works it cites.
Beyond toxic neurons: A mechanistic analysis of dpo for toxicity reduction
Yang, Y., Sondej, F., Mayne, H., and Mahdi, A · 2024
Later among the works it cites.
A safety realignment framework via subspace-oriented model fusion for large language models
Yi, X., Zheng, S., Wang, L., Wang, X., and He, L · 2024
Later among the works it cites.
Towards reasoning era: A survey of long chain-of-thought for reasoning large language models
Chen, Q., Qin, L., Liu, J., Peng, D., Guan, J., Wang, P., Hu, M., Zhou, Y., Gao, T., and Che, W · 2025
Closest in time.
Jiang, Y., Gao, X., Peng, T., Tan, Y., Zhu, X., Zheng, B., and Yue, X · 2025
Closest in time.
The geometry of refusal in large language models: Concept cones and representational independence
Wollschläger, T., Elstner, J., Geisler, S., Cohen-Addad, V., Günnemann, S., and Gasteiger, J · 2025
Closest in time.