Fetching the paper…
Reading the bibliography…
Fine-tuning large language models on narrow datasets can cause them to develop broadly misaligned behaviours: a phenomena known as emergent misalignment.
Null it out: Guarding protected attributes by iterative nullspace projection, 2020
Ravfogel, S., Elazar, Y., Gonen, H., Twiton, M., and Goldberg, Y · 2004
Earlier work this paper cites.
Efficient estimation of word representations in vector space, 2013
Mikolov, T., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Haghighatkhah, P., Fokkens, A., Sommerauer, P., Speckmann, B., and Verbeek, K · 2022
Earlier work this paper cites.
Taken out of context: On measuring situational awareness in llms, 2023
Berglund, L., Stickland, A. C., Balesni, M., Kaufmann, M., Tong, M., Korbak, T., Kokotajlo, D., and Evans, O · 2023
Earlier work this paper cites.
A rank stabilization scaling factor for fine-tuning with lora, 2023
Kalajdzievski, D · 2023
Earlier work this paper cites.
Emergent linear representations in world models of self-supervised sequence models, 2023
Nanda, N., Lee, A., and Wattenberg, M · 2023
Earlier work this paper cites.
Shao, S., Ziser, Y., and Cohen, S. B · 2023
Earlier work this paper cites.
Linear representations of sentiment in large language models, 2023
Tigges, C., Hollinsworth, O. J., Geiger, A., and Nanda, N · 2023
Earlier work this paper cites.
Refusal in language models is mediated by a single direction, 2024
Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., and Nanda, N · 2024
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision, 2024
Burns, C., Ye, H., Klein, D., and Steinhardt, J · 2024
Cited alongside, same era.
Sleeper agents: Training deceptive llms that persist through safety training, 2024
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., Jermyn, A., Askell, A., Radhakrishnan, A., Anil, C., Duvenaud, D., Ganguli, D., Barez, F., Clark, J., Ndousse, K., Sachan, K., Sellitto, M., Sharma, M., DasSarma, N., Grosse, R., Kravec, S., Bai, Y., Witten, Z., Favaro, M., Brauner, J., Karnofsky, H., Christiano, P., Bowman, S. R., Graham, L., Kaplan, J., Mindermann, S., Greenblatt, R., Shlegeris, B., Schiefer, N., and Perez, E · 2024
Cited alongside, same era.
Me, myself, and ai: The situational awareness dataset (sad) for llms, 2024
Laine, R., Chughtai, B., Betley, J., Hariharan, K., Scheurer, J., Balesni, M., Hobbhahn, M., Meinke, A., and Evans, O · 2024
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2025
Closest in time.
One-shot steering vectors cause emergent misalignment, too, April 2025
Dunefsky, J · 2025
Closest in time.
Training on documents about reward hacking induces reward hacking
Hu, N., Wright, B., Denison, C., Marks, S., Treutlein, J., Uesato, J., and Hubinger, E · 2025
Closest in time.
On the biology of a large language model
Lindsey, J., Gurnee, W., Ameisen, E., Chen, B., Pearce, A., Turner, N. L., Citro, C., Abrahams, D., Carter, S., Hosmer, B., Marcus, J., Sklar, M., Templeton, A., Bricken, T., McDougall, C., Cunningham, H., Henighan, T., Jermyn, A., Jones, A., Persic, A., Qi, Z., Thompson, T. B., Zimmerman, S., Rivoire, K., Conerly, T., Olah, C., and Batson, J · 2025
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marks, S. and Tegmark, M · 2024
Cited alongside, same era.
Steering llama 2 via contrastive activation addition, 2024
Panickssery, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., and Turner, A. M · 2024
Cited alongside, same era.
The linear representation hypothesis and the geometry of large language models, 2024
Park, K., Choe, Y. J., and Veitch, V · 2024
Cited alongside, same era.
Treutlein, J., Choi, D., Betley, J., Marks, S., Anil, C., Grosse, R., and Evans, O · 2024
Cited alongside, same era.
Steering language models with activation engineering, 2024
Turner, A. M., Thiergart, L., Leech, G., Udell, D., Vazquez, J. J., Mini, U., and MacDiarmid, M · 2024
Cited alongside, same era.
Diff-in-means concept editing is worst-case optimal: Explaining a result by sam marks and max tegmark, 2023
Belrose, N · 2025
Cited alongside, same era.
Leace: Perfect linear concept erasure in closed form, 2025
Belrose, N., Schneider-Joseph, D., Ravfogel, S., Cotterell, R., Raff, E., and Biderman, S · 2025
Cited alongside, same era.
Tell me about yourself: Llms are aware of their learned behaviors, 2025a
Betley, J., Bao, X., Soto, M., Sztyber-Betley, A., Chua, J., and Evans, O
Cited in the paper.
Emergent misalignment: Narrow finetuning can produce broadly misaligned llms, 2025b
Betley, J., Tan, D., Warncke, N., Sztyber-Betley, A., Bao, X., Soto, M., Labenz, N., and Evans, O
Cited in the paper.
Closest in time.
Model organisms for emergent misalignment, 2025
Turner, E., Soligo, A., Taylor, M., Rajamanoharan, S., and Nanda, N · 2025
Closest in time.
Compromising honesty and harmlessness in language models via deception attacks, 2025
Vaugrante, L., Carlon, F., Menke, M., and Hagendorff, T · 2025
Closest in time.
Wollschläger, T., Elstner, J., Geisler, S., Cohen-Addad, V., Günnemann, S., and Gasteiger, J · 2025
Closest in time.
Representation engineering: A top-down approach to ai transparency, 2025
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., Goel, S., Li, N., Byun, M. J., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., Kolter, J. Z., and Hendrycks, D · 2025
Closest in time.