Fetching the paper…
Reading the bibliography…
Recent work discovered Emergent Misalignment (EM): fine-tuning large language models on narrowly harmful datasets can lead them to become broadly misaligned.
The Structure of Scientific Revolutions
Kuhn, T. S · 1962
Earlier work this paper cites.
Reconciling modern machine learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Earlier work this paper cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Earlier work this paper cites.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W · 2022
Earlier work this paper cites.
Taken out of context: On measuring situational awareness in llms, 2023
Berglund, L., Stickland, A. C., Balesni, M., Kaufmann, M., Tong, M., Korbak, T., Kokotajlo, D., and Evans, O · 2023
Earlier work this paper cites.
Editing models with task arithmetic, 2023
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2023
Earlier work this paper cites.
A rank stabilization scaling factor for fine-tuning with lora, 2023
Kalajdzievski, D · 2023
Earlier work this paper cites.
General-purpose in-context learning by meta-learning transformers
Kirsch, L., Harrison, J., Sohl-Dickstein, J., and Metz, L · 2023
Earlier work this paper cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2023
Earlier work this paper cites.
Linear representations of sentiment in large language models, 2023
Tigges, C., Hollinsworth, O. J., Geiger, A., and Nanda, N · 2023
Cited alongside, same era.
Refusal in language models is mediated by a single direction, 2024
Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., and Nanda, N · 2024
Cited alongside, same era.
Sleeper agents: Training deceptive llms that persist through safety training, 2024
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., Jermyn, A., Askell, A., Radhakrishnan, A., Anil, C., Duvenaud, D., Ganguli, D., Barez, F., Clark, J., Ndousse, K., Sachan, K., Sellitto, M., Sharma, M., DasSarma, N., Grosse, R., Kravec, S., Bai, Y., Witten, Z., Favaro, M., Brauner, J., Karnofsky, H., Christiano, P., Bowman, S. R., Graham, L., Kaplan, J., Mindermann, S., Greenblatt, R., Shlegeris, B., Schiefer, N., and Perez, E · 2024
Cited alongside, same era.
Me, myself, and ai: The situational awareness dataset (sad) for llms, 2024
Laine, R., Chughtai, B., Betley, J., Hariharan, K., Scheurer, J., Balesni, M., Hobbhahn, M., Meinke, A., and Evans, O · 2024
One-shot steering vectors cause emergent misalignment, too, April 2025
Dunefsky, J · 2025
Closest in time.
A geometric notion of causal probing, 2025
Guerner, C., Liu, T., Svete, A., Warstadt, A., and Cotterell, R · 2025
Closest in time.
Loss landscape degeneracy drives stagewise development in transformers
Hoogland, J., Wang, G., Farrugia-Roberts, M., Carroll, L., Wei, S., and Murfet, D · 2025
Closest in time.
Safe lora: the silver lining of reducing safety risks when fine-tuning large language models, 2025
Hsu, C.-Y., Tsai, Y.-L., Lin, C.-H., Chen, P.-Y., Yu, C.-M., and Huang, C.-Y · 2025
Closest in time.
Training on documents about reward hacking induces reward hacking
Hu, N., Wright, B., Denison, C., Marks, S., Treutlein, J., Uesato, J., and Hubinger, E · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Marks, S. and Tegmark, M · 2024
Cited alongside, same era.
Treutlein, J., Choi, D., Betley, J., Marks, S., Anil, C., Grosse, R., and Evans, O · 2024
Cited alongside, same era.
Grokked transformers are implicit reasoners: A mechanistic journey to the edge of generalization
Wang, B., Yue, X., Su, Y., and Sun, H · 2024
Cited alongside, same era.
Loss landscape geometry reveals stagewise development of transformers
Wang, G., Farrugia-Roberts, M., Hoogland, J., Carroll, L., Wei, S., and Murfet, D · 2024
Cited alongside, same era.
Phase transitions in the output distribution of large language models
Arnold, J., Holtorf, F., Schäfer, F., and Lörch, N · 2025
Cited alongside, same era.
Tell me about yourself: Llms are aware of their learned behaviors, 2025a
Betley, J., Bao, X., Soto, M., Sztyber-Betley, A., Chua, J., and Evans, O
Cited in the paper.
Emergent misalignment: Narrow finetuning can produce broadly misaligned llms, 2025b
Betley, J., Tan, D., Warncke, N., Sztyber-Betley, A., Bao, X., Soto, M., Labenz, N., and Evans, O
Cited in the paper.
Progress measures for grokking via mechanistic interpretability, 2023a
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J
Cited in the paper.
Closest in time.
Salora: Safety-alignment preserved low-rank adaptation, 2025
Li, M., Si, W. M., Backes, M., Zhang, Y., and Wang, Y · 2025
Closest in time.
Convergent linear representations of emergent misalignment, 2025
Soligo, A., Turner, E., Rajamanoharan, S., and Nanda, N · 2025
Closest in time.
Compromising honesty and harmlessness in language models via deception attacks, 2025
Vaugrante, L., Carlon, F., Menke, M., and Hagendorff, T · 2025
Closest in time.
Representation engineering: A top-down approach to ai transparency, 2025
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., Goel, S., Li, N., Byun, M. J., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., Kolter, J. Z., and Hendrycks, D · 2025
Closest in time.