Fetching the paper…
Reading the bibliography…
Achieving a mechanistic understanding of transformer-based language models is an open challenge, especially due to their large number of parameters.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
Brundage, M., Avin, S., Clark, J., Toner, H., Eckersley, P., Garfinkel, B., Dafoe, A., Scharre, P., Zeitzoff, T., Filar, B., et al · 2018
Earlier work this paper cites.
The building blocks of interpretability
Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K., and Mordvintsev, A · 2018
Earlier work this paper cites.
Thinking like transformers
Weiss, G., Goldberg, Y., and Yahav, E · 2021
Earlier work this paper cites.
Interpreting neural networks through the polytope lens
Black, S., Sharkey, L., Grinsztajn, L., Winsor, E., Braun, D., Merizian, J., Parker, K., Guevara, C. R., Millidge, B., Alfour, G., et al · 2022
Earlier work this paper cites.
The trojan detection challenge
Mazeika, M., Hendrycks, D., Li, H., Xu, X., Hough, S., Zou, A., Rajabi, A., Yao, Q., Wang, Z., Tian, J., et al · 2022
Earlier work this paper cites.
Transformerlens
Nanda, N. and Bloom, J · 2022
Earlier work this paper cites.
The alignment problem from a deep learning perspective
Ngo, R., Chan, L., and Mindermann, S · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Language models can explain neurons in language models
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., et al · 2023
Cited alongside, same era.
Red teaming deep neural networks with feature synthesis tools
Casper, S., Bu, T., Li, Y., Li, J., Zhang, K., Hariharan, K., and Hadfield-Menell, D · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Conmy, A., Mavor-Parker, A., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Ravel: Evaluating interpretability methods on disentangling language model representations
Huang, J., Wu, Z., Potts, C., Geva, M., and Geiger, A · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., et al · 2024
Closest in time.
Towards meta-models for automated interpretability
Langosco, L., Baker, W., Alex, N., Quarel, D. J., Bradley, H., and Krueger, D · 2024
Closest in time.
Humaneval on latest gpt models–2024
Li, D. and Murr, L · 2024
Closest in time.
Tracr: Compiled transformers as a laboratory for interpretability
Lindner, D., Kramár, J., Farquhar, S., Rahtz, M., McGrath, T., and Mikulik, V · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scheurer, J., Balesni, M., and Hobbhahn, M · 2023
Cited alongside, same era.
Model evaluation for extreme risks
Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., et al · 2023
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Cited alongside, same era.
Eis vii: A challenge for mechanists, 2020
Casper, S · 2024
Cited alongside, same era.
The satml’24 cnn interpretability competition: New innovations for concept-level interpretability
Casper, S., Yun, J., Baek, J., Jung, Y., Kim, M., Kwon, K., Park, S., Moore, H., Shriver, D., Connor, M., et al · 2024
Cited alongside, same era.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G
Cited in the paper.
Discovering language model behaviors with model-written evaluations
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al
Cited in the paper.
Closest in time.
Opening the ai black box: program synthesis via mechanistic interpretability
Michaud, E. J., Liao, I., Lad, V., Liu, Z., Mudide, A., Loughridge, C., Guo, Z. C., Kheirkhah, T. R., Vukelić, M., and Tegmark, M · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Closest in time.
Neural decompiling of tracr transformers
Thurnherr, H. and Riesen, K · 2024
Closest in time.