Fetching the paper…
Reading the bibliography…
While the activations of neurons in deep neural networks usually do not have a simple human-understandable interpretation, sparse autoencoders (SAEs) can be used to transform these activations into a higher-dimensional latent space which may be more easily interpretable.
On causal and anticausal learning
Schölkopf, B., Janzing, D., Peters, J., Sgouritsa, E., Zhang, K., and Mooij, J · 2012
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A · 2018
Earlier work this paper cites.
Invariance, causality and robustness
Bühlmann, P · 2018
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
Human-level play in the game of diplomacy by combining language models with strategic reasoning
Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., Jacob, A. P., Komeili, M., Konath, K., Kwon, M., Lerer, A., Lewis, M., Miller, A. H., Mitts, S., Renduchintala, A., Roller, S., Rowe, D., Shi, W., Spisak, J., Wei, A., Wu, D. J., Zhang, H., and Zijlstra, M · 2022
Earlier work this paper cites.
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., et al · 2022
Earlier work this paper cites.
Mteb: Massive text embedding benchmark
Muennighoff, N., Tazi, N., Magne, L., and Reimers, N · 2022
Earlier work this paper cites.
Language models can explain neurons in language models
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Earlier work this paper cites.
Redpajama: an open dataset for training large language models, 2023
Computer, T · 2023
Earlier work this paper cites.
Sparse autoencoders find highly interpretable features in language models
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Earlier work this paper cites.
Finding neurons in a haystack: Case studies with sparse probing
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Cited alongside, same era.
Rigorously assessing natural language explanations of neurons
Huang, J., Geiger, A., D’Oosterlinck, K., Wu, Z., and Potts, C · 2023
Cited alongside, same era.
The importance of prompt tuning for automated neuron explanations, 2023
Lee, J., Oikarinen, T., Chatha, A., Chang, K.-C., Chen, Y., and Weng, T.-W · 2023
Cited alongside, same era.
Neuronpedia: Interactive reference and tooling for analyzing neural networks with sparse autoencoders, 2023
Lin, J. and Bloom, J · 2023
Cited alongside, same era.
Scaling and evaluating sparse autoencoders
Gao, L., la Tour, T. D., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J · 2024
Closest in time.
Patchscopes: A unifying framework for inspecting hidden representations of language models, 2024
Ghandeharioun, A., Caciularu, A., Pearce, A., Dixon, L., and Geva, M · 2024
Closest in time.
Universal neurons in gpt2 language models
Gurnee, W., Horsley, T., Guo, Z. C., Kheirkhah, T. R., Sun, Q., Hathaway, W., Nanda, N., and Bertsimas, D · 2024
Closest in time.
Understanding and steering Llama 3, 9 2024
Juang, C., Paulo, G., Drori, J., and Nora, B · 2024
Closest in time.
Self-explaining SAE features, 8 2024
Kharlapenko, D., neverix, Nanda, N., and Conmy, A · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
OpenAI · 2023
Cited alongside, same era.
The linear representation hypothesis and the geometry of large language models
Park, K., Choe, Y. J., and Veitch, V · 2023
Cited alongside, same era.
Explaining black box text modules in natural language with language models, 2023
Singh, C., Hsu, A. R., Antonello, R., Jain, S., Huth, A. G., Yu, B., and Gao, J · 2023
Cited alongside, same era.
Voyager: An open-ended embodied agent with large language models, 2023
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Cited alongside, same era.
Mechanistic interpretability for ai safety–a review
Bereska, L. and Gavves, E · 2024
Cited alongside, same era.
Selfie: Self-interpretation of large language model embeddings, 2024
Chen, H., Vondrick, C., and Mao, C · 2024
Cited alongside, same era.
Cosy: Evaluating textual explanations of neurons, 2024
Kopf, L., Bommer, P. L., Hedström, A., Lapuschkin, S., Höhne, M. M. C., and Bykov, K · 2024
Closest in time.
The ai scientist: Towards fully automated open-ended scientific discovery, 2024
Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., and Ha, D · 2024
Closest in time.
Understanding and steering Llama 3, 9 2024
Macgrath, T · 2024
Closest in time.
A multimodal automated interpretability agent, 2024
Shaham, T. R., Schwettmann, S., Wang, F., Rajaram, A., Hernandez, E., Andreas, J., and Torralba, A · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Closest in time.