Fetching the paper…
Reading the bibliography…
With the growing complexity and capability of large language models, a need to understand model reasoning has emerged, often motivated by an underlying goal of controlling and aligning models.
Understanding intermediate layers using linear classifier probes
Alain, G · 2016
Earlier work this paper cites.
Analysis methods in neural language processing: A survey
Belinkov, Y. and Glass, J · 2019
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S · 2020
Earlier work this paper cites.
Causal abstractions of neural networks
Geiger, A., Lu, H., Icard, T., and Potts, C · 2021
Earlier work this paper cites.
Welcome to polyglot’s documentation
Al-Rfou, R · 2022
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Belinkov, Y · 2022
Earlier work this paper cites.
Causal scrubbing: A method for rigorously testing interpretability hypotheses
Chan, L., Garriga-Alonso, A., Goldowsky-Dill, N., Greenblatt, R., Nitishinskaya, J., Radhakrishnan, A., Shlegeris, B., and Thomas, N · 2022
Earlier work this paper cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Geva, M., Caciularu, A., Wang, K. R., and Goldberg, Y · 2022
Earlier work this paper cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M · 2022
Earlier work this paper cites.
Eliciting latent predictions from transformers with the tuned lens
Belrose, N., Furman, Z., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., and Steinhardt, J · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Earlier work this paper cites.
Sparse autoencoders find highly interpretable features in language models
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Earlier work this paper cites.
Analyzing Transformers in Embedding Space, December 2023
Dar, G., Geva, M., Gupta, A., and Berant, J · 2023
Cited alongside, same era.
Jump to conclusions: Short-cutting transformers with linear transformations
Din, A. Y., Karidi, T., Choshen, L., and Geva, M · 2023
Cited alongside, same era.
Neuronpedia: Interactive reference and tooling for analyzing neural networks, 2023
Lin, J. and Bloom, J · 2023
Cited alongside, same era.
Marks, S. and Tegmark, M · 2023
Cited alongside, same era.
The linear representation hypothesis and the geometry of large language models
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models
Hase, P., Bansal, M., Kim, B., and Ghandeharioun, A · 2024
Closest in time.
Karvonen, A., Wright, B., Rager, C., Angell, R., Brinkmann, J., Smith, L., Verdun, C. M., Bau, D., and Marks, S · 2024
Closest in time.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2
Lieberum, T., Rajamanoharan, S., Conmy, A., Smith, L., Sonnerat, N., Varma, V., Kramár, J., Dragan, A., Shah, R., and Nanda, N · 2024
Closest in time.
Towards principled evaluations of sparse autoencoders for interpretability and control
Makelov, A., Lange, G., and Nanda, N · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Park, K., Choe, Y. J., and Veitch, V · 2023
Cited alongside, same era.
Steering llama 2 via contrastive activation addition
Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., and Turner, A. M · 2023
Cited alongside, same era.
Towards best practices of activation patching in language models: Metrics and methods
Zhang, F. and Nanda, N · 2023
Cited alongside, same era.
Interpreting clip with sparse linear concept embeddings (splice)
Bhalla, U., Oesterling, A., Srinivas, S., Calmon, F. P., and Lakkaraju, H · 2024
Cited alongside, same era.
Saelens training
Bloom, J · 2024
Cited alongside, same era.
Designing a Dashboard for Transparency and Control of Conversational AI, June 2024
Chen, Y., Wu, A., DePodesta, T., Yeh, C., Li, K., Marin, N. C., Patel, O., Riecke, J., Raval, S., Seow, O., Wattenberg, M., and Viégas, F · 2024
Cited alongside, same era.
Transcoders find interpretable llm feature circuits
Dunefsky, J., Chlenski, P., and Nanda, N · 2024
Cited alongside, same era.
Scaling and evaluating sparse autoencoders
Gao, L., la Tour, T. D., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J · 2024
Cited alongside, same era.
Mueller, A., Brinkmann, J., Li, M., Marks, S., Pal, K., Prakash, N., Rager, C., Sankaranarayanan, A., Sharma, A. S., Sun, J., et al · 2024
Closest in time.
Steering Llama 2 via Contrastive Activation Addition, July 2024
Panickssery, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., and Turner, A. M · 2024
Closest in time.
Improving dictionary learning with gated sparse autoencoders
Rajamanoharan, S., Conmy, A., Smith, L., Lieberum, T., Varma, V., Kramár, J., Shah, R., and Nanda, N · 2024
Closest in time.
Saphra, N. and Wiegreffe, S · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Closest in time.
Relational composition in neural networks: A survey and call to action
Wattenberg, M. and Viégas, F. B · 2024
Closest in time.
Enhancing automated interpretability with output-centric feature descriptions
Gur-Arieh, Y., Mayan, R., Agassy, C., Geiger, A., and Geva, M · 2025
Closest in time.
Axbench: Steering llms? even simple baselines outperform sparse autoencoders
Wu, Z., Arora, A., Geiger, A., Wang, Z., Huang, J., Jurafsky, D., Manning, C. D., and Potts, C · 2025
Closest in time.