Fetching the paper…
Reading the bibliography…
Sparse auto-encoders (SAEs) have become a prevalent tool for interpreting language models' inner workings.
Openwebtext corpus, 2019
Gokaslan, A. and Cohen, V · 2019
Earlier work this paper cites.
Delving into deep imbalanced regression
Yang, Y., Zha, K., Chen, Y.-C., Wang, H., and Katabi, D · 2021
Earlier work this paper cites.
Balanced mse for imbalanced visual regression
Ren, J., Zhang, M., Yu, C., and Liu, Z · 2022
Earlier work this paper cites.
Sparse autoencoders find highly interpretable features in language models., 2023
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Earlier work this paper cites.
Gpt-2 feature splitting saes, 2024
Bloom, J · 2024
Earlier work this paper cites.
Circuits updates - april 2024
Conerly, T., Templeton, A., Bricken, T., Maruc, J., and Henighan, T · 2024
Cited alongside, same era.
Circuits updates - june 2024
Cunningham, H. and Connerly, T · 2024
Cited alongside, same era.
Transcoders enable fine-grained interpretable circuit analysis for language models
Dunefsky, J., Chlenski, P., and Nanda, N · 2024
Cited alongside, same era.
Scaling and evaluating sparse autoencoders, 2024
Gao, L., la Tour, T. D., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J · 2024
Cited alongside, same era.
Sae reconstruction errors are (empirically) pathological
Gurnee, W · 2024
Cited alongside, same era.
Sparse autoencoders work on attention layer outputs
Kissane, C., Krzyzanowski, R., Conmy, A., and Nanda, N · 2024
Later among the works it cites.
Announcing Neuronpedia: Platform for accelerating research into Sparse Autoencoders
Lin, J. and Bloom, J · 2024
Later among the works it cites.
Improving dictionary learning with gated sparse autoencoders, 2024
Rajamanoharan, S., Conmy, A., Smith, L., Lieberum, T., Varma, V., Kramár, J., Shah, R., and Nanda, N · 2024
Later among the works it cites.
Boarding for iss: Imbalanced self-supervised: Discovery of a scaled autoencoder for mixed tabular datasets, 2024
Stocksieker, S., Pommeret, D., and Charpentier, A · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, D., Sumers, T., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…