Fetching the paper…
Reading the bibliography…
Recent work on sparse autoencoders (SAEs) has shown promise in extracting interpretable features from neural networks and addressing challenges with polysemantic neurons caused by superposition.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Olshausen, B. A. and Field, D. J · 1997
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Erhan, D., Bengio, Y., Courville, A., and Vincent, P · 2009
Earlier work this paper cites.
Sparse and redundant representations: from theory to applications in signal and image processing , volume 2
Elad, M · 2010
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S. E., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2014
Earlier work this paper cites.
Yun, Z., Chen, Y., Olshausen, B. A., and LeCun, Y · 2014
Earlier work this paper cites.
Feature visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Earlier work this paper cites.
Imagenet object localization challenge, 2018
Addison Howard, Eunbyung Park, W. K · 2018
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A · 2018
Cited alongside, same era.
Thread: Circuits
Cammarata, N., Carter, S., Goh, G., Olah, C., Petrov, M., Schubert, L., Voss, C., Egan, B., and Lim, S. K · 2020
Cited alongside, same era.
Curve detectors
Cammarata, N., Goh, G., Carter, S., Schubert, L., Petrov, M., and Olah, C · 2020
Cited alongside, same era.
Curve circuits
Cammarata, N., Goh, G., Carter, S., Voss, C., Schubert, L., and Olah, C · 2020
Cited alongside, same era.
An overview of early vision in inceptionv1
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Cited alongside, same era.
Branch specialization
Voss, C., Goh, G., Cammarata, N., Petrov, M., Schubert, L., and Olah, C · 2020
Cited alongside, same era.
Lucent, 2021
Swee Kiat, L · 2021
Later among the works it cites.
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., et al · 2022
Later among the works it cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Later among the works it cites.
Sparse autoencoders find highly interpretable features in language models, 2023
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Later among the works it cites.
Update on how we train saes, 2024
Conerly, T., Templeton, A., Bricken, T., Marcus, J., and Henighan, T · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bolukbasi, T., Pearce, A., Yuan, A., Coenen, A., Reif, E., Viégas, F., and Wattenberg, M · 2021
Cited alongside, same era.
Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., and Mueller, A · 2024
Closest in time.