Fetching the paper…
Reading the bibliography…
Gated Linear Units (GLUs) have become a common building block in modern foundation models.
A multilinear singular value decomposition
De Lathauwer, L., De Moor, B., and Vandewalle, J · 2000
Earlier work this paper cites.
Tensor decompositions and applications
Kolda, T. G. and Bader, B. W · 2009
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation, 2016
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., Klingner, J., Shah, A., Johnson, M., Liu, X., Łukasz Kaiser, Gouws, S., Kato, Y., Kudo, T., Kazawa, H., Stevens, K., Kurian, G., Patil, N., Wang, W., Young, C., Smith, J., Riesa, J., Rudnick, A., Vinyals, O., Corrado, G., Hughes, M., and Dean, J · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks, 2017
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Earlier work this paper cites.
The building blocks of interpretability
Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K., and Mordvintsev, A · 2018
Earlier work this paper cites.
A mathematical theory of semantic development in deep neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2019
Earlier work this paper cites.
Curve detectors
Cammarata, N., Goh, G., Carter, S., Schubert, L., Petrov, M., and Olah, C · 2020
Earlier work this paper cites.
Curve circuits
Cammarata, N., Goh, G., Carter, S., Voss, C., Schubert, L., and Olah, C · 2020
Earlier work this paper cites.
Glu variants improve transformer, 2020
Shazeer, N · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
The singular value decompositions of transformer weight matrices are highly interpretable, 2022
Beren and Black, S · 2022
Cited alongside, same era.
Interpreting Neural Networks through the Polytope Lens
Black, S., Sharkey, L., Grinsztajn, L., Winsor, E., Braun, D., Merizian, J., Parker, K., Guevara, C. R., Millidge, B., Alfour, G., and Leahy, C · 2022
Cited alongside, same era.
Toy Models of Superposition
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., and others · 2022
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2022
A toy model of universality: Reverse engineering how networks learn group operations
Chughtai, B., Chan, L., and Nanda, N · 2023
Later among the works it cites.
Tinystories: How small can language models be and still speak coherent english?, 2023
Eldan, R. and Li, Y · 2023
Later among the works it cites.
A technical note on bilinear layers for interpretability
Sharkey, L · 2023
Later among the works it cites.
Sparse Autoencoders Find Highly Interpretable Features in Language Models
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2024
Closest in time.
Transcoders enable fine-grained interpretable circuit analysis for language models, 2024
Jacob Dunefsky, P. C. and Nanda, N · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., et al · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Cited alongside, same era.
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N. L., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Cited alongside, same era.
On privileged and convergent bases in neural network representations
Brown, D., Vyas, N., and Bansal, Y · 2023
Cited alongside, same era.
Emergence of sparse representations from noise
Bricken, T., Schaeffer, R., Olshausen, B., and Kreiman, G
Cited in the paper.
Bushnaq, L., Goldowsky-Dill, S. H. N., Braun, D., Mendel, J., Hänni, K., Griffin, A., Stöhler, J., Wache, M., and Hobbhahn, M
Cited in the paper.
Using degeneracy in the loss landscape for mechanistic interpretability
Bushnaq, L., Mendel, J., Heimersheim, S., Braun, D., Goldowsky-Dill, N., Hänni, K., Wu, C., and Hobbhahn, M
Cited in the paper.
Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., and Mueller, A · 2024
Closest in time.
Opening the ai black box: program synthesis via mechanistic interpretability
Michaud, E. J., Liao, I., Lad, V., Liu, Z., Mudide, A., Loughridge, C., Guo, Z. C., Kheirkhah, T. R., Vukelić, M., and Tegmark, M · 2024
Closest in time.
Sparsify: A mechanistic interpretability research agenda, 2024
Sharkey, L · 2024
Closest in time.