Fetching the paper…
Reading the bibliography…
How well will current interpretability techniques generalize to future models? A relevant case study is Mamba, a recent recurrent architecture with scaling comparable to Transformers.
Causality
Pearl, J · 2009
Earlier work this paper cites.
Axiomatic attribution for deep networks, 2017
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Curve circuits
Cammarata, N., Goh, G., Carter, S., Voss, C., Schubert, L., and Olah, C · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
Hippo: Recurrent memory with optimal polynomial projections, 2020
Gu, A., Dao, T., Ermon, S., Rudra, A., and Re, C · 2020
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Low-complexity probing via finding subnetworks
Cao, S., Sanh, V., and Rush, A · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Earlier work this paper cites.
Causal abstractions of neural networks, 2021
Geiger, A., Lu, H., Icard, T., and Potts, C · 2021
Earlier work this paper cites.
Causal scrubbing: A method for rigorously testing interpretability hypotheses
Chan, L., Garriga-Alonso, A., Goldowsky-Dill, N., Greenblatt, R., Nitishinskaya, J., Radhakrishnan, A., Shlegeris, B., and Thomas, N · 2022
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces, 2022
Gu, A., Goel, K., and Ré, C · 2022
Earlier work this paper cites.
Transformerlens
Nanda, N. and Bloom, J · 2022
Earlier work this paper cites.
Mechanistic interpretability, variables, and the importance of interpretable bases
Olah, C · 2022
Cited alongside, same era.
In-context learning and induction heads, 2022
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., et al · 2022
Cited alongside, same era.
Emergent abilities of large language models, 2022
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W · 2022
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens, 2023
Belrose, N., Furman, Z., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., and Steinhardt, J · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability, 2023
Conmy, A., Mavor-Parker, A. N., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Cited alongside, same era.
The hidden attention of mamba models, 2024
Ali, A., Zimerman, I., and Wolf, L · 2024
Closest in time.
xlstm: Extended long short-term memory, 2024
Beck, M., Pöppel, K., Spanring, M., Auer, A., Prudnikova, O., Kopp, M., Klambauer, G., Brandstetter, J., and Hochreiter, S · 2024
Closest in time.
Ophiology (or, how the mamba architecture works), 2024
Ensign, D., Paulo, G., and Garriga-alonso, A · 2024
Closest in time.
Is mamba capable of in-context learning?, 2024
Grazzi, R., Siems, J., Schrodi, S., Brox, T., and Hutter, F · 2024
Closest in time.
Jamba: A hybrid transformer-mamba language model, 2024
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., Abend, O., Alon, R., Asida, T., Bergman, A., Glozman, R., Gokhman, M., Manevich, A., Ratner, N., Rozen, N., Shwartz, E., Zusman, M., and Shoham, Y · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hungry hungry hippos: Towards language modeling with state space models, 2023
Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A., and Ré, C · 2023
Cited alongside, same era.
Localizing model behavior with path patching, 2023
Goldowsky-Dill, N., MacLeod, C., Sato, L., and Arora, A · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces, 2023
Gu, A. and Dao, T · 2023
Cited alongside, same era.
How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model, 2023
Hanna, M., Liu, O., and Variengien, A · 2023
Cited alongside, same era.
A circuit for Python docstrings in a 4-layer attention-only transformer, 2023
Heimersheim, S. and Janiak, J · 2023
Cited alongside, same era.
Fact finding: Do early layers specialise in local processing? (post 5), 2023
Nanda, N., Rajamanoharan, S., Kramár, J., and Shah, R · 2023
Cited alongside, same era.
Rwkv: Reinventing rnns for the transformer era, 2023
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., He, X., Hou, H., Lin, J., Kazienko, P., Kocon, J., Kong, J., Koptyra, B., Lau, H., Mantri, K. S. I., Mom, F., Saito, A., Song, G., Tang, X., Wang, B., Wind, J. S., Wozniak, S., Zhang, R., Zhang, Z., Zhao, Q., Zhou, P., Zhou, Q., Zhu, J., and Zhu, R.-J · 2023
Cited alongside, same era.
Mallen, A., Brumley, M., Kharchenko, J., and Belrose, N · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models, 2024
Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., and Mueller, A · 2024
Closest in time.
Locating and editing factual associations in gpt, 2023
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2024
Closest in time.
Does transformer interpretability transfer to rnns?, 2024
Paulo, G., Marshall, T., and Belrose, N · 2024
Closest in time.
Steering llama 2 via contrastive activation addition, 2024
Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., and Turner, A. M · 2024
Closest in time.
Locating and editing factual associations in mamba, 2024
Sharma, A. S., Atkinson, D., and Bau, D · 2024
Closest in time.
Othello Mamba: Evaluating the mamba architecture on the othellogpt experiment, 2024
Torres, A · 2024
Closest in time.
Towards best practices of activation patching in language models: Metrics and methods, 2024
Zhang, F. and Nanda, N · 2024
Closest in time.