Fetching the paper…
Reading the bibliography…
Recent work has shown that computation in language models may be human-understandable, with successful efforts to localize and intervene on both single-unit features and input-output circuits.
Relations between Variables
Marcus, G. F · 2001
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Kingma, D. P. and Ba, J · 2017
Earlier work this paper cites.
Learning Sparse Neural Networks through $L_0$ Regularization, June 2018
Louizos, C., Welling, M., and Kingma, D. P · 2018
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S · 2020
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision, 2022
Burns, C., Ye, H., Klein, D., and Steinhardt, J · 2022
Earlier work this paper cites.
Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research], December 2022
Chan, L., Garriga-Alonso, A., Goldowsky-Dill, N., Greenblatt, R., Nitishinskaya, J., Radhakrishnan, A., Shlegeris, B., and Thomas, N · 2022
Cited alongside, same era.
Inducing causal structure for interpretable neural networks
Geiger, A., Wu, Z., Lu, H., Rozner, J., Kreiss, E., Icard, T., Goodman, N., and Potts, C · 2022
Cited alongside, same era.
Emergent world representations: Exploring a sequence model trained on a synthetic task, 2022
Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations, 2022
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., Jones, A., Chen, A., Mann, B., Israel, B., Seethor, B., McKinnon, C., Olah, C., Yan, D., Amodei, D., Amodei, D., Drain, D., Li, D., Tran-Johnson, E., Khundadze, G., Kernion, J., Landis, J., Kerr, J., Mueller, J., Hyun, J., Landau, J., Ndousse, K., Goldberg, L., Lovitt, L., Lucas, M., Sellitto, M., Zhang, M., Kingsland, N., Elhage, N., Joseph, N., Mercado, N., DasSarma, N., Rausch, O., Larson, R., McCandlish, S., Johnston, S., Kravec, S., Showk, S. E., Lanham, T., Telleen-Lawton, T., Brown, T., Henighan, T., Hume, T., Bai, Y., Hatfield-Dodds, Z., Clark, J., Bowman, S. R., Askell, A., Grosse, R., Hernandez, D., Ganguli, D., Hubinger, E., Schiefer, N., and Kaplan, J · 2022
Characterizing manipulation from ai systems, 2023
Carroll, M., Chan, A., Ashton, H., and Krueger, D · 2023
Closest in time.
Towards automated circuit discovery for mechanistic interpretability, 2023
Conmy, A., Mavor-Parker, A. N., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Closest in time.
Finding alignments between interpretable causal variables and distributed neural representations
Geiger, A., Wu, Z., Potts, C., Icard, T., and Goodman, N. D · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Closest in time.
Interpretability at scale: Identifying causal mechanisms in alpaca, 2023
Wu, Z., Geiger, A., Potts, C., and Goodman, N. D · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Räukur, T., Ho, A., Casper, S., and Hadfield-Menell, D · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y
Cited in the paper.
Locating and editing factual associations in GPT
Meng, K., Bau, D., Andonian, A., and Belinkov, Y
Cited in the paper.