In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Cited alongside, same era.
Choose your programming copilot: A comparison of the program synthesis performance of github copilot and genetic programming
Sobania, D., Briesch, M., and Rothlauf, F · 2022
Cited alongside, same era.
Chess as a testbed for language model state tracking
Toshniwal, S., Wiseman, S., Livescu, K., and Gimpel, K · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Original
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Cited alongside, same era.
Language models can explain neurons in language models
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Cited alongside, same era.
Can transformers learn the greatest common divisor?
Original
Charton, F · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Original
Conmy, A., Mavor-Parker, A. N., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Cited alongside, same era.
Interpretable machine learning for science with pysr and symbolicregression. jl
Original
Cranmer, M · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Original
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Original
Gu, A. and Dao, T · 2023
Cited alongside, same era.