Successor heads: Recurring, interpretable attention heads in the wild
Original
Gould, R., Ong, E., Ogden, G., and Conmy, A · 2023
Later among the works it cites.
Does circuit analysis interpretability scale? evidence from multiple choice capabilities in chinchilla
Original
Lieberum, T., Rahtz, M., Kramár, J., Nanda, N., Irving, G., Shah, R., and Mikulik, V · 2023
Later among the works it cites.
Transformers need glasses! information over-squashing in language tasks
Barbero, F., Banino, A., Kapturowski, S., Kumaran, D., Madeira Araújo, J., Vitvitskyi, O., Pascanu, R., and Veličković, P · 2024
Closest in time.
Improving fine-grained understanding in image-text pre-training
Bica, I., Ilic, A., Bauer, M., Erdogan, G., Bošnjak, M., Kaplanis, C., Gritsenko, A. A., Minderer, M., Blundell, C., Pascanu, R., and Mitrovic, J · 2024
Closest in time.
MVSFormer++: Revealing the devil in transformer’s details for multi-view stereo
Cao, C., Ren, X., and Fu, Y · 2024
Closest in time.
Asynchronous algorithmic alignment with cocycles
Dudzik, A. J., von Glehn, T., Pascanu, R., and Veličković, P · 2024
Closest in time.
Your context is not an array: Unveiling random access limitations in transformers
Original
Ebrahimi, M., Panchal, S., and Memisevic, R · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Original
Gemma Team, Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al · 2024
Closest in time.
How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Hanna, M., Liu, O., and Variengien, A · 2024
Closest in time.
Flax: A neural network library and ecosystem for JAX, 2024
Heek, J., Levskaya, A., Oliver, A., Ritter, M., Rondepierre, B., Steiner, A., and van Zee, M · 2024
Closest in time.
Aero: Softmax-only llms for efficient private inference
Original
Jha, N. K. and Reagen, B · 2024
Closest in time.
Penzai + Treescope: A toolkit for interpreting, visualizing, and editing models as data
Johnson, D. D · 2024
Closest in time.
Interpreting attention layer outputs with sparse autoencoders
Original
Kissane, C., Krzyzanowski, R., Bloom, J. I., Conmy, A., and Nanda, N · 2024
Closest in time.
The CLRS-Text Algorithmic Reasoning Language Benchmark
Original
Markeeva, L., McLeish, S., Ibarz, B., Bounsi, W., Kozlova, O., Vitvitskyi, A., Blundell, C., Goldstein, T., Schwarzschild, A., and Veličković, P · 2024
Closest in time.
Theory, analysis, and best practices for sigmoid self-attention
Original
Ramapuram, J., Danieli, F., Dhekane, E., Weers, F., Busbridge, D., Ablin, P., Likhomanenko, T., Digani, J., Gu, Z., Shidani, A., and Webb, R · 2024
Closest in time.
Retrieval head mechanistically explains long-context factuality
Original
Wu, W., Wang, Y., Xiao, G., Peng, H., and Fu, Y · 2024
Closest in time.
Entropix: Entropy Based Sampling and Parallel CoT Decoding, 2024
xjdr and doomslide · 2024
Closest in time.
Selective attention improves transformer
Leviathan, Y., Kalman, M., and Matias, Y · 2025
Closest in time.
Scaling stick-breaking attention: An efficient implementation and in-depth study
Tan, S., Yang, S., Courville, A., Panda, R., and Shen, Y · 2025
Closest in time.
What makes a good feedforward computational graph?
Original
Vitvitskyi, A., Araújo, J. G., Lackenby, M., and Veličković, P · 2025
Closest in time.
Differential transformer
Ye, T., Dong, L., Xia, Y., Sun, Y., Zhu, Y., Huang, G., and Wei, F · 2025
Closest in time.