Fetching the paper…
Reading the bibliography…
Recent studies in interpretability have explored the inner workings of transformer models trained on tasks across various domains, often discovering that these networks naturally develop highly structured representations.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning · 2019
Earlier work this paper cites.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Christopher D Manning, Kevin Clark, John Hewitt, et al · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, et al · 2021
Earlier work this paper cites.
Nelson Elhage, Tristan Hume, Catherine Olsson, et al · 2022
Earlier work this paper cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, David Bau, et al · 2022
Earlier work this paper cites.
Acquisition of chess knowledge in alphazero
Thomas McGrath, Andrei Kapishnikov, Nenad Tomašev, et al · 2022
Earlier work this paper cites.
Transformerlens
Neel Nanda and Joseph Bloom · 2022
Earlier work this paper cites.
In-context Learning and Induction Hea ds
Catherine Olsson, Nelson Elhage, Neel Nanda, et al · 2022
Earlier work this paper cites.
Sparse Relational Reasoning with Object-Centric Representations, July 2022
Alex F. Spies, Alessandra Russo, and Murray Shanahan · 2022
Cited alongside, same era.
Eliciting Latent Predictions from Transformers with the Tuned Lens
Nora Belrose, Zach Furman, Logan Smith, et al · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Brian Chen, et al · 2023
Cited alongside, same era.
Sparse Autoencoders Find Highly Interpretable Features in Language Models, October 2023
Hoagy Cunningham, Aidan Ewart, Logan Riggs, et al · 2023
Cited alongside, same era.
A configurable library for generating and manipulating maze datasets
Zhengfu He, Xuyang Ge, Qiong Tang, et al · 2024
Closest in time.
Linearly structured world representations in maze-solving transformers
Michael I. Ivanitskiy, Alex F. Spies, Tilman Räuker, et al · 2024
Closest in time.
Evidence of learned look-ahead in a chess-playing neural network
Erik Jenner, Shreyas Kapur, Vasil Georgiev, et al · 2024
Closest in time.
Anthropic circuits Updates - January 2024, January 2024
Adam Jermyn and Adly Templeton · 2024
Closest in time.
Emergent world models and latent variable estimation in chess-playing language models
Adam Karvonen · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael I. Ivanitskiy, Rusheb Shah, Alex F Spies, et al · 2023
Cited alongside, same era.
Tom Lieberum, Matthew Rahtz, János Kramár, et al · 2023
Cited alongside, same era.
Actually, othello-gpt has a linear emergent world representation
Neel Nanda · 2023
Cited alongside, same era.
Future lens: Anticipating subsequent tokens from a single hidden state
Koyena Pal, Jiuding Sun, Andrew Yuan, et al · 2023
Cited alongside, same era.
Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2023
Cited alongside, same era.
A mechanistic analysis of a transformer trained on a symbolic multi-step reasoning task
Jannik Brinkmann, Abhay Sheshadri, Victor Levoso, et al · 2024
Cited alongside, same era.
Measuring progress in dictionary learning for language model interpretability with board game models
Adam Karvonen, Benjamin Wright, Can Rager, et al · 2024
Closest in time.
SAE Visualizer
Callum McDougall · 2024
Closest in time.
A philosophical introduction to language models - part ii: The way forward, 2024
Raphaël Millière and Cameron Buckner · 2024
Closest in time.
Evaluating cognitive maps and planning in large language models with cogeval
Ida Momennejad, Hosein Hasanbeig, Felipe Vieira Frujeri, et al · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.