Fetching the paper…
Reading the bibliography…
Recent works have shown that transformers can solve contextual reasoning tasks by internally executing computational graphs called circuits.
Semantic subspace learning with conditional significance vectors
Tripathi et al · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov et al · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov et al · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Mikolov et al · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi et al · 2015
Earlier work this paper cites.
Ba et al · 2016
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Sanjeev et al · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani et al · 2017
Earlier work this paper cites.
Improving lexical choice in neural machine translation
Toan Nguyen and David Chiang · 2018
Earlier work this paper cites.
Root mean square layer normalization
Biao Zhang and Rico Sennrich · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin et al · 2019
Earlier work this paper cites.
Transformers without tears: Improving the normalization of self-attention
Nguyen et al · 2019
Earlier work this paper cites.
Understanding and improving layer normalization
Xu et al · 2019
Earlier work this paper cites.
Visualizing and measuring the geometry of bert
Coenen et al · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
Hewitt et al · 2019
Earlier work this paper cites.
How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings
Kawin Ethayarajh · 2019
Earlier work this paper cites.
Deep learning using rectified linear units (relu), 2019
Abien Fred Agarap · 2019
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
A primer in bertology: What we know about how bert works
Rogers et al · 2020
Earlier work this paper cites.
Thread: Circuits
Cammarata et al · 2020
Earlier work this paper cites.
On layer normalization in the transformer architecture
Xiong et al · 2020
Earlier work this paper cites.
Query-key normalization for transformers
Henry et al · 2020
Cited alongside, same era.
Finding universal grammatical relations in multilingual BERT
Chi et al · 2020
Cited alongside, same era.
Array programming with NumPy
Harris et al · 2020
Cited alongside, same era.
Accurate prediction of protein structures and interactions using a three-track neural network
Baek et al · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Elhage et al · 2021
Cited alongside, same era.
Normformer: Improved transformer pretraining with extra normalization, 2021
Shleifer et al · 2021
Cited alongside, same era.
Localizing model behavior with path patching, 2023
Goldowsky-Dill et al · 2023
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
Dehghani et al · 2023
Later among the works it cites.
A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis, 2023
Stolfo et al · 2023
Later among the works it cites.
Localizing model behavior with path patching
Goldowsky-Dill et al · 2023
Later among the works it cites.
The transient nature of emergent in-context learning in transformers
Singh et al · 2023
Later among the works it cites.
On the expressivity role of layernorm in transformers’ attention
Brody et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Incorporating Residual and Normalization Layers into Analysis of Masked Language Models
Kobayashi et al · 2021
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Dong et al · 2021
Cited alongside, same era.
Transformers with competitive ensembles of independent mechanisms
Lambd et al · 2021
Cited alongside, same era.
Isotropy in the contextual embedding space: Clusters and manifolds
Cai et al · 2021
Cited alongside, same era.
The low-dimensional linear geometry of contextualized word representations
Evan Hernandez and Jacob Andreas · 2021
Cited alongside, same era.
In-context learning and induction heads
Olsson et al · 2022
Cited alongside, same era.
Later among the works it cites.
Traveling words: A geometric interpretation of transformers
Raul Molina · 2023
Later among the works it cites.
The linear representation hypothesis and the geometry of large language models
Park et al · 2023
Later among the works it cites.
Investigating semantic subspaces of transformer sentence embeddings through linear structural probing
Dmitry Nikolaev and Sebastian Padó · 2023
Later among the works it cites.
Finding alignments between interpretable causal variables and distributed neural representations
Geiger et al · 2023
Later among the works it cites.
Accurate structure prediction of biomolecular interactions with alphafold 3
Abramson et al · 2024
Closest in time.
Transforming the bootstrap: Using transformers to compute scattering amplitudes in planar n = 4 super yang-mills theory, 2024
Cai et al · 2024
Closest in time.
A primer on the inner workings of transformer-based language models, 2024
Ferrando et al · 2024
Closest in time.
Small-scale proxies for large-scale transformer training instabilities
Wortsman et al · 2024
Closest in time.
Information flow routes: Automatically interpreting language models at scale
Javier Ferrando and Elena Voita · 2024
Closest in time.
What needs to go right for an induction head? a mechanistic study of in-context learning circuits and their formation, 2024
Singh et al · 2024
Closest in time.
On the role of attention masks and layernorm in transformers, 2024
Wu et al · 2024
Closest in time.
Towards understanding how transformer perform multi-step reasoning with matching operation, 2024
Wang et al · 2024
Closest in time.
When can transformers reason with abstract symbols?
Boix-Adserà et al · 2024
Closest in time.
On the origins of linear representations in large language models, 2024
Jiang et al · 2024
Closest in time.
Uncovering hidden geometry in transformers via disentangling position and context, 2024
Song et al · 2024
Closest in time.
Linearity of relation decoding in transformer language models
Hernandez et al · 2024
Closest in time.