Fetching the paper…
Reading the bibliography…
Variable binding -- the ability to associate variables with values -- is fundamental to symbolic computation and cognition.
Relational reasoning and generalization using nonsymbolic neural networks
Geiger, A., Carstensen, A., Frank, M. C., and Potts, C · 1939
Earlier work this paper cites.
Tensor product variable binding and the representation of symbolic structures in connectionist systems
Smolensky, P · 1990
Earlier work this paper cites.
Rethinking eliminative connectionism
Marcus, G. F · 1998
Earlier work this paper cites.
Generalization, rules, and neural networks: A simulation of Marcus et. al
Elman, J. L · 1999
Earlier work this paper cites.
Rule learning by seven-month-old infants
Marcus, G. F., Vijayan, S., Rao, S. B., and Vishton, P. M · 1999
Earlier work this paper cites.
The Algebraic Mind: Integrating Connectionism and Cognitive Science
Marcus, G. F · 2001
Earlier work this paper cites.
Memory and the Computational Brain: Why Cognitive Science will Transform Neuroscience
Gallistel, C. and King, A · 2011
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Fixing weight decay regularization in adam
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection for dimension reduction
McInnes, L., Healy, J., and Melville, J · 2018
Earlier work this paper cites.
A review of computational models of basic rule learning: The neural-symbolic debate and beyond
Alhama, R. G. and Zuidema, W · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Thread: Circuits
Cammarata, N., Carter, S., Goh, G., Olah, C., Petrov, M., Schubert, L., Voss, C., Egan, B., and Lim, S. K · 2020
Cited alongside, same era.
How do decisions emerge across layers in neural models? Interpretation with differentiable masking
Cao, N. D., Schlichtkrull, M. S., Aziz, W., and Titov, I · 2020
Cited alongside, same era.
Neural natural language inference models partially embed theories of lexical entailment and negation
Geiger, A., Richardson, K., and Potts, C · 2020
Cited alongside, same era.
Investigating gender bias in language models using causal mediation analysis
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S · 2020
Cited alongside, same era.
Are neural nets modular? inspecting functional modularity through differentiable weight masks
Csordás, R., van Steenkiste, S., and Schmidhuber, J · 2021
Cited alongside, same era.
Causal abstractions of neural networks
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Later among the works it cites.
Unveiling transformers with LEGO: A synthetic reasoning task, 2023
Zhang, Y., Backurs, A., Bubeck, S., Eldan, R., Gunasekar, S., and Wagner, T · 2023
Later among the works it cites.
Representational Analysis of Binding in Language Models
Dai, Q., Heinzerling, B., and Inui, K · 2024
Later among the works it cites.
How do language models bind entities in context?
Feng, J. and Steinhardt, J · 2024
Later among the works it cites.
How do language models bind entities in context?
Feng, J. and Steinhardt, J · 2024
Later among the works it cites.
Monitoring Latent World States in Language Models with Propositional Probes, June 2024
Feng, J., Russell, S., and Steinhardt, J · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Geiger, A., Lu, H., Icard, T., and Potts, C · 2021
Cited alongside, same era.
Predicting inductive biases of pre-trained models
Lovering, C., Jha, R., Linzen, T., and Pavlick, E · 2021
Cited alongside, same era.
Sparse interventions in language models with differentiable masking
Cao, N. D., Schmid, L., Hupkes, D., and Titov, I · 2022
Cited alongside, same era.
Causal scrubbing, a method for rigorously testing interpretability hypotheses
Chan, L., Garriga-Alonso, A., Goldwosky-Dill, N., Greenblatt, R., Nitishinskaya, J., Radhakrishnan, A., Shlegeris, B., and Thomas, N · 2022
Cited alongside, same era.
Towards understanding grokking: An effective theory of representation learning
Liu, Z., Kitouni, O., Nolte, N. S., Michaud, E., Tegmark, M., and Williams, M · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Meng, K., Bau, D., Andonian, A. J., and Belinkov, Y · 2022
Cited alongside, same era.
Discovering variable binding circuitry with desiderata, 2023
Davies, X., Nadeau, M., Prakash, N., Shaham, T. R., and Bau, D · 2023
Cited alongside, same era.
Causal abstraction: A theoretical foundation for mechanistic interpretability, 2024
Geiger, A., Ibeling, D., Zur, A., Chaudhary, M., Chauhan, S., Huang, J., Arora, A., Wu, Z., Goodman, N., Potts, C., and Icard, T · 2024
Later among the works it cites.
Hashhop: Long context evaluation
Magic · 2024
Later among the works it cites.
Mueller, A., Brinkmann, J., Li, M., Marks, S., Pal, K., Prakash, N., Rager, C., Sankaranarayanan, A., Sharma, A. S., Sun, J., Todd, E., Bau, D., and Belinkov, Y · 2024
Later among the works it cites.
Fine-tuning enhances existing mechanisms: A case study on entity tracking
Prakash, N., Shaham, T. R., Haklay, T., Belinkov, Y., and Bau, D · 2024
Later among the works it cites.
Mechanistic?
Saphra, N. and Wiegreffe, S · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Later among the works it cites.
Open problems in mechanistic interpretability, 2025
Sharkey, L., Chughtai, B., Batson, J., Lindsey, J., Wu, J., Bushnaq, L., Goldowsky-Dill, N., Heimersheim, S., Ortega, A., Bloom, J., Biderman, S., Garriga-Alonso, A., Conmy, A., Nanda, N., Rumbelow, J., Wattenberg, M., Schoots, N., Miller, J., Michaud, E. J., Casper, S., Tegmark, M., Saunders, W., Bau, D., Todd, E., Geiger, A., Geva, M., Hoogland, J., Murfet, D., and McGrath, T · 2025
Closest in time.