Fetching the paper…
Reading the bibliography…
Many neural nets appear to represent data as linear combinations of "feature vectors." Algorithms for discovering these vectors have seen impressive recent success.
Principles of neurodynamics. perceptrons and the theory of brain mechanisms
Rosenblatt, F · 1961
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Fodor, J. A. and Pylyshyn, Z. W · 1988
Earlier work this paper cites.
Implications of recursive distributed representations
Pollack, J · 1988
Earlier work this paper cites.
Local vs. distributed coding
Thorpe, S · 1989
Earlier work this paper cites.
Tensor product variable binding and the representation of symbolic structures in connectionist systems
Smolensky, P · 1990
Earlier work this paper cites.
Distributed representations and nested compositional structure
Plate, T. A · 1994
Earlier work this paper cites.
Fully distributed representation
Kanerva, P. et al · 1997
Earlier work this paper cites.
A common framework for distributed representation schemes for compositional structure
Plate, T · 1997
Earlier work this paper cites.
Learning distributed representations of concepts using linear relational embedding
Paccanaro, A. and Hinton, G. E · 2001
Earlier work this paper cites.
Representing word meaning and order information in a composite holographic lexicon
Jones, M. N. and Mewhort, D. J · 2007
Earlier work this paper cites.
Vector-based models of semantic composition
Mitchell, J. and Lapata, M · 2008
Earlier work this paper cites.
Hyperdimensional computing: An introduction to computing in distributed representation with high-dimensional random vectors
Kanerva, P · 2009
Earlier work this paper cites.
Composition in distributional models of semantics
Mitchell, J. and Lapata, M · 2010
Earlier work this paper cites.
Estimating linear models for compositional distributional semantics
Zanzotto, F., Korkontzelos, I., Fallucchi, F., Manandhar, S., et al · 2010
Earlier work this paper cites.
The neural binding problem (s)
Feldman, J · 2013
Earlier work this paper cites.
Representing objects, relations, and sequences
Gallant, S. I. and Okaywe, T. W · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Mikolov, T., Yih, W.-t., and Zweig, G · 2013
Earlier work this paper cites.
Parsing with compositional vector grammars
Socher, R., Bauer, J., Manning, C. D., and Ng, A. Y · 2013
Earlier work this paper cites.
Exploring the symbolic/subsymbolic continuum: A case study of raam
Blank, D. S., Meeden, L. A., and Marshall, J. B · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
Socher, R., Karpathy, A., Le, Q. V., Manning, C. D., and Ng, A. Y · 2014
Cited alongside, same era.
Recursive neural networks can learn logical semantics
Bowman, S., Potts, C., and Manning, C. D · 2015
Cited alongside, same era.
A latent variable model approach to pmi-based word embeddings
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A · 2016
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Rnns implicitly implement tensor product representations
McCoy, R. T., Linzen, T., Dunbar, E., and Smolensky, P · 2018
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., et al · 2023
Later among the works it cites.
Sparse autoencoders find highly interpretable features in language models
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Later among the works it cites.
How do language models bind entities in context?
Feng, J. and Steinhardt, J · 2023
Later among the works it cites.
Linearity of relation decoding in transformer language models
Hernandez, E., Sharma, A. S., Haklay, T., Meng, K., Wattenberg, M., Andreas, J., Belinkov, Y., and Bau, D · 2023
Later among the works it cites.
Is this the subspace you are looking for? an interpretability illusion for subspace activation patching
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andreas, J · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
Hewitt, J. and Manning, C. D · 2019
Cited alongside, same era.
Do attention heads in bert track syntactic dependencies?
Phang, J., Bordia, S., Bowman, S. R., et al · 2019
Cited alongside, same era.
Visualizing and measuring the geometry of bert
Reif, E., Yuan, A., Wattenberg, M., Viegas, F. B., Coenen, A., Pearce, A., and Kim, B · 2019
Cited alongside, same era.
Finding universal grammatical relations in multilingual bert
Chi, E. A., Hewitt, J., and Manning, C. D · 2020
Cited alongside, same era.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Manning, C. D., Clark, K., Hewitt, J., Khandelwal, U., and Levy, O · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Cited alongside, same era.
Makelov, A., Lange, G., Geiger, A., and Nanda, N · 2023
Later among the works it cites.
Marks, S. and Tegmark, M · 2023
Later among the works it cites.
Distributed representations: Composition and superposition
Olah, C · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al · 2023
Later among the works it cites.
Designing a dashboard for transparency and control of conversational ai
Chen, Y., Wu, A., DePodesta, T., Yeh, C., Li, K., Marin, N. C., Patel, O., Riecke, J., Raval, S., Seow, O., et al · 2024
Closest in time.
Toward a mathematical framework for computation in superposition
Dmitry Vaintrob, Jake Mendel, K · 2024
Closest in time.
Not all language model features are linear
Engels, J., Liao, I., Michaud, E. J., Gurnee, W., and Tegmark, M · 2024
Closest in time.
Monitoring latent world states in language models with propositional probes
Feng, J., Russell, S., and Steinhardt, J · 2024
Closest in time.
Sparse autoencoders work on attention layer outputs
Kissane, C., Robertzk, Conmy, A., and Nanda, N · 2024
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
Li, K., Patel, O., Viégas, F., Pfister, H., and Wattenberg, M · 2024
Closest in time.
What’s up with llms representing xors of arbitrary features?
Marks, S · 2024
Closest in time.
Talking heads: Understanding inter-layer communication in transformer language models
Merullo, J., Eickhoff, C., and Pavlick, E · 2024
Closest in time.
Fine-tuning enhances existing mechanisms: A case study on entity tracking
Prakash, N., Shaham, T. R., Haklay, T., Belinkov, Y., and Bau, D · 2024
Closest in time.
Improving dictionary learning with gated sparse autoencoders
Rajamanoharan, S., Conmy, A., Smith, L., Lieberum, T., Varma, V., Kramár, J., Shah, R., and Nanda, N · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Closest in time.