Fetching the paper…
Reading the bibliography…
Despite rapid adoption and deployment of large language models (LLMs), the internal computations of these models remain opaque and poorly understood.
Gabor filter-based edge detection
Rajiv Mehrotra, Kameswara Rao Namuduri, and Nagarajan Ranganathan · 1992
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Bruno A Olshausen and David J Field · 1996
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Bruno A Olshausen and David J Field · 1997
Earlier work this paper cites.
Spikes: exploring the neural code
Fred Rieke, David Warland, Rob de Ruyter Van Steveninck, and William Bialek · 1999
Earlier work this paper cites.
Sparse coding and decorrelation in primary visual cortex during natural vision
William E Vinje and Jack L Gallant · 2000
Earlier work this paper cites.
Information processing with population codes
Alexandre Pouget, Peter Dayan, and Richard Zemel · 2000
Earlier work this paper cites.
Modular representations of odorants in the glomerular layer of the rat olfactory bulb and the effects of stimulus concentration
Brett A Johnson and Michael Leon · 2000
Earlier work this paper cites.
Redundancy reduction revisited
Horace Barlow · 2001
Earlier work this paper cites.
Invariant visual representation by single neurons in the human brain
R Quian Quiroga, Leila Reddy, Gabriel Kreiman, Christof Koch, and Itzhak Fried · 2005
Earlier work this paper cites.
Independent codes for spatial and episodic memory in hippocampal neuronal ensembles
Stefan Leutgeb, Jill K Leutgeb, Carol A Barnes, Edvard I Moser, Bruce L McNaughton, and May-Britt Moser · 2005
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn · 2005
Earlier work this paper cites.
Compressed sensing
David L Donoho · 2006
Earlier work this paper cites.
Sparse representation of sounds in the unanesthetized auditory cortex
Tomáš Hromádka, Michael R DeWeese, and Anthony M Zador · 2008
Earlier work this paper cites.
Design principles of biological circuits
Uri Alon · 2009
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell · 2014
Earlier work this paper cites.
Mutual information between discrete and continuous data sets
Brian C Ross · 2014
Earlier work this paper cites.
Neural population coding: combining insights from microscopic and mass signals
Stefano Panzeri, Jakob H Macke, Joachim Gross, and Christoph Kayser · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Multifaceted feature visualization: Uncovering the different types of features learned by each neuron in deep neural networks, 2016
Anh Nguyen, Jason Yosinski, and Jeff Clune · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever · 2017
Earlier work this paper cites.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties, 2018
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni · 2018
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass · 2018
Earlier work this paper cites.
Kedar Dhamdhere, Mukund Sundararajan, and Qiqi Yan · 2018
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2018
Earlier work this paper cites.
Disentangling by factorising
Hyunjik Kim and Andriy Mnih · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Interactive supercomputing on 40,000 cores for machine learning and data analysis
Albert Reuther, Jeremy Kepner, Chansup Byun, Siddharth Samsi, William Arcand, David Bestor, Bill Bergeron, Vijay Gadepally, Michael Houle, Matthew Hubbell, Michael Jones, Anna Klein, Lauren Milechin, Julia Mullen, Andrew Prout, Antonio Rosa, Charles Yee, and Peter Michaleas · 2018
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Designing and interpreting probes with control tasks
J Hewitt and P Liang · 2019
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al · 2019
Earlier work this paper cites.
What is one grain of sand in the desert? analyzing individual neurons in deep nlp models
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass · 2019
Earlier work this paper cites.
On interpretability and feature representations: an analysis of the sentiment neuron
Jonathan Donnelly and Adam Roegiest · 2019
Earlier work this paper cites.
Visualizing and measuring the geometry of bert, 2019
Andy Coenen, Emily Reif, Ann Yuan, Been Kim, Adam Pearce, Fernanda Viégas, and Martin Wattenberg · 2019
Cited alongside, same era.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Cited alongside, same era.
Understanding the role of individual units in a deep neural network
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba · 2020
Cited alongside, same era.
Curve circuits
Nick Cammarata, Gabriel Goh, Shan Carter, Chelsea Voss, Ludwig Schubert, and Chris Olah · 2020
Cited alongside, same era.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas · 2020
Cited alongside, same era.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy · 2020
Finding skill neurons in pre-trained transformer-based language models
Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou, Zhiyuan Liu, and Juanzi Li · 2022
Later among the works it cites.
Linguistic correlation analysis: Discovering salient neurons in deepnlp models
Nadir Durrani, Fahim Dalvi, and Hassan Sajjad · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sparse regression: Scalable algorithms and empirical performance
Dimitris Bertsimas, Jean Pauphilet, and Bart Van Parys · 2020
Cited alongside, same era.
A primer in bertology: What we know about how bert works, 2020
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2020
Cited alongside, same era.
Finding experts in transformer models
Xavier Suau, Luca Zappella, and Nicholas Apostoloff · 2020
Cited alongside, same era.
Analyzing individual neurons in pre-trained language models
Nadir Durrani, Hassan Sajjad, Fahim Dalvi, and Yonatan Belinkov · 2020
Cited alongside, same era.
Intrinsic probing through dimension selection
Lucas Torroba Hennigen, Adina Williams, and Ryan Cotterell · 2020
Cited alongside, same era.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov · 2020
Cited alongside, same era.
Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg · 2022
Later among the works it cites.
Analyzing transformers in embedding space
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant · 2022
Later among the works it cites.
Engineering monosemanticity in toy models
Adam S Jermyn, Nicholas Schiefer, and Evan Hubinger · 2022
Later among the works it cites.
Taking features out of superposition with sparse autoencoders, 2022
Lee Sharkey, Dan Braun, and Beren Millidge · 2022
Later among the works it cites.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Later among the works it cites.
Neuroscope: A website for mechanistic interpretability of language models, 2022
Neel Nanda · 2022
Later among the works it cites.
Natural abstractions: Key claims, theorems, and critiques, 2022
Lawrence Chan, Leon Lang, and Erik Jenner · 2022
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Later among the works it cites.
Is power-seeking ai an existential risk?
Joseph Carlsmith · 2022
Later among the works it cites.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Later among the works it cites.
Transformerlens, 2022
Neel Nanda · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Impossibility theorems for feature attribution
Blair Bilodeau, Natasha Jaques, Pang Wei Koh, and Been Kim · 2022
Later among the works it cites.
On the relationship between explanation and prediction: A causal view
Amir-Hossein Karimi, Krikamol Muandet, Simon Kornblith, Bernhard Schölkopf, and Been Kim · 2022
Later among the works it cites.
Interpreting neural networks through the polytope lens, 2022
Sid Black, Lee Sharkey, Leo Grinsztajn, Eric Winsor, Dan Braun, Jacob Merizian, Kip Parker, Carlos Ramón Guevara, Beren Millidge, Gabriel Alfour, and Connor Leahy · 2022
Later among the works it cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2022
Later among the works it cites.
Localizing model behavior with path patching
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
A toy model of universality: Reverse engineering how networks learn group operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Closest in time.
The alignment problem from a deep learning perspective, 2023
Richard Ngo, Lawrence Chan, and Sören Mindermann · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Closest in time.
Privileged bases in the transformer residual stream
Nelson Elhage, Robert Lasenby, and Christopher Olah · 2023
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling, 2023
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal · 2023
Closest in time.
The quantization model of neural scaling
Eric J Michaud, Ziming Liu, Uzay Girit, and Max Tegmark · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun · 2023
Closest in time.
Distributed representations: Composition & superposition, 2023
Chris Olah · 2023
Closest in time.
A comprehensive mechanistic interpretability explainer & glossary, 2023
Neel Nanda · 2023
Closest in time.
Superposition, memorization, and double descent
Tom Henighan, Shan Carter, Tristan Hume, Nelson Elhage, Robert Lasenby, Stanislav Fort, Nicholas Schiefer, and Christopher Olah · 2023
Closest in time.
Circuits updates — may 2023: Attention head superposition, 2023
Adam Jermyn, Chris Olah, and Tom Henighan · 2023
Closest in time.
Actually, othello-gpt has a linear emergent world model, Mar 2023
Neel Nanda · 2023
Closest in time.