Fetching the paper…
Reading the bibliography…
Mechanistic Interpretability aims to reverse engineer the algorithms implemented by neural networks by studying their weights and activations.
Estimating the dimension of a model
Gideon Schwarz · 1978
Earlier work this paper cites.
Statistical Learning Theory
Vladimir N. Vapnik · 1998
Earlier work this paper cites.
Spontaneous evolution of modularity and network motifs
Nadav Kashtan and Uri Alon · 2005
Earlier work this paper cites.
Algebraic geometry and statistical learning theory , volume 25
Sumio Watanabe · 2009
Earlier work this paper cites.
The evolutionary origins of modularity
Jeff Clune, Jean-Baptiste Mouret, and Hod Lipson · 2012
Earlier work this paper cites.
Transformer Feed-Forward Layers Are Key-Value Memories, September 2021
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2012
Earlier work this paper cites.
Machine Learning: A Probabilistic Perspective
Kevin P Murphy · 2012
Earlier work this paper cites.
A widely applicable bayesian information criterion
Sumio Watanabe · 2013
Earlier work this paper cites.
Why neurons mix: high dimensionality for higher cognition
Stefano Fusi, Earl K Miller, and Mattia Rigotti · 2016
Earlier work this paper cites.
A kronecker-factored approximate fisher matrix for convolution layers, 2016
Roger Grosse and James Martens · 2016
Earlier work this paper cites.
Multifaceted feature visualization: Uncovering the different types of features learned by each neuron in deep neural networks, 2016
Anh Nguyen, Jason Yosinski, and Jeff Clune · 2016
Earlier work this paper cites.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
Visualizing the loss landscape of neural nets, 2018
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Earlier work this paper cites.
Kronecker-factored curvature approximations for recurrent neural networks
James Martens, Jimmy Ba, and Matt Johnson · 2018
Earlier work this paper cites.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Earlier work this paper cites.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape, 2019
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred A. Hamprecht · 2019
Earlier work this paper cites.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2019
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Eigendamage: Structured pruning in the kronecker-factored eigenbasis, 2019
Chaoqi Wang, Roger Grosse, Sanja Fidler, and Guodong Zhang · 2019
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature, 2020
James Martens and Roger Grosse · 2020
Cited alongside, same era.
Singular learning theory iv: the rlct
Daniel Murfet · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Cited alongside, same era.
Deep learning is singular, and that’s good
Susan Wei, Daniel Murfet, Mingming Gong, Hui Li, Jesse Gell-Redman, and Thomas Quella · 2022
Later among the works it cites.
Dslt 1. the rlct measures the effective dimension of neural networks, Jun 2023
Liam Carroll · 2023
Later among the works it cites.
Dynamical versus bayesian phase transitions in a toy model of superposition
Zhongtian Chen, Edmund Lau, Jake Mendel, Susan Wei, and Daniel Murfet · 2023
Later among the works it cites.
Towards automated circuit discovery for mechanistic interpretability, 2023
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Later among the works it cites.
Studying large language model generalization with influence functions, 2023
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, Evan Hubinger, Kamilė Lukošiūtė, Karina Nguyen, Nicholas Joseph, Sam McCandlish, Jared Kaplan, and Samuel R. Bowman · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Phase transitions in neural networks
Liam Carrol · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
Clusterability in neural networks, 2021
Daniel Filan, Stephen Casper, Shlomi Hod, Cody Wild, Andrew Critch, and Stuart Russell · 2021
Cited alongside, same era.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Cited alongside, same era.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata †, Chelsea Voss †, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Cited alongside, same era.
Is sgd a bayesian sampler? well, almost
Chris Mingard, Guillermo Valle-Pérez, Joar Skalse, and Ard A Louis · 2021
Cited alongside, same era.
Neural networks generalise because of this one weird trick
Jesse Hoogland · 2023
Later among the works it cites.
Towards developmental interpretability, Jul 2023
Jesse Hoogland, Alexander Gietelink Oldenziel, Daniel Murfet, and Stan van Wingerden · 2023
Later among the works it cites.
Quantifying degeneracy in singular models via the learning coefficient
Edmund Lau, Daniel Murfet, and Susan Wei · 2023
Later among the works it cites.
Seeing is believing: Brain-inspired modular training for mechanistic interpretability, 2023
Ziming Liu, Eric Gan, and Max Tegmark · 2023
Later among the works it cites.
Locating and editing factual associations in gpt, 2023
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability, 2023
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Later among the works it cites.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks, 2023
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2023
Later among the works it cites.
Consideration on the learning efficiency of multiple-layered neural networks with linear units
Miki Aoyagi · 2024
Closest in time.
Lucius Bushnaq, Stefan Heimersheim, Nicholas Goldowsky-Dill, Dan Braun, Jake Mendel, Kaarel Hänni, Avery Griffin, Jörn Stöhler, Magdalena Wache, and Marius Hobbhahn · 2024
Closest in time.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2024
Closest in time.
The developmental landscape of in-context learning, 2024
Jesse Hoogland, George Wang, Matthew Farrugia-Roberts, Liam Carroll, Susan Wei, and Daniel Murfet · 2024
Closest in time.