2022

Engineering Monosemanticity in Toy Models

Jermyn, Adam S., Schiefer, Nicholas, Hubinger, Evan

Understand

In some neural networks, individual neurons correspond to natural ``features'' in the input.

  • Such \emph{monosemantic} neurons are of great help in interpretability studies, as they can be cleanly understood.
  • In this work we report preliminary attempts to engineer monosemanticity in toy models.
  • We find that models can be made more monosemantic without increasing the loss by just changing which local minimum the training process finds.

Built on

  • Reducing BERT pre-training time from 3 days to 76 minutes

    Original

    Yang You, Jing Li, Jonathan Hseu, Xiaodan Song, James Demmel, and Cho-Jui Hsieh · 1904

    Earlier work this paper cites.

  • Zipf’s word frequency law in natural language: A critical review and future directions

    Steven T. Piantadosi · 2014

    Earlier work this paper cites.

  • Feature visualization

    Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017

    Earlier work this paper cites.

Similar

  • Zoom in: An introduction to circuits

    Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020

    Cited alongside, same era.

  • Softmax linear units

    Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah · 2022

    Cited alongside, same era.

Then

  • Toy models of superposition

    Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022

    Closest in time.

  • Polysemanticity and capacity in neural networks, 2022

    Adam Scherlis, Kshitij Sachan, Adam S. Jermyn, Joe Benton, and Buck Shlegeris · 2022

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…