Fetching the paper…
Reading the bibliography…
We introduce methods for discovering and applying sparse feature circuits.
Counterfactuals
David K. Lewis · 1973
Earlier work this paper cites.
Identifiability and exchangeability for direct and indirect effects, Identifiability and Exchangeability for Direct and Indirect Effects
James M. Robins and Sander Greenland · 1992
Earlier work this paper cites.
Learning Factorial Codes by Predictability Minimization, Learning Factorial Codes by Predictability Minimization
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Direct and indirect effects
Judea Pearl · 2001
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Disentangling factors of variation via generative entangling
Guillaume Desjardins, Aaron Courville, and Yoshua Bengio · 2012
Earlier work this paper cites.
k-sparse autoencoders, k-Sparse Autoencoders
Alireza Makhzani and Brendan J. Frey · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization, Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Infogan: interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Understanding disentangling in β \beta -VAE, 2017, Understanding disentangling in β \beta -VAE
Christopher P. Burgess, Irina Higgins, Arka Pal, Loic Matthey, Nick Watters, Guillaume Desjardins, and Alexander Lerchner · 2017
Earlier work this paper cites.
beta-VAE: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Isolating sources of disentanglement in variational autoencoders, 2018, Isolating Sources of Disentanglement in Variational Autoencoders
Tian Qi Chen, Xuechen Li, Roger Grosse, and David Duvenaud · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Earlier work this paper cites.
Disentangling by factorising
Hyunjik Kim and Andriy Mnih · 2018
Earlier work this paper cites.
Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study, Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study
John R. Zech, Marcus A. Badgeley, Manway Liu, Anthony B. Costa, Joseph J. Titano, and Eric Karl Oermann · 2018
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai · 2019
Earlier work this paper cites.
Distributionally robust language modeling
Yonatan Oren, Shiori Sagawa, Tatsunori B. Hashimoto, and Percy Liang · 2019
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Earlier work this paper cites.
Learning from failure: Training debiased classifier from biased classifier
Junhyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee, and Jinwoo Shin · 2020
Earlier work this paper cites.
The hessian penalty: A weak prior for unsupervised disentanglement
William Peebles, John Peebles, Jun-Yan Zhu, Alexei A. Efros, and Antonio Torralba · 2020
Earlier work this paper cites.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg · 2020
Earlier work this paper cites.
Distributionally robust neural networks
Shiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, and Percy Liang · 2020
Earlier work this paper cites.
Towards debiasing NLU models from unknown biases
Prasetya Ajie Utama, Nafise Sadat Moosavi, and Iryna Gurevych · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2020
Earlier work this paper cites.
Double-hard debias: Tailoring word embeddings for gender bias mitigation
Tianlu Wang, Xi Victoria Lin, Nazneen Fatema Rajani, Bryan McCann, Vicente Ordonez, and Caiming Xiong · 2020
Earlier work this paper cites.
Environment inference for invariant learning
Elliot Creager, Joern-Henrik Jacobsen, and Richard Zemel · 2021
Earlier work this paper cites.
Causal analysis of syntactic agreement mechanisms in neural language models
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Cited alongside, same era.
Just train twice: Improving group robustness without training group information
Evan Z Liu, Behzad Haghgoo, Annie S Chen, Aditi Raghunathan, Pang Wei Koh, Shiori Sagawa, Percy Liang, and Chelsea Finn · 2021
Cited alongside, same era.
Explaining neural networks by decoding layer activations
Johannes Schneider and Michalis Vlachos · 2021
Cited alongside, same era.
Increasing robustness to spurious correlations using forgettable examples
Yadollah Yaghoobzadeh, Soroush Mehri, Remi Tachet des Combes, T. J. Hazen, and Alessandro Sordoni · 2021
Cited alongside, same era.
Coping with label shift via distributionally robust optimisation
Jingzhao Zhang, Aditya Krishna Menon, Andreas Veit, Srinadh Bhojanapalli, Sanjiv Kumar, and Suvrit Sra · 2021
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Shielded representations: Protecting sensitive attributes through iterative gradient-based projection
Shadi Iskander, Kira Radinsky, and Yonatan Belinkov · 2023
Later among the works it cites.
Last layer re-training is sufficient for robustness to spurious correlations
Polina Kirichenko, Pavel Izmailov, and Andrew Gordon Wilson · 2023
Later among the works it cites.
Neuronpedia: Interactive reference and tooling for analyzing neural networks, 2023, Neuronpedia: Interactive Reference and Tooling for Analyzing Neural Networks
Johnny Lin and Joseph Bloom · 2023
Later among the works it cites.
The quantization model of neural scaling
Eric J Michaud, Ziming Liu, Uzay Girit, and Max Tegmark · 2023
Later among the works it cites.
Open source replication & commentary on Anthropic’s dictionary learning paper, 2023, Open Source Replication & Commentary on Anthropic’s Dictionary Learning Paper
Neel Nanda · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yonatan Belinkov · 2022
Cited alongside, same era.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Cited alongside, same era.
Inducing causal structure for interpretable neural networks
Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah Goodman, and Christopher Potts · 2022
Cited alongside, same era.
Exploring linear feature disentanglement for neural networks
T. He, Z. Li, Y. Gong, Y. Yao, X. Nie, and Y. Yin · 2022
Cited alongside, same era.
Simple data balancing achieves competitive worst-group-accuracy
Badr Youbi Idrissi, Martin Arjovsky, Mohammad Pezeshki, and David Lopez-Paz · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
Spread spurious attribute: Improving worst-group accuracy with spurious attribute estimation, 2022
Junhyun Nam, Jaehyung Kim, Jaeho Lee, and Jinwoo Shin · 2022
Cited alongside, same era.
Later among the works it cites.
Fact finding: Attempting to reverse-engineer factual recall on the neuron level, 2023, Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level
Neel Nanda, Senthooran Rajamanoharan, János Kramár, and Rohin Shah · 2023
Later among the works it cites.
Label-free concept bottleneck models
Tuomas Oikarinen, Subhro Das, Lam M. Nguyen, and Tsui-Wei Weng · 2023
Later among the works it cites.
BLIND: Bias removal with no demographics
Hadas Orgad and Yonatan Belinkov · 2023
Later among the works it cites.
Attribution patching outperforms automated circuit discovery
Aaquib Syed, Can Rager, and Arthur Conmy · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Later among the works it cites.
Robust and interpretable medical image classifiers via concept bottleneck models
An Yan, Yu Wang, Yiwu Zhong, Zexue He, Petros Karypis, Zihan Wang, Chengyu Dong, Amilcare Gentili, Chun-Nan Hsu, Jingbo Shang, and Julian McAuley · 2023
Later among the works it cites.
Characterizing mechanisms for factual recall in language models
Qinan Yu, Jack Merullo, and Ellie Pavlick · 2023
Later among the works it cites.
Representation engineering: A top-down approach to AI transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al · 2023
Later among the works it cites.
Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs
Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L Leavitt, and Naomi Saphra · 2024
Closest in time.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2024
Closest in time.
Interpreting CLIP’s image representation via text-based decomposition
Yossi Gandelsman, Alexei A. Efros, and Jacob Steinhardt · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2024
Closest in time.
Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms
Michael Hanna, Sandro Pezzelle, and Yonatan Belinkov · 2024
Closest in time.
The unreasonable effectiveness of easy training data for hard tasks
Peter Hase, Mohit Bansal, Peter Clark, and Sarah Wiegreffe · 2024
Closest in time.
Leveraging prototypical representations for mitigating social bias without demographic information
Shadi Iskander, Kira Radinsky, and Yonatan Belinkov · 2024
Closest in time.
AtP*: An efficient and scalable method for localizing llm behaviour to components
János Kramár, Tom Lieberum, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024
Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca Dragan, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Aaron Mueller, Jannik Brinkmann, Millicent Li, Samuel Marks, Koyena Pal, Nikhil Prakash, Can Rager, Aruna Sankaranarayanan, Arnab Sen Sharma, Jiuding Sun, Eric Todd, David Bau, and Yonatan Belinkov · 2024
Closest in time.
The alignment problem from a deep learning perspective
Richard Ngo, Lawrence Chan, and Sören Mindermann · 2024
Closest in time.
Fine-tuning enhances existing mechanisms: A case study on entity tracking
Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size, 2024
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, Anton Tsitsulin, Nino Vieillard, Piotr Stanczyk, Sertan Girgin, Nikola Momchev, Matt Hoffman, Shantanu Thakoor, Jean-Bastien Grill, Behnam Neyshabur, Olivier Bachem, Alanna Walton, Aliaksei Severyn, Alicia Parrish, Aliya Ahmad, Allen Hutchison, Alvin Abdagic, Amanda Carl, Amy Shen, Andy Brock, Andy Coenen, Anthony Laforge, Antonia Paterson, Ben Bastian, Bilal Piot, Bo Wu, Brandon Royal, Charlie Chen, Chintu Kumar, Chris Perry, Chris Welty, Christopher A. Choquette-Choo, Danila Sinopalnikov, David Weinberger, Dimple Vijaykumar, Dominika Rogozińska, Dustin Herbison, Elisa Bandy, Emma Wang, Eric Noland, Erica Moreira, Evan Senter, Evgenii Eltyshev, Francesco Visin, Gabriel Rasskin, Gary Wei, Glenn Cameron, Gus Martins, Hadi Hashemi, Hanna Klimczak-Plucińska, Harleen Batra, Harsh Dhand, Ivan Nardini, Jacinda Mein, Jack Zhou, James Svensson, Jeff Stanway, Jetha Chan, Jin Peng Zhou, Joana Carrasqueira, Joana Iljazi, Jocelyn Becker, Joe Fernandez, Joost van Amersfoort, Josh Gordon, Josh Lipschultz, Josh Newlan, Ju yeong Ji, Kareem Mohamed, Kartikeya Badola, Kat Black, Katie Millican, Keelin McDonell, Kelvin Nguyen, Kiranbir Sodhia, Kish Greene, Lars Lowe Sjoesund, Lauren Usui, Laurent Sifre, Lena Heuermann, Leticia Lago, Lilly McNealus, Livio Baldini Soares, Logan Kilpatrick, Lucas Dixon, Luciano Martins, Machel Reid, Manvinder Singh, Mark Iverson, Martin Görner, Mat Velloso, Mateo Wirth, Matt Davidow, Matt Miller, Matthew Rahtz, Matthew Watson, Meg Risdal, Mehran Kazemi, Michael Moynihan, Ming Zhang, Minsuk Kahng, Minwoo Park, Mofi Rahman, Mohit Khatwani, Natalie Dao, Nenshad Bardoliwalla, Nesh Devanathan, Neta Dumai, Nilay Chauhan, Oscar Wahltinez, Pankil Botarda, Parker Barnes, Paul Barham, Paul Michel, Pengchong Jin, Petko Georgiev, Phil Culliton, Pradeep Kuppala, Ramona Comanescu, Ramona Merhej, Reena Jana, Reza Ardeshir Rokni, Rishabh Agarwal, Ryan Mullins, Samaneh Saadat, Sara Mc Carthy, Sarah Perrin, Sébastien M. R. Arnold, Sebastian Krause, Shengyang Dai, Shruti Garg, Shruti Sheth, Sue Ronstrom, Susan Chan, Timothy Jordan, Ting Yu, Tom Eccles, Tom Hennigan, Tomas Kocisky, Tulsee Doshi, Vihan Jain, Vikas Yadav, Vilobh Meshram, Vishal Dharmadhikari, Warren Barkley, Wei Wei, Wenming Ye, Woohyun Han, Woosuk Kwon, Xiang Xu, Zhe Shen, Zhitao Gong, Zichuan Wei, Victor Cotruta, Phoebe Kirk, Anand Rao, Minh Giang, Ludovic Peran, Tris Warkentin, Eli Collins, Joelle Barral, Zoubin Ghahramani, Raia Hadsell, D. Sculley, Jeanine Banks, Anca Dragan, Slav Petrov, Oriol Vinyals, Jeff Dean, Demis Hassabis, Koray Kavukcuoglu, Clement Farabet, Elena Buchatskaya, Sebastian Borgeaud, Noah Fiedel, Armand Joulin, Kathleen Kenealy, Robert Dadashi, and Alek Andreev · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet, Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Closest in time.
Function vectors in large language models
Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau · 2024
Closest in time.