Fetching the paper…
Reading the bibliography…
Understanding and mitigating the potential risks associated with foundation models (FMs) hinges on developing effective interpretability methods.
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang. 2019 · 1911
Earlier work this paper cites.
Convergent and discriminant validation by the multitrait-multimethod matrix
Donald T Campbell and Donald W Fiske. 1959 · 1959
Earlier work this paper cites.
Matching pursuits with time-frequency dictionaries
Stéphane G Mallat and Zhifeng Zhang. 1993 · 1993
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Bruno A Olshausen and David J Field. 1996 · 1996
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Bruno A Olshausen and David J Field. 1997 · 1997
Earlier work this paper cites.
Causal mediation analysis for interpreting neural nlp: The case of gender bias
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Simas Sakenis, Jason Huang, Yaron Singer, and Stuart Shieber. 2020 · 2004
Earlier work this paper cites.
Compressed sensing
David L Donoho. 2006 · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov. 2006 · 2006
Earlier work this paper cites.
Sparse deep belief net model for visual area v2
Honglak Lee, Chaitanya Ekanadham, and Andrew Ng. 2007 · 2007
Earlier work this paper cites.
Tilted empirical risk minimization
Tian Li, Ahmad Beirami, Maziar Sanjabi, and Virginia Smith. 2020 · 2007
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson and Hugo Zaragoza. 2009 · 2009
Earlier work this paper cites.
Deep learning of representations: Looking forward
Yoshua Bengio. 2013 · 2013
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Quoc V Le. 2013 · 2013
Earlier work this paper cites.
Zero-bias autoencoders and the benefits of co-adapting features
Kishore Konda, Roland Memisevic, and David Krueger. 2014 · 2014
Earlier work this paper cites.
Sparse modeling for image and vision processing
Julien Mairal, Francis Bach, Jean Ponce, et al. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Decoupled weight decay regularization
I Loshchilov. 2017 · 2017
Earlier work this paper cites.
A characterization of guesswork on swiftly tilting curves
Ahmad Beirami, Robert Calderbank, Mark M Christiansen, Ken R Duffy, and Muriel Médard. 2018 · 2018
Earlier work this paper cites.
Isolating sources of disentanglement in variational autoencoders
Ricky TQ Chen, Xuechen Li, Roger B Grosse, and David K Duvenaud. 2018 · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. 2018 · 2018
Earlier work this paper cites.
Disentangling by factorising
Hyunjik Kim and Andriy Mnih. 2018 · 2018
Earlier work this paper cites.
Confounding variables can degrade generalization performance of radiological deep learning models
John R Zech, Marcus A Badgeley, Manway Liu, Anthony B Costa, Joseph J Titano, and Eric K Oermann. 2018 · 2018
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019 · 2019
Earlier work this paper cites.
Disentangling disentanglement in variational autoencoders
Emile Mathieu, Tom Rainforth, Nana Siddharth, and Yee Whye Teh. 2019 · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020 · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. 2020 · 2020
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan. 2020 · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021 · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. 2021 · 2021
Cited alongside, same era.
Towards a theoretical framework of out-of-distribution generalization
Haotian Ye, Chuanlong Xie, Tianle Cai, Ruichen Li, Zhenguo Li, and Liwei Wang. 2021 · 2021
Cited alongside, same era.
Zeyu Yun, Yubei Chen, Bruno A Olshausen, and Yann LeCun. 2021 · 2021
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov. 2022 · 2022
Representation engineering: A top-down approach to ai transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. 2023 · 2023
Later among the works it cites.
Examining language model performance with reconstructed activations using sparse autoencoders
Evan Anders and Joseph Bloom. 2024 · 2024
Closest in time.
arxiv physics dataset
Anonymous. 2024 · 2024
Closest in time.
Claude 3.5 sonnet
Anthropic. 2024 · 2024
Closest in time.
Gemma-2b-residual-stream-saes
John Bloom. 2024 · 2024
Closest in time.
Identifying functionally important features with end-to-end sparse dictionary learning
Dan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, and Lee Sharkey. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Softmax linear units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah. 2022a · 2022
Cited alongside, same era.
Unsupervised dense information retrieval with contrastive learning
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022 · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
The alignment problem from a deep learning perspective
Richard Ngo, Lawrence Chan, and Sören Mindermann. 2022 · 2022
Cited alongside, same era.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. 2022 · 2022
Cited alongside, same era.
Taking features out of superposition with sparse autoencoders
Lee Sharkey, Dan Braun, and Beren Millidge. 2022 · 2022
Cited alongside, same era.
Taking features out of superposition with sparse autoencoders, 2022
Lee Sharkey, Dan Braun, and Beren Millidge. 2023 · 2022
Cited alongside, same era.
Closest in time.
Experiments with an alternative method to promote sparsity in sparse autoencoders
Eoin Farrell. 2024 · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. 2024 · 2024
Closest in time.
Openwebtext corpus
Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex. 2019 · 2024
Closest in time.
arxiv physics instruct tune 30k dataset
Algorithmic Research Group. 2024 · 2024
Closest in time.
The unreasonable effectiveness of easy training data for hard tasks
Peter Hase, Mohit Bansal, Peter Clark, and Sarah Wiegreffe. 2024 · 2024
Closest in time.
Tanh penalty in dictionary learning
Adam Jermyn, Adly Templeton, Joshua Batson, and Trenton Bricken. 2024 · 2024
Closest in time.
Understanding sae features with the logit lens
Johnny Lin Joseph Bloom. 2024 · 2024
Closest in time.
Open source automated interpretability for sparse autoencoder features
Caden Juang, Gonçalo Paulo, Jacob Drori, and Nora Belrose. 2024 · 2024
Closest in time.
Pile toxicity balanced dataset
Tomek Korbak. 2024 · 2024
Closest in time.
Atp*: An efficient and scalable method for localizing llm behaviour to components
János Kramár, Tom Lieberum, Rohin Shah, and Neel Nanda. 2024 · 2024
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2024 · 2024
Closest in time.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2
Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca Dragan, Rohin Shah, and Neel Nanda. 2024 · 2024
Closest in time.
Towards principled evaluations of sparse autoencoders for interpretability and control
Aleksandar Makelov, George Lange, and Neel Nanda. 2024 · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller. 2024 · 2024
Closest in time.
Transformer debugger
Dan Mossing, Steven Bills, Henk Tillman, Tom Dupré la Tour, Nick Cammarata, Leo Gao, Joshua Achiam, Catherine Yeh, Jan Leike, Jeff Wu, et al. 2024 · 2024
Closest in time.
Disentangling dense embeddings with sparse autoencoders
Charles O’Neill, Christine Ye, Kartheik Iyer, and John F Wu. 2024 · 2024
Closest in time.
Improving dictionary learning with gated sparse autoencoders
Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, János Kramár, Rohin Shah, and Neel Nanda. 2024 · 2024
Closest in time.
Improving sae’s by sqrt()-ing l1 and removing lowest activating features
Logan Riggs and Jannik Brinkmann. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. 2024 · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, Alex Tamkin, Esin Durmus, Tristan Hume, Francesco Mosconi, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan. 2024 · 2024
Closest in time.
Addressing feature suppression in saes
Benjamin Wright and Lee Sharkey. 2024 · 2024
Closest in time.