Fetching the paper…
Reading the bibliography…
As large language models (LLMs) see wider real-world use, understanding and mitigating their unsafe behaviors is critical.
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al. 2019 · 1905
Earlier work this paper cites.
Relatif: Identifying explanatory training samples via relative influence
Elnaz Barshan, Marc-Etienne Brunet, and Gintare Karolina Dziugaite. 2020 · 1909
Earlier work this paper cites.
The influence curve and its role in robust estimation
Frank R Hampel. 1974 · 1974
Earlier work this paper cites.
Characterizations of an empirical influence function for detecting influential cases in regression
R Dennis Cook and Sanford Weisberg. 1980 · 1980
Earlier work this paper cites.
Direct and indirect effects
Judea Pearl. 2001 · 2001
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
Supervised dictionary learning
Julien Mairal, Jean Ponce, Guillermo Sapiro, Andrew Zisserman, and Francis Bach. 2008 · 2008
Earlier work this paper cites.
Captum: A unified and generic model interpretability library for pytorch
Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, et al. 2020 · 2009
Earlier work this paper cites.
Alireza Makhzani and Brendan Frey. 2013 · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomáš Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013 · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek. 2015 · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. 2018 · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Representer point selection for explaining deep neural networks
Chih-Kuan Yeh, Joon Kim, Ian En-Hsu Yen, and Pradeep K Ravikumar. 2018 · 2018
Earlier work this paper cites.
Neural network attributions: a causal perspective
Aditya Chattopadhyay, Piyushi Manupriya, Anirban Sarkar, and Vineeth N Balasubramanian. 2019 · 2019
Earlier work this paper cites.
What is one grain of sand in the desert? analyzing individual neurons in deep nlp models
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass. 2019 · 2019
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou. 2019 · 2019
Earlier work this paper cites.
Visual analytics in deep learning: An interrogative survey for the next frontiers
Fred Hohman, Minsuk Kahng, Robert Pienta, and Duen Horng Chau. 2019 · 2019
Earlier work this paper cites.
Towards efficient data valuation based on the shapley value
Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nick Hynes, Nezihe Merve Gürel, Bo Li, Ce Zhang, Dawn Song, and Costas J Spanos. 2019 · 2019
Earlier work this paper cites.
Sanvis: Visual analytics for understanding self-attention networks
Cheonbok Park, Inyoup Na, Yongjang Jo, Sungbok Shin, Jaehyo Yoo, Bum Chul Kwon, Jian Zhao, Hyungjong Noh, Yeonsoo Lee, and Jaegul Choo. 2019 · 2019
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Earlier work this paper cites.
Seq2seq-vis: A visual debugging tool for sequence-to-sequence models
Hendrik Strobelt, Sebastian Gehrmann, Michael Behrisch, Adam Perer, Hanspeter Pfister, and Alexander M. Rush. 2019 · 2019
Earlier work this paper cites.
A multiscale visualization of attention in the transformer model
Jesse Vig. 2019 · 2019
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Earlier work this paper cites.
Quantifying attention flow in transformers
Samira Abnar and Willem Zuidema. 2020 · 2020
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
Vitaly Feldman and Chiyuan Zhang. 2020 · 2020
Earlier work this paper cites.
Explaining black box predictions and unveiling data artifacts through influence functions
Xiaochuang Han, Byron C. Wallace, and Yulia Tsvetkov. 2020 · 2020
Earlier work this paper cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Earlier work this paper cites.
Interpreting interpretability: Understanding data scientists’ use of interpretability tools for machine learning
Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna Wallach, and Jennifer Wortman Vaughan. 2020 · 2020
Earlier work this paper cites.
Attention is not only a weight: Analyzing transformers with vector norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020 · 2020
Earlier work this paper cites.
Interpreting gpt: The logit lens
nostalgebraist. 2020 · 2020
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan. 2020 · 2020
Earlier work this paper cites.
Not all unlabeled data are equal: Learning to weight data in semi-supervised learning
Zhongzheng Ren, Raymond Yeh, and Alexander Schwing. 2020 · 2020
Earlier work this paper cites.
The language interpretability tool: Extensible, interactive visualizations and analysis for NLP models
Ian Tenney, James Wexler, Jasmijn Bastings, Tolga Bolukbasi, Andy Coenen, Sebastian Gehrmann, Ellen Jiang, Mahima Pushkarna, Carey Radebaugh, Emily Reif, and Ann Yuan. 2020 · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020 · 2020
Earlier work this paper cites.
Influence functions in deep learning are fragile
Samyadeep Basu, Phil Pope, and Soheil Feizi. 2021 · 2021
Earlier work this paper cites.
The role of interactive visualization in fostering trust in ai
Emma Beauxis-Aussalet, Michael Behrisch, Rita Borgo, Duen Horng Chau, Christopher Collins, David Ebert, Mennatallah El-Assady, Alex Endert, Daniel A Keim, Jörn Kohlhammer, et al. 2021 · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021 · 2021
Earlier work this paper cites.
Explaining by removing: A unified framework for model explanation
Ian Covert, Scott Lundberg, and Su-In Lee. 2021 · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. 2021 · 2021
Earlier work this paper cites.
Causal analysis of syntactic agreement mechanisms in neural language models
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov. 2021 · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. 2021 · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Earlier work this paper cites.
FastIF: Scalable influence functions for efficient model interpretation and debugging
Han Guo, Nazneen Rajani, Peter Hase, Mohit Bansal, and Caiming Xiong. 2021 · 2021
Earlier work this paper cites.
Influence tuning: Demoting spurious correlations via instance attribution and instance-driven updates
Xiaochuang Han and Yulia Tsvetkov. 2021 · 2021
Earlier work this paper cites.
Contrastive explanations for model interpretability
Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi, and Yoav Goldberg. 2021 · 2021
Earlier work this paper cites.
Incorporating Residual and Normalization Layers into Analysis of Masked Language Models
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2021 · 2021
Earlier work this paper cites.
Keywordmap: Attention-based visual exploration for keyword analysis
Yamei Tu, Jiayi Xu, and Han-Wei Shen. 2021 · 2021
Earlier work this paper cites.
Dodrio: Exploring transformer models with interactive visualization
Zijie J. Wang, Robert Turko, and Duen Horng Chau. 2021 · 2021
Earlier work this paper cites.
Towards tracing knowledge in language models back to the training data
Ekin Akyurek, Tolga Bolukbasi, Frederick Liu, Binbin Xiong, Ian Tenney, Jacob Andreas, and Kelvin Guu. 2022 · 2022
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov. 2022 · 2022
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2022 · 2022
Earlier work this paper cites.
Faithful reasoning using large language models
Antonia Creswell and Murray Shanahan. 2022 · 2022
Earlier work this paper cites.
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, et al. 2022 · 2022
Earlier work this paper cites.
Towards opening the black box of neural machine translation: Source and target interpretations of the transformer
Javier Ferrando, Gerard I. Gállego, Belen Alastruey, Carlos Escolano, and Marta R. Costa-jussà. 2022a · 2022
Earlier work this paper cites.
Measuring the mixing of contextual information in the transformer
Javier Ferrando, Gerard I. Gállego, and Marta R. Costa-jussà. 2022b · 2022
Earlier work this paper cites.
LM-debugger: An interactive tool for inspection and intervention in transformer-based language models
Mor Geva, Avi Caciularu, Guy Dar, Paul Roit, Shoval Sadde, Micah Shlain, Bar Tamir, and Yoav Goldberg. 2022a · 2022
Earlier work this paper cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022b · 2022
Earlier work this paper cites.
Xiaochuang Han and Yulia Tsvetkov. 2022 · 2022
Earlier work this paper cites.
Datamodels: Predicting predictions from training data
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry. 2022 · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022 · 2022
Earlier work this paper cites.
Resolving training biases via influence-based data relabeling
Shuming Kong, Yanyan Shen, and Linpeng Huang. 2022 · 2022
Earlier work this paper cites.
Rethinking explainability as a dialogue: A practitioner’s perspective
Himabindu Lakkaraju, Dylan Slack, Yuxin Chen, Chenhao Tan, and Sameer Singh. 2022 · 2022
Earlier work this paper cites.
Post-hoc interpretability for neural nlp: A survey
Andreas Madsen, Siva Reddy, and Sarath Chandar. 2022 · 2022
Earlier work this paper cites.
Few-shot self-rationalization with natural language prompts
Ana Marasovic, Iz Beltagy, Doug Downey, and Matthew Peters. 2022 · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov. 2022 · 2022
Earlier work this paper cites.
GlobEnc: Quantifying global token attribution by incorporating the whole encoder layer in transformers
Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, and Mohammad Taher Pilehvar. 2022 · 2022
Earlier work this paper cites.
Neuroscope: A website for mechanistic interpretability of language models
Neel Nanda. 2022 · 2022
Earlier work this paper cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2022 · 2022
Earlier work this paper cites.
Interpretable proof generation via iterative backward reasoning
Hanhao Qu, Yu Cao, Jun Gao, Liang Ding, and Ruifeng Xu. 2022 · 2022
Earlier work this paper cites.
Scaling up influence functions
Andrea Schioppa, Polina Zablotskaia, David Vilar, and Artem Sokolov. 2022 · 2022
Earlier work this paper cites.
Taking features out of superposition with sparse autoencoders
Lee Sharkey, Dan Braun, and Beren Millidge. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022 · 2022
Earlier work this paper cites.
Davinz: Data valuation using deep neural networks at initialization
Zhaoxuan Wu, Yao Shu, and Bryan Kian Hsiang Low. 2022 · 2022
Earlier work this paper cites.
The unreliability of explanations in few-shot prompting for textual reasoning
Xi Ye and Greg Durrett. 2022 · 2022
Earlier work this paper cites.
First is better than last for language data influence
Chih-Kuan Yeh, Ankur Taly, Mukund Sundararajan, Frederick Liu, and Pradeep Ravikumar. 2022 · 2022
Earlier work this paper cites.
Interpreting language models with contrastive explanations
Kayo Yin and Graham Neubig. 2022 · 2022
Earlier work this paper cites.
Legal prompting: Teaching a language model to think like a lawyer
Fangyi Yu, Lee Quartey, and Frank Schilder. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Exploring evaluation methods for interpretable machine learning: A survey
Nourah Alangari, Mohamed El Bachir Menai, Hassan Mathkour, and Ibrahim Almosallam. 2023 · 2023
Earlier work this paper cites.
The internal state of an llm knows when it’s lying
Amos Azaria and Tom Mitchell. 2023 · 2023
Earlier work this paper cites.
Eliciting latent predictions from transformers with the tuned lens
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023 · 2023
Earlier work this paper cites.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. 2023 · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Chris Olah. 2023 · 2023
Earlier work this paper cites.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. 2023 · 2023
Earlier work this paper cites.
Selection-inference: Exploiting large language models for interpretable logical reasoning
Antonia Creswell, Murray Shanahan, and Irina Higgins. 2023 · 2023
Earlier work this paper cites.
Detecting and mitigating hallucinations in machine translation: Model internal workings alone do well, sentence similarity Even better
David Dale, Elena Voita, Loic Barrault, and Marta R. Costa-jussà. 2023 · 2023
Earlier work this paper cites.
Analyzing transformers in embedding space
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. 2023 · 2023
Earlier work this paper cites.
Discovering variable binding circuitry with desiderata
Xander Davies, Max Nadeau, Nikhil Prakash, Tamar Rott Shaham, and David Bau. 2023 · 2023
Earlier work this paper cites.
Atman: understanding transformer predictions through memory efficient attention manipulation
Björn Deiseroth, Mayukh Deb, Samuel Weinbach, Manuel Brack, Patrick Schramowski, and Kristian Kersting. 2023 · 2023
Earlier work this paper cites.
Jump to conclusions: Short-cutting transformers with linear transformations
Alexander Yom Din, Taelin Karidi, Leshem Choshen, and Mor Geva. 2023 · 2023
Earlier work this paper cites.
Rather a nurse than a physician - contrastive explanations under investigation
Oliver Eberle, Ilias Chalkidis, Laura Cabello, and Stephanie Brandl. 2023 · 2023
Earlier work this paper cites.
Sequential integrated gradients: a simple but effective method for explaining language models
Joseph Enguehard. 2023 · 2023
Earlier work this paper cites.
Explaining how transformers use context to build predictions
Javier Ferrando, Gerard I. Gállego, Ioannis Tsiamas, and Marta R. Costa-jussà. 2023 · 2023
Earlier work this paper cites.
Neuron to graph: Interpreting language model neurons at scale
Alex Foote, Neel Nanda, Esben Kran, Ioannis Konstas, Shay Cohen, and Fazl Barez. 2023 · 2023
Cited alongside, same era.
Is ChatGPT a good causal reasoner? a comprehensive evaluation
Jinglong Gao, Xiao Ding, Bing Qin, and Ting Liu. 2023 · 2023
Cited alongside, same era.
Deepdecipher: Accessing and investigating neuron activation in large language models
Albert Garde, Esben Kran, and Fazl Barez. 2023 · 2023
Cited alongside, same era.
Dissecting recall of factual associations in auto-regressive language models
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023 · 2023
Cited alongside, same era.
Localizing model behavior with path patching
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora. 2023 · 2023
Interpreting attention layer outputs with sparse autoencoders
Connor Kissane, Robert Krzyzanowski, Joseph Isaac Bloom, Arthur Conmy, and Neel Nanda. 2024 · 2024
Later among the works it cites.
Analyzing feed-forward blocks in transformers through the lens of attention maps
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2024 · 2024
Later among the works it cites.
Atp*: An efficient and scalable method for localizing llm behaviour to components
János Kramár, Tom Lieberum, Rohin Shah, and Neel Nanda. 2024 · 2024
Later among the works it cites.
Datainf: Efficiently estimating data influence in loRA-tuned LLMs and diffusion models
Yongchan Kwon, Eric Wu, Kevin Wu, and James Zou. 2024 · 2024
Later among the works it cites.
A mechanistic understanding of alignment algorithms: a case study on dpo and toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, and Rada Mihalcea. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Studying large language model generalization with influence functions
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, et al. 2023 · 2023
Cited alongside, same era.
Finding neurons in a haystack: Case studies with sparse probing
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. 2023 · 2023
Cited alongside, same era.
Simfluence: Modeling the influence of individual training examples by simulating training runs
Kelvin Guu, Albert Webson, Ellie Pavlick, Lucas Dixon, Ian Tenney, and Tolga Bolukbasi. 2023 · 2023
Cited alongside, same era.
How does gpt-2 compute greater-than? interpreting mathematical abilities in a pre-trained language model
Michael Hanna, Ollie Liu, and Alexandre Variengien. 2023 · 2023
Cited alongside, same era.
Understanding transformer memorization recall through idioms
Adi Haviv, Ido Cohen, Jacob Gidron, Roei Schuster, Yoav Goldberg, and Mor Geva. 2023 · 2023
Cited alongside, same era.
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen-Chuan Chang. 2023 · 2023
Cited alongside, same era.
VISIT: Visualizing and interpreting the semantic information flow of transformers
Shahar Katz and Yonatan Belinkov. 2023 · 2023
Cited alongside, same era.
No two devils alike: Unveiling distinct mechanisms of fine-tuning attacks
Chak Tou Leong, Yi Cheng, Kaishuai Xu, Jian Wang, Hanlin Wang, and Wenjie Li. 2024 · 2024
Later among the works it cites.
Still no lie detector for language models: Probing empirical and conceptual roadblocks
Benjamin A Levinstein and Daniel A Herrmann. 2024 · 2024
Later among the works it cites.
Evaluating readability and faithfulness of concept-based explanations
Meng Li, Haoran Jin, Ruixuan Huang, Zhihao Xu, Defu Lian, Zijia Lin, Di Zhang, and Xiting Wang. 2024c · 2024
Later among the works it cites.
AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap
Q. Vera Liao and Jennifer Wortman Vaughan. 2024 · 2024
Later among the works it cites.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2
Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca Dragan, Rohin Shah, and Neel Nanda. 2024 · 2024
Later among the works it cites.
Towards understanding jailbreak attacks in LLMs: A representation space analysis
Yuping Lin, Pengfei He, Han Xu, Yue Xing, Makoto Yamada, Hui Liu, and Jiliang Tang. 2024 · 2024
Later among the works it cites.
Sparse crosscoders for cross-layer features and model diffing
Jack Lindsey, Adly Templeton, Jonathan Marcus, Thomas Conerly, Joshua Batson, and Christopher Olah. 2024 · 2024
Later among the works it cites.
On the universal truthfulness hyperplane inside LLMs
Junteng Liu, Shiqi Chen, Yu Cheng, and Junxian He. 2024a · 2024
Later among the works it cites.
Universal response and emergence of induction in llms
Niclas Luick. 2024 · 2024
Later among the works it cites.
Mechanistic insights: Circuit transformations across input and fine-tuning landscapes
Chiyu Ma, Lin Shi, Ollie Liu, Wenhua Liang, Jiaqi Gan, Ming Cheng, Willie Neiswanger, and Soroush Vosoughi. 2024 · 2024
Later among the works it cites.
Towards principled evaluations of sparse autoencoders for interpretability and control
Aleksandar Makelov, George Lange, and Neel Nanda. 2024 · 2024
Later among the works it cites.
What is ai interpretability?
Amanda McGrath and Alexandra Jonker. 2024 · 2024
Later among the works it cites.
Selfcheck: Using LLMs to zero-shot check their own step-by-step reasoning
Ning Miao, Yee Whye Teh, and Tom Rainforth. 2024 · 2024
Later among the works it cites.
Explaining large language models decisions using shapley values
Behnam Mohammadi. 2024 · 2024
Later among the works it cites.
A glitch in the matrix? locating and detecting language model grounding with fakepedia
Giovanni Monea, Maxime Peyrard, Martin Josifoski, Vishrav Chaudhary, Jason Eisner, Emre Kiciman, Hamid Palangi, Barun Patra, and Robert West. 2024 · 2024
Later among the works it cites.
Reasoning beyond bias: A study on counterfactual prompting and chain of thought reasoning
Kyle Moore, Jesse Roberts, Thao Pham, and Douglas Fisher. 2024 · 2024
Later among the works it cites.
Attribution patching: Activation patching at industrial scale
Neel Nanda. 2024 · 2024
Later among the works it cites.
Steering language model refusal with sparse autoencoders
Kyle O’Brien, David Majercak, Xavier Fernandes, Richard Edgar, Jingya Chen, Harsha Nori, Dean Carignan, Eric Horvitz, and Forough Poursabzi-Sangde. 2024 · 2024
Later among the works it cites.
Disentangling dense embeddings with sparse autoencoders
Charles O’Neill, Christine Ye, Kartheik Iyer, and John F Wu. 2024 · 2024
Later among the works it cites.
Llms know more than they show: On the intrinsic representation of llm hallucinations
Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart, Idan Szpektor, Hadas Kotek, and Yonatan Belinkov. 2024 · 2024
Later among the works it cites.
Automatically interpreting millions of features in large language models
Gonçalo Paulo, Alex Mallen, Caden Juang, and Nora Belrose. 2024 · 2024
Later among the works it cites.
Navigating the safety landscape: Measuring risks in finetuning large language models
Sheng Y Peng, Pin-Yu Chen, Matthew Hull, and Duen H Chau. 2024 · 2024
Later among the works it cites.
Significance of chain of thought in gender bias mitigation for english-dravidian machine translation
Lavanya Prahallad and Radhika Mamidi. 2024 · 2024
Later among the works it cites.
Fine-tuning enhances existing mechanisms: A case study on entity tracking
Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau. 2024b · 2024
Later among the works it cites.
Model internals-based answer attribution for trustworthy retrieval-augmented generation
Jirui Qi, Gabriele Sarti, Raquel Fernández, and Arianna Bisazza. 2024 · 2024
Later among the works it cites.
Towards tracing trustworthiness dynamics: Revisiting pre-training period of large language models
Chen Qian, Jie Zhang, Wei Yao, Dongrui Liu, Zhenfei Yin, Yu Qiao, Yong Liu, and Jing Shao. 2024 · 2024
Later among the works it cites.
Mambalrp: Explaining selective state space sequence models
Farnoush Rezaei Jafari, Grégoire Montavon, Klaus-Robert Müller, and Oliver Eberle. 2024 · 2024
Later among the works it cites.
Steering llama 2 via contrastive activation addition
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner. 2024 · 2024
Later among the works it cites.
Llm potentiality and awareness: a position paper from the perspective of trustworthy and responsible ai modeling
Iqbal H Sarker. 2024 · 2024
Later among the works it cites.
Quantifying the plausibility of context reliance in neural machine translation
Gabriele Sarti, Grzegorz Chrupała, Malvina Nissim, and Arianna Bisazza. 2024 · 2024
Later among the works it cites.
Guobin Shen, Dongcheng Zhao, Yiting Dong, Xiang He, and Yi Zeng. 2024 · 2024
Later among the works it cites.
Eliciting uncertainty in chain-of-thought to mitigate bias against forecasting harmful user behaviors
Anthony Sicilia and Malihe Alikhani. 2024 · 2024
Later among the works it cites.
Enhancing adversarial attacks through chain of thought
Jingbo Su. 2024 · 2024
Later among the works it cites.
Attribution patching outperforms automated circuit discovery
Aaquib Syed, Can Rager, and Arthur Conmy. 2024 · 2024
Later among the works it cites.
Interpreting pretrained language models via concept bottlenecks
Zhen Tan, Lu Cheng, Song Wang, Bo Yuan, Jundong Li, and Huan Liu. 2024 · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, Alex Tamkin, Esin Durmus, Tristan Hume, Francesco Mosconi, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan. 2024 · 2024
Later among the works it cites.
Assessing the brittleness of safety alignment via pruning and low-rank modifications
Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang, and Peter Henderson. 2024 · 2024
Later among the works it cites.
Enhancing training data attribution for large language models with fitting error consideration
Kangxi Wu, Liang Pang, Huawei Shen, and Xueqi Cheng. 2024a · 2024
Later among the works it cites.
LESS: Selecting influential data for targeted instruction tuning
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. 2024 · 2024
Later among the works it cites.
Defensive prompt patch: A robust and interpretable defense of llms against jailbreak attacks
Chen Xiong, Xiangyu Qi, Pin-Yu Chen, and Tsung-Yi Ho. 2024 · 2024
Later among the works it cites.
SaySelf: Teaching LLMs to express confidence with self-reflective rationales
Tianyang Xu, Shujin Wu, Shizhe Diao, Xiaoze Liu, Xingyao Wang, Yangyi Chen, and Jing Gao. 2024a · 2024
Later among the works it cites.
Enhancing semantic consistency of large language models through model editing: An interpretability-oriented approach
Jingyuan Yang, Dapeng Chen, Yajing Sun, Rongjun Li, Zhiyong Feng, and Wei Peng. 2024a · 2024
Later among the works it cites.
Attentionviz: A global view of transformer attention
Catherine Yeh, Yida Chen, Aoyu Wu, Cynthia Chen, Fernanda Viégas, and Martin Wattenberg. 2024 · 2024
Later among the works it cites.
Mechanistic understanding and mitigation of language model non-factual hallucinations
Lei Yu, Meng Cao, Jackie CK Cheung, and Yue Dong. 2024b · 2024
Later among the works it cites.
Attention satisfies: A constraint-satisfaction lens on factual errors of language models
Mert Yuksekgonul, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar, Ranjita Naik, Hamid Palangi, Ece Kamar, and Besmira Nushi. 2024 · 2024
Later among the works it cites.
Defending large language models against jailbreak attacks via layer-specific editing
Wei Zhao, Zhe Li, Yige Li, Ye Zhang, and Jun Sun. 2024d · 2024
Later among the works it cites.
How alignment and jailbreak work: Explain LLM safety through intermediate hidden states
Zhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu, Fei Huang, and Yongbin Li. 2024 · 2024
Later among the works it cites.
Locking down the finetuned llms safety
Minjun Zhu, Linyi Yang, Yifan Wei, Ningyu Zhang, and Yue Zhang. 2024 · 2024
Later among the works it cites.
Samir Abdaljalil, Filippo Pallucchini, Andrea Seveso, Hasan Kurban, Fabio Mercorio, and Erchin Serpedin. 2025 · 2025
Closest in time.
Circuit tracing: Revealing computational graphs in language models
Emmanuel Ameisen, Jack Lindsey, Adam Pearce, Wes Gurnee, Nicholas L. Turner, Brian Chen, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum McDougall, Hoagy Cunningham, Thomas Henighan, Adam Jermyn, Andy Jones, Andrew Persic, Zhenyi Qi, T. Ben Thompson, Sam Zimmerman, Kelley Rivoire, Thomas Conerly, Chris Olah, and Joshua Batson. 2025 · 2025
Closest in time.
Roberto Araya. 2025 · 2025
Closest in time.
A close look at decomposition-based xai-methods for transformer language models
Leila Arras, Bruno Puri, Patrick Kahardipraja, Sebastian Lapuschkin, and Wojciech Samek. 2025 · 2025
Closest in time.
Language models can predict their own behavior
Dhananjay Ashok and Jonathan May. 2025 · 2025
Closest in time.
Mechanistic permutability: Match features across layers
Nikita Balagansky, Ian Maksimov, and Daniil Gavrilov. 2025 · 2025
Closest in time.
Steering large language model activations in sparse spaces
Reza Bayat, Ali Rahimi-Kalahroudi, Mohammad Pezeshki, Sarath Chandar, and Pascal Vincent. 2025 · 2025
Closest in time.
Tell me about yourself: Llms are aware of their learned behaviors
Jan Betley, Xuchan Bao, Martín Soto, Anna Sztyber-Betley, James Chua, and Owain Evans. 2025 · 2025
Closest in time.
Looking inward: Language models can learn about themselves by introspection
Felix Jedidja Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, and Owain Evans. 2025 · 2025
Closest in time.
Reasoning-grounded natural language explanations for language models
Vojtech Cahlik, Rodrigo Alves, and Pavel Kordik. 2025 · 2025
Closest in time.
On behalf of the stakeholders: Trends in NLP model interpretability in the era of LLMs
Nitay Calderon and Roi Reichart. 2025 · 2025
Closest in time.
Scalable influence and fact tracing for large language model pretraining
Tyler A. Chang, Dheeraj Rajagopal, Tolga Bolukbasi, Lucas Dixon, and Ian Tenney. 2025 · 2025
Closest in time.
Attributive reasoning for hallucination diagnosis of large language models
Yuyan Chen, Zehao Li, Shuangjie You, Zhengyu Chen, Jingwen Chang, Yi Zhang, Weinan Dai, Qingpei Guo, and Yanghua Xiao. 2025 · 2025
Closest in time.
Domaino1s: Guiding llm reasoning for explainable answers in high-stakes domains
Xu Chu, Zhijie Tan, Hanlin Xue, Guanyu Wang, Tong Mo, and Weiping Li. 2025 · 2025
Closest in time.
Cram: Credibility-aware attention modification in llms for combating misinformation in rag
Boyi Deng, Wenjie Wang, Fengbin Zhu, Qifan Wang, and Fuli Feng. 2025 · 2025
Closest in time.
Jacobian sparse autoencoders: Sparsify computations, not just activations
Lucy Farnik, Tim Lawson, Conor Houghton, and Laurence Aitchison. 2025 · 2025
Closest in time.
Do i know this entity? knowledge awareness and hallucinations in language models
Javier Ferrando, Oscar Balcells Obeso, Senthooran Rajamanoharan, and Neel Nanda. 2025 · 2025
Closest in time.
Ahmed Frikha, Muhammad Reza Ar Razi, Krishna Kanth Nakka, Ricardo Mendes, Xue Jiang, and Xuebing Zhou. 2025 · 2025
Closest in time.
Sparse autoencoder features for classifications and transferability
Jack Gallifant, Shan Chen, Kuleen Sasse, Hugo Aerts, Thomas Hartvigsen, and Danielle S Bitterman. 2025 · 2025
Closest in time.
Internal activation as the polar star for steering unsafe llm behavior
Peixuan Han, Cheng Qian, Xiusi Chen, Yuji Zhang, Denghui Zhang, and Heng Ji. 2025 · 2025
Closest in time.
Towards llm guardrails via sparse representation steering
Zeqing He, Zhibo Wang, Huiyu Xu, and Kui Ren. 2025 · 2025
Closest in time.
Refusal behavior in large language models: A nonlinear perspective
Fabian Hildebrandt, Andreas Maier, Patrick Krauss, and Achim Schilling. 2025 · 2025
Closest in time.
How llms learn: Tracing internal representations with sparse autoencoders
Tatsuro Inaba, Kentaro Inui, Yusuke Miyao, Yohei Oseki, Benjamin Heinzerling, and Yu Takagi. 2025 · 2025
Closest in time.
Comt: Chain-of-medical-thought reduces hallucination in medical report generation
Yue Jiang, Jiawei Chen, Dingkang Yang, Mingcheng Li, Shunli Wang, Tong Wu, Ke Li, and Lihua Zhang. 2025 · 2025
Closest in time.
Don’t forget it! conditional sparse autoencoder clamping works for unlearning
Matthew Khoriaty, Andrii Shportko, Gustavo Mercier, and Zach Wood-Doughty. 2025 · 2025
Closest in time.
Mixhd: A method for detecting hallucinations based on the internal state and output probability of large language models
Chuang Li, Bingnan Xing, Dongdong Huo, Qihui Zhou, Zhen Xu, and Yu Wang. 2025a · 2025
Closest in time.
On the biology of a large language model
Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum McDougall, Hoagy Cunningham, Thomas Henighan, Adam Jermyn, Andy Jones, Andrew Persic, Zhenyi Qi, T. Ben Thompson, Sam Zimmerman, Kelley Rivoire, Thomas Conerly, Chris Olah, and Joshua Batson. 2025 · 2025
Closest in time.
Safety at scale: A comprehensive survey of large model safety
Xingjun Ma, Yifeng Gao, Yixu Wang, Ruofan Wang, Xin Wang, Ye Sun, Yifan Ding, Hengyuan Xu, Yunhao Chen, Yunhan Zhao, Hanxun Huang, Yige Li, Jiaming Zhang, Xiang Zheng, Yang Bai, Zuxuan Wu, Xipeng Qiu, Jingfeng Zhang, Yiming Li, Xudong Han, Haonan Li, Jun Sun, Cong Wang, Jindong Gu, Baoyuan Wu, Siheng Chen, Tianwei Zhang, Yang Liu, Mingming Gong, Tongliang Liu, Shirui Pan, Cihang Xie, Tianyu Pang, Yinpeng Dong, Ruoxi Jia, Yang Zhang, Shiqing Ma, Xiangyu Zhang, Neil Gong, Chaowei Xiao, Sarah Erfani, Tim Baldwin, Bo Li, Masashi Sugiyama, Dacheng Tao, James Bailey, and Yu-Gang Jiang. 2025 · 2025
Closest in time.
Noiser: Bounded input perturbations for attributing large language models
Mohammad Reza Ghasemi Madani, Aryo Pradipta Gema, Gabriele Sarti, Yu Zhao, Pasquale Minervini, and Andrea Passerini. 2025 · 2025
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller. 2025 · 2025
Closest in time.
Promptaid: Visual prompt exploration, perturbation, testing and iteration for large language models
Aditi Mishra, Bretho Danzy, Utkarsh Soni, Anjana Arunkumar, Jinbin Huang, Bum Chul Kwon, and Chris Bryan. 2025 · 2025
Closest in time.
Saro: Enhancing llm safety through reasoning-based alignment
Yutao Mou, Yuxiao Luo, Shikun Zhang, and Wei Ye. 2025 · 2025
Closest in time.
Efficient dictionary learning with switch sparse autoencoders
Anish Mudide, Joshua Engels, Eric J Michaud, Max Tegmark, and Christian Schroeder de Witt. 2025 · 2025
Closest in time.
Decoding dark matter: Specialized sparse autoencoders for interpreting rare concepts in foundation models
Aashiq Muhamed, Mona T. Diab, and Virginia Smith. 2025 · 2025
Closest in time.
Melissa Kazemi Rad, Huy Nghiem, Andy Luo, Sahil Wadhwa, Mohammad Sorower, and Stephen Rawls. 2025 · 2025
Closest in time.
Think or step-by-step? unzipping the black box in zero-shot prompts
Nikta Gohari Sadr, Sangmitra Madhusudan, and Ali Emami. 2025 · 2025
Closest in time.
Open problems in mechanistic interpretability
Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeff Wu, Lucius Bushnaq, Nicholas Goldowsky-Dill, Stefan Heimersheim, Alejandro Ortega, Joseph Bloom, Stella Biderman, Adria Garriga-Alonso, Arthur Conmy, Neel Nanda, Jessica Rumbelow, Martin Wattenberg, Nandi Schoots, Joseph Miller, Eric J. Michaud, Stephen Casper, Max Tegmark, William Saunders, David Bau, Eric Todd, Atticus Geiger, Mor Geva, Jesse Hoogland, Daniel Murfet, and Tom McGrath. 2025 · 2025
Closest in time.
Route sparse autoencoder to interpret large language models
Wei Shi, Sihang Li, Tao Liang, Mingyang Wan, Guojun Ma, Xiang Wang, and Xiangnan He. 2025 · 2025
Closest in time.
Charlotte Siska and Anush Sankaran. 2025 · 2025
Closest in time.
Interpretable steering of large language models with feature guided activation additions
Samuel Soo, Wesley Teng, Chandrasekaran Balaganesh, Tan Guoxian, and Ming YAN. 2025 · 2025
Closest in time.
Concept bottleneck large language models
Chung-En Sun, Tuomas Oikarinen, Berk Ustun, and Tsui-Wei Weng. 2025 · 2025
Closest in time.
Xue Tan, Hao Luan, Mingyu Luo, Xiaoyan Sun, Ping Chen, and Jun Dai. 2025 · 2025
Closest in time.
An investigation of large language models and their vulnerabilities in spam detection
Qiyao Tang and Xiangyang Li. 2025 · 2025
Closest in time.
Finding sparse autoencoder representations of errors in cot prompting
Justin Theodorus, V Swaytha, Shivani Gautam, Adam Ward, Mahir Shah, Cole Blondin, and Kevin Zhu. 2025 · 2025
Closest in time.
Large language models: A comprehensive survey on architectures, applications, and challenges
Vinod Veeramachaneni. 2025 · 2025
Closest in time.
Using mechanistic interpretability to craft adversarial attacks against large language models
Thomas Winninger, Boussad Addad, and Katarzyna Kapusta. 2025 · 2025
Closest in time.
Interpreting and steering llms with mutual information-based explanations on sparse autoencoders
Xuansheng Wu, Jiayi Yuan, Wenlin Yao, Xiaoming Zhai, and Ninghao Liu. 2025 · 2025
Closest in time.
Safety is not only about refusal: Reasoning-enhanced fine-tuning for interpretable llm safety
Yuyou Zhang, Miao Li, William Han, Yihang Yao, Zhepeng Cen, and Ding Zhao. 2025 · 2025
Closest in time.
LLM neurosurgeon: Targeted knowledge removal in LLMs using sparse autoencoders
Dylan Zhou, Kunal Patil, Yifan Sun, Karthik lakshmanan, Senthooran Rajamanoharan, and Arthur Conmy. 2025a · 2025
Closest in time.
Better explain transformers by illuminating important information
Linxin Song, Yan Cui, Ao Luo, Freddy Lecue, and Irene Li. 2024 · 2062
Closest in time.
Entailer: Answering questions with faithful and truthful chains of reasoning
Oyvind Tafjord, Bhavana Dalvi Mishra, and Peter Clark. 2022 · 2093
Closest in time.