Fetching the paper…
Reading the bibliography…
Interpretable machine learning has exploded as an area of interest over the last decade, sparked by the rise of increasingly large datasets and deep neural networks.
Classification and Regression Trees
L. Breiman, J. H. Friedman, R. A. Olshen, and C. J. Stone · 1984
Earlier work this paper cites.
Generalized additive models
Trevor Hastie and Robert Tibshirani · 1986
Earlier work this paper cites.
Induction of decision trees
J. Ross Quinlan · 1986
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Accurate intelligible models with pairwise interactions
Yin Lou, Rich Caruana, Johannes Gehrke, and Giles Hooker · 2013
Earlier work this paper cites.
Understanding neural networks through deep visualization
Jason Yosinski, Jeff Clune, Anh Nguyen, Thomas Fuchs, and Hod Lipson · 2015
Earlier work this paper cites.
Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission
Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad · 2015
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
European union regulations on algorithmic decision-making and a” right to explanation”
Bryce Goodman and Seth Flaxman · 2016
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Generating visual explanations
Lisa Anne Hendricks, Zeynep Akata, Marcus Rohrbach, Jeff Donahue, Bernt Schiele, and Trevor Darrell · 2016
Earlier work this paper cites.
Supersparse linear integer models for optimized medical scoring systems
Berk Ustun and Cynthia Rudin · 2016
Earlier work this paper cites.
A roadmap for a rigorous science of interpretability
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory Sayres · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Explaining nonlinear classification decisions with deep taylor decomposition
Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert Müller · 2017
Earlier work this paper cites.
GAN dissection: Visualizing and understanding generative adversarial networks
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Bolei Zhou, Joshua B Tenenbaum, William T Freeman, and Antonio Torralba · 2018
Earlier work this paper cites.
Distill-and-Compare: Auditing black-box models using transparent model distillation
Sarah Tan, Rich Caruana, Giles Hooker, and Yin Lou · 2018
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom · 2018
Earlier work this paper cites.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni · 2018
Earlier work this paper cites.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen · 2018
Earlier work this paper cites.
Definitions, methods, and applications in interpretable machine learning
W. James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu · 2019
Earlier work this paper cites.
Interpretable machine learning
Christoph Molnar · 2019
Earlier work this paper cites.
The challenge of crafting intelligible intelligence
Daniel S Weld and Gagan Bansal · 2019
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 2019
Earlier work this paper cites.
Sarthak Jain and Byron C Wallace · 2019
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher · 2019
Earlier work this paper cites.
Incorporating priors with feature attribution on text classification
Frederick Liu and Besim Avci · 2019
Earlier work this paper cites.
What does BERT look at? An analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel · 2020
Earlier work this paper cites.
Interpreting interpretability: understanding data scientists’ use of interpretability tools for machine learning
Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna Wallach, and Jennifer Wortman Vaughan · 2020
Earlier work this paper cites.
Interpretation of nlp models through input marginalization
Siwon Kim, Jihun Yi, Eunji Kim, and Sungroh Yoon · 2020
Earlier work this paper cites.
REALM: Retrieval-augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang · 2020
Earlier work this paper cites.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas · 2020
Earlier work this paper cites.
Interpretable machine learning: Fundamental principles and 10 grand challenges
Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong · 2021
Earlier work this paper cites.
Adaptive wavelet distillation from neural networks through interpretations
Wooseok Ha, Chandan Singh, Francois Lanusse, Srigokul Upadhyayula, and Bin Yu · 2021
Earlier work this paper cites.
Learning from learning machines: a new generation of AI technology to meet the needs of science
Luca Pion-Tonachini, Kristofer Bouchard, Hector Garcia Martin, Sean Peisert, W Bradley Holtz, Anil Aswani, Dipankar Dwivedi, Haruko Wainwright, Ghanshyam Pilania, Benjamin Nachman, et al · 2021
Earlier work this paper cites.
Does the whole exceed its parts? the effect of AI explanations on complementary team performance
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld · 2021
Earlier work this paper cites.
Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm, and Katja Filippova · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Integrated directional gradients: Feature interaction attribution for neural nlp models
Sandipan Sikdar, Parantapa Bhattacharya, and Kieran Heese · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2021
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Earlier work this paper cites.
imodels: a python package for fitting interpretable models
Chandan Singh, Keyan Nasseri, Yan Shuo Tan, Tiffany Tang, and Bin Yu · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Rev: information-theoretic evaluation of free-text rationales
Hanjie Chen, Faeze Brahman, Xiang Ren, Yangfeng Ji, Yejin Choi, and Swabha Swayamdipta · 2022
Earlier work this paper cites.
Can language models learn from explanations in context?
Andrew K Lampinen, Ishita Dasgupta, Stephanie CY Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L McClelland, Jane X Wang, and Felix Hill · 2022
Earlier work this paper cites.
Locating and editing factual knowledge in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Earlier work this paper cites.
Fast model editing at scale, 2022
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning · 2022
Earlier work this paper cites.
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic · 2022
Earlier work this paper cites.
Multi-modal molecule structure-text model for text-based retrieval and editing
Shengchao Liu, Weili Nie, Chengpeng Wang, Jiarui Lu, Zhuoran Qiao, Ling Liu, Jian Tang, Chaowei Xiao, and Anima Anandkumar · 2022
Cited alongside, same era.
Rethinking explainability as a dialogue: A practitioner’s perspective
Himabindu Lakkaraju, Dylan Slack, Yuxin Chen, Chenhao Tan, and Sameer Singh · 2022
Cited alongside, same era.
The unreliability of explanations in few-shot prompting for textual reasoning
Xi Ye and Greg Durrett · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou · 2022
Cited alongside, same era.
Text and patterns: For effective chain of thought, it takes two to tango
Larger language models do in-context learning differently
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al · 2023
Later among the works it cites.
Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Lidén, Zhou Yu, Weizhu Chen, and Jianfeng Gao · 2023
Later among the works it cites.
Unifying corroborative and contributive attributions in large language models
Theodora Worledge, Judy Hanwen Shen, Nicole Meister, Caleb Winston, and Carlos Guestrin · 2023
Later among the works it cites.
Text embeddings reveal (almost) as much as text
John X Morris, Volodymyr Kuleshov, Vitaly Shmatikov, and Alexander M Rush · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aman Madaan and Amir Yazdanbakhsh · 2022
Cited alongside, same era.
Towards understanding chain-of-thought prompting: An empirical study of what matters
Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, and Huan Sun · 2022
Cited alongside, same era.
Measuring and narrowing the compositionality gap in language models, 2022
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A. Smith, and Mike Lewis · 2022
Cited alongside, same era.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, et al · 2022
Cited alongside, same era.
Natural language descriptions of deep visual features
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Cited alongside, same era.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Cited alongside, same era.
What can transformers learn in-context? A case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant · 2022
Cited alongside, same era.
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, Zico Kolter, and Dan Hendrycks · 2023
Later among the works it cites.
Eliciting latent predictions from transformers with the tuned lens
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Later among the works it cites.
Finding neurons in a haystack: Case studies with sparse probing
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas · 2023
Later among the works it cites.
How do language models bind entities in context?
Jiahai Feng and Jacob Steinhardt · 2023
Later among the works it cites.
Circuit component reuse across tasks in transformer language models
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick · 2023
Later among the works it cites.
Tom Lieberum, Matthew Rahtz, János Kramár, Geoffrey Irving, Rohin Shah, and Vladimir Mikulik · 2023
Later among the works it cites.
Interpretability at scale: Identifying causal mechanisms in Alpaca
Zhengxuan Wu, Atticus Geiger, Christopher Potts, and Noah D. Goodman · 2023
Later among the works it cites.
What algorithms can transformers learn? a study in length generalization
Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Josh Susskind, Samy Bengio, and Preetum Nakkiran · 2023
Later among the works it cites.
Studying large language model generalization with influence functions
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, et al · 2023
Later among the works it cites.
Large language models struggle to learn long-tail knowledge
Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel · 2023
Later among the works it cites.
Sources of hallucination by large language models on inference tasks
Nick McKenna, Tianyi Li, Liang Cheng, Mohammad Javad Hosseini, Mark Johnson, and Mark Steedman · 2023
Later among the works it cites.
Tell your model where to attend: Post-hoc attention steering for LLMs
Qingru Zhang, Chandan Singh, Liyuan Liu, Xiaodong Liu, Bin Yu, Jianfeng Gao, and Tuo Zhao · 2023
Later among the works it cites.
The truth is in there: Improving reasoning in language models with layer-selective rank reduction
Pratyusha Sharma, Jordan T Ash, and Dipendra Misra · 2023
Later among the works it cites.
Victor Dibia · 2023
Later among the works it cites.
Benchmarking large language models as AI research agents
Qian Huang, Jian Vora, Percy Liang, and Jure Leskovec · 2023
Later among the works it cites.
Table-GPT: Table-tuned gpt for diverse table tasks
Peng Li, Yeye He, Dror Yashar, Weiwei Cui, Song Ge, Haidong Zhang, Danielle Rifinski Fainman, Dongmei Zhang, and Surajit Chaudhuri · 2023
Later among the works it cites.
Towards foundation models for learning on tabular data
Han Zhang, Xumeng Wen, Shun Zheng, Wei Xu, and Jiang Bian · 2023
Later among the works it cites.
Generative table pre-training empowers models for tabular prediction
Tianping Zhang, Shaowen Wang, Shuicheng Yan, Jian Li, and Qian Liu · 2023
Later among the works it cites.
LLMs understand glass-box models, discover surprises, and suggest repairs
Benjamin J Lengerich, Sebastian Bordt, Harsha Nori, Mark E Nunnally, Yin Aphinyanaphongs, Manolis Kellis, and Rich Caruana · 2023
Later among the works it cites.
MaNtLE: Model-agnostic natural language explainer
Rakesh R Menon, Kerem Zaman, and Shashank Srivastava · 2023
Later among the works it cites.
Augmenting interpretable models with large language models during training
Chandan Singh, Armin Askari, Rich Caruana, and Jianfeng Gao · 2023
Later among the works it cites.
Denis Jered McInerney, Geoffrey Young, Jan-Willem van de Meent, and Byron C Wallace · 2023
Later among the works it cites.
Designing LLM chains by adapting techniques from crowdsourcing workflows
Madeleine Grunde-McLaughlin, Michelle S Lam, Ranjay Krishna, Daniel S Weld, and Jeffrey Heer · 2023
Later among the works it cites.
Tree prompting: efficient task adaptation without fine-tuning
John X Morris, Chandan Singh, Alexander M Rush, Jianfeng Gao, and Yuntian Deng · 2023
Later among the works it cites.
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang · 2023
Later among the works it cites.
Self-Refine: Iterative refinement with self-feedback, 2023
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark · 2023
Later among the works it cites.
Self-verification improves few-shot clinical information extraction
Zelalem Gero, Chandan Singh, Hao Cheng, Tristan Naumann, Michel Galley, Jianfeng Gao, and Hoifung Poon · 2023
Later among the works it cites.
Augmented language models: a survey
Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, et al · 2023
Later among the works it cites.
Explaining patterns in data with language models via interpretable autoprompting, 2023
Chandan Singh, John X. Morris, Jyoti Aneja, Alexander M. Rush, and Jianfeng Gao · 2023
Later among the works it cites.
Goal-driven explainable clustering via language descriptions
Zihan Wang, Jingbo Shang, and Ruiqi Zhong · 2023
Later among the works it cites.
TopicGPT: A prompt-based topic modeling framework
Chau Minh Pham, Alexander Hoyle, Simeng Sun, and Mohit Iyyer · 2023
Later among the works it cites.
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr · 2023
Later among the works it cites.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman · 2023
Later among the works it cites.
Lost in the middle: How language models use long contexts
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang · 2023
Later among the works it cites.
Benchmarking and improving generator-validator consistency of language models
Xiang Lisa Li, Vaishnavi Shrivastava, Siyan Li, Tatsunori Hashimoto, and Percy Liang · 2023
Later among the works it cites.
Large language models for automated open-domain scientific hypotheses discovery
Zonglin Yang, Xinya Du, Junxian Li, Jie Zheng, Soujanya Poria, and Erik Cambria · 2023
Later among the works it cites.
Multi-modal molecule structure–text model for text-based retrieval and editing
Shengchao Liu, Weili Nie, Chengpeng Wang, Jiarui Lu, Zhuoran Qiao, Ling Liu, Jian Tang, Chaowei Xiao, and Animashree Anandkumar · 2023
Later among the works it cites.
Mathematical discoveries from program search with large language models
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M. Pawan Kumar, Emilien Dupont, Francisco J. R. Ruiz, Jordan Ellenberg, Pengming Wang, Omar Fawzi, Pushmeet Kohli, and Alhussein Fawzi · 2023
Later among the works it cites.
Bridging the Human-AI knowledge gap: Concept discovery and transfer in alphazero
Lisa Schut, Nenad Tomasev, Tom McGrath, Demis Hassabis, Ulrich Paquet, and Been Kim · 2023
Later among the works it cites.
Eliciting human preferences with language models
Belinda Z Li, Alex Tamkin, Noah Goodman, and Jacob Andreas · 2023
Later among the works it cites.
Recommender AI agent: Integrating large language models for interactive recommendations
Xu Huang, Jianxun Lian, Yuxuan Lei, Jing Yao, Defu Lian, and Xing Xie · 2023
Later among the works it cites.
Relying on the unreliable: The impact of language models’ reluctance to express uncertainty, 2024
Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, and Maarten Sap · 2024
Closest in time.
PatchScope: A unifying framework for inspecting hidden representations of language models, 2024
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva · 2024
Closest in time.
In-context language learning: Arhitectures and algorithms, 2024
Ekin Akyürek, Bailin Wang, Yoon Kim, and Jacob Andreas · 2024
Closest in time.
LLMCheckup: Conversational examination of large language models via interpretability tools
Qianli Wang, Tatiana Anikina, Nils Feldhus, Josef van Genabith, Leonhard Hennig, and Sebastian Möller · 2024
Closest in time.
A comprehensive survey of hallucination mitigation techniques in large language models
SM Tonmoy, SM Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das · 2024
Closest in time.
Towards consistent natural-language explanations via explanation-consistency finetuning, 2024
Yanda Chen, Chandan Singh, Xiaodong Liu, Simiao Zuo, Bin Yu, He He, and Jianfeng Gao · 2024
Closest in time.
Deductive closure training of language models for coherence, accuracy, and updatability
Afra Feyza Akyürek, Ekin Akyürek, Leshem Choshen, Derry Wijaya, and Jacob Andreas · 2024
Closest in time.
Towards conversational diagnostic ai
Tao Tu, Anil Palepu, Mike Schaekermann, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin, Nenad Tomasev, et al · 2024
Closest in time.