Fetching the paper…
Reading the bibliography…
Artificial Intelligence (AI) has continued to achieve tremendous success in recent times.
Hierarchical structure in perceptual representation
Palmer S · 1977
Earlier work this paper cites.
Objects, parts, and categories
Tversky B and Hemenway K · 1984
Earlier work this paper cites.
The copycat project: An experiment in nondeterminism and creative analogies
Hofstadter D · 1984
Earlier work this paper cites.
Induction of decision trees
Quinlan J · 1986
Earlier work this paper cites.
Generalized additive models: Some applications
Hastie T and Tibshirani R · 1987
Earlier work this paper cites.
The apache iii prognostic system: Risk prediction of hospital mortality for critically iii hospitalized adults
Knaus W · 1991
Earlier work this paper cites.
Improved use of continuous attributes in c4.5
Quinlan J · 1996
Earlier work this paper cites.
Generating accurate rule sets without global optimization
Frank E and Witten I · 1998
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Uses and misuses of the correlation coefficient
Onwuegbuzie A and Daniel L · 1999
Earlier work this paper cites.
When to control for covariates? panel asymptotics for estimates of treatment effects
Angrist J and Hahn J · 2004
Earlier work this paper cites.
Stable and efficient multiple smoothing parameter estimation for generalized additive models
Wood S · 2004
Earlier work this paper cites.
Commonality analysis: Partitioning variance to facilitate better understanding of data
Zientek L and Thompson B · 2006
Earlier work this paper cites.
Commonality analysis: Partitioning variance to facilitate better understanding of data
Linda Reichwein Zientek and Bruce Thompson · 2006
Earlier work this paper cites.
Component selection and smoothing in multivariate nonparametric regression
Lin Y and Zhang H · 2006
Earlier work this paper cites.
True to the model or true to the data?
Chen H, Janizek J, Lundberg S, and Lee S.-I · 2006
Earlier work this paper cites.
High-dimensional additive modeling
Meier L, Van De S, Geer P, and Bühlmann · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky A and Hinton G · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng J, Dong W, Socher R, Li L, Li K, and Fei-Fei L · 2009
Earlier work this paper cites.
Simple means to improve the interpretability of regression coefficients
Schielzeth H · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair V and Hinton G · 2010
Earlier work this paper cites.
Regression commonality analysis: A technique for quantitative theory building
Nimon K and Reio T · 2011
Earlier work this paper cites.
Classification and regression trees
Loh W · 2011
Earlier work this paper cites.
Evaluation and improvement of interpretability for self-explainable part-prototype networks
Huang Q · 2011
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Wah C, Branson S, Welinder P, Perona P, and Belongie S · 2011
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie · 2011
Earlier work this paper cites.
Accurate intelligible models with pairwise interactions
Lou Y, Caruana R, Gehrke J, and Hooker G · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan K, Vedaldi A, and Zisserman A · 2013
Earlier work this paper cites.
Visualizing and Understanding Convolutional Networks
Zeiler M and Fergus R · 2013
Earlier work this paper cites.
The community earth system model: a framework for collaborative research
Hurrell J · 2013
Earlier work this paper cites.
Conceptnet 5: A large semantic network for relational knowledge
Speer R and Havasi C · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma D · 2013
Earlier work this paper cites.
Using commonality analysis in multiple regressions: a tool to decompose regression effects in the face of multicollinearity
Ray-Mukherjee J, Nimon K, Mukherjee S, Morris D, Slotow R, and Hamer M · 2014
Earlier work this paper cites.
Inference on treatment effects after selection among high-dimensional controls
Belloni A, Chernozhukov V, and Hansen C · 2014
Earlier work this paper cites.
Fifty years of classification and regression trees
Loh W · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Springenberg J, Dosovitskiy A, Brox T, and Riedmiller M · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau D, Cho K, and Bengio Y · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach S, Binder A, Montavon G, Klauschen F, Müller K, and Samek W · 2015
Earlier work this paper cites.
The Mythos of Model Interpretability
Lipton Z · 2016
Earlier work this paper cites.
Sparse partially linear additive models
Lou Y, Bien J, Caruana R, and Gehrke J · 2016
Earlier work this paper cites.
Supersparse linear integer models for optimized medical scoring systems
Ustun B and Rudin C · 2016
Earlier work this paper cites.
Explaining the predictions of any classifier
Ribeiro M, Singh S, and Guestrin C · 2016
Earlier work this paper cites.
Not Just a Black Box: Learning Important Features Through Propagating Activation Differences
Shrikumar A, Greenside P, Shcherbina A, and Kundaje A · 2016
Earlier work this paper cites.
Learning deep features for discriminative localization
Zhou B, Khosla A, Lapedriza A, Oliva A, and Torralba A · 2016
Earlier work this paper cites.
Retain: An interpretable predictive model for healthcare using reverse time attention mechanism
Edward Choi, Mohammad Taha Bahadori, Jimeng Sun, Joshua Kulas, Andy Schuetz, and Walter Stewart · 2016
Earlier work this paper cites.
Not just a black box: Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, Anna Shcherbina, and Anshul Kundaje · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks D and Gimpel K · 2016
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Johnson A · 2016
Earlier work this paper cites.
Accountability of ai under the law: The role of explanation
Doshi-Velez F · 2017
Earlier work this paper cites.
Linear regression with a randomly censored covariate: Application to an alzheimer’s study
Atem F, Qian J, Maye J, Johnson K, and Betensky R · 2017
Earlier work this paper cites.
Generalized additive models
Hastie T · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg S and Lee S.-I · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan M, Taly A, and Teh Y · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju R, Cogswell M, Das A, Vedantam R, Parikh D, and Batra D · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Smilkov D, Thorat N, Kim B, Viégas F, and Wattenberg M · 2017
Earlier work this paper cites.
Explaining nonlinear classification decisions with deep taylor decomposition
Montavon G, Lapuschkin S, Binder A, Samek W, and Müller K · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Shrikumar A, Greenside P, and Kundaje A · 2017
Earlier work this paper cites.
What is relevant in a text document?’: An interpretable machine learning approach
Arras L, Horn F, Montavon G, Müller K, and Samek W · 2017
Earlier work this paper cites.
Explaining recurrent neural network predictions in sentiment analysis
Arras L, Montavon G, Müller K, and Samek W · 2017
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Samek W, Binder A, Montavon G, Lapuschkin S, and Müller K · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani A · 2017
Earlier work this paper cites.
Top-down neural attention by excitation backprop
Zhang J, Bargal S, Lin Z, Brandt J, Shen X, and Sclaroff S · 2017
Earlier work this paper cites.
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
Wachter S, Mittelstadt B, and Russell C · 2017
Earlier work this paper cites.
Places: A 10 million image database for scene recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Man against machine: diagnostic performance of a deep learning convolutional neural network for dermoscopic melanoma recognition in comparison to 58 dermatologists
Haenssle H · 2018
Earlier work this paper cites.
The mythos of model interpretability
Lipton Z · 2018
Earlier work this paper cites.
Explaining explanations: An overview of interpretability of machine learning
Gilpin L, Bau D, Yuan B, Bajwa A, Specter M, and Kagal L · 2018
Earlier work this paper cites.
Methods for interpreting and understanding deep neural networks
Montavon G, Samek W, and Müller K · 2018
Earlier work this paper cites.
Increasing transparency in algorithmic-decision-making with explainable ai
Waltl B and Vogl R · 2018
Earlier work this paper cites.
Balancing the trade-off between accuracy and interpretability in software defect prediction
Mori T and Uchihira N · 2018
Earlier work this paper cites.
Peeking inside the black-box: A survey on explainable artificial intelligence (xai)
Adadi A and Berrada M · 2018
Earlier work this paper cites.
Inference in linear regression models with many covariates and heteroscedasticity
Cattaneo M, Jansson M, and Newey W · 2018
Earlier work this paper cites.
Local explanation methods for deep neural networks lack sensitivity to parameter values
Adebayo J, Gilmer J, Goodfellow I, and Kim B · 2018
Earlier work this paper cites.
Audiomnist: Exploring explainable artificial intelligence for audio analysis on a simple benchmark
Becker S, Vielhaben J, Ackermann M, Müller K, Lapuschkin S, and Samek W · 2018
Earlier work this paper cites.
Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks
Chattopadhay A, Sarkar A, Howlader P, and Balasubramanian V · 2018
Earlier work this paper cites.
Rise: Randomized input sampling for explanation of black-box models
Petsiuk V, Das A, and Saenko K · 2018
Earlier work this paper cites.
A theoretical explanation for perplexing behaviors of backpropagationbased visualizations
Nie W, Zhang Y, and Krause A · 2018
Earlier work this paper cites.
Sanity checks for saliency maps
Adebayo J, Gilmer J, Muelly M, Goodfellow I, Hardt M, and Garnett R · 2018
Earlier work this paper cites.
Disan: Directional self-attention network for rnn/cnn-free language understanding
Shen T, Jiang J, Zhou T, Pan S, Long G, and Zhang C · 2018
Earlier work this paper cites.
Learn to pay attention
Jetley S, Lord N, Lee N, and Torr P · 2018
Earlier work this paper cites.
Attention u-net: Learning where to look for the pancreas
Oktay O · 2018
Earlier work this paper cites.
Interpretable credit application predictions with counterfactual explanations
Grath R · 2018
Earlier work this paper cites.
Deep learning for case-based reasoning through prototypes: A neural network that explains its predictions
Li O, Liu H, Chen C, and Rudin C · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim B · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin J, Chang M, Lee K, and Toutanova K · 2018
Earlier work this paper cites.
Did the model understand the question?
Mudrakarta P, Taly A, Sundararajan M, and Dhamdhere K · 2018
Earlier work this paper cites.
Deep contextualized word representations
Peters M · 2018
Earlier work this paper cites.
Dissecting contextual word embeddings: Architecture and representation
Peters M, Neumann M, Zettlemoyer L, and Yih W · 2018
Earlier work this paper cites.
Places: A 10 million image database for scene recognition
Zhou B, Lapedriza A, Khosla A, Oliva A, and Torralba A · 2018
Earlier work this paper cites.
Development and validation of deep learning-based automatic detection algorithm for malignant pulmonary nodules on chest radiographs
Nam J · 2019
Earlier work this paper cites.
Deep learning outperformed 136 of 157 dermatologists in a head-to-head dermoscopic melanoma image classification task
Brinker T · 2019
Earlier work this paper cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Rudin C · 2019
Earlier work this paper cites.
Unmasking clever hans predictors and assessing what machines really learn
Lapuschkin S, Wäldchen S, Binder A, Montavon G, Samek W, and Müller K · 2019
Earlier work this paper cites.
Balancing accuracy and interpretability of machine learning approaches for radiation treatment outcomes modeling
Luo Y, Tseng H.-H, Cui S, Wei L, Ten Haken R, and Naqa I · 2019
Earlier work this paper cites.
Axiomatic interpretability for multiclass additive models
Zhang X, Lou Y, Tan S, Chajewska U, Koch P, and Caruana R · 2019
Earlier work this paper cites.
Decision tree underfitting in mining of gene expression data. an evolutionary multi-test tree approach
Czajkowski M and Kretowski M · 2019
Earlier work this paper cites.
Optimal sparse decision trees
Hu X, Rudin C, and Seltzer M · 2019
Earlier work this paper cites.
The (un)reliability of saliency methods
Kindermans P · 2019
Earlier work this paper cites.
Understanding individual decisions of cnns via contrastive backpropagation
Gu J, Yang Y, and Tresp V · 2019
Earlier work this paper cites.
Explaining convolutional neural networks using softmax gradient layer-wise relevance propagation
Iwana B, Kuroki R, and Uchida S · 2019
Earlier work this paper cites.
A benchmark for interpretability methods in deep neural networks
Hooker S, Erhan D, Kindermans P.-J, and Brain K · 2019
Earlier work this paper cites.
Developing the sensitivity of lime for better machine learning explanation
Lee E, Braines D, Stiffler M, Hudler A, and Spie Ed · 2019
Earlier work this paper cites.
Dlime: A deterministic local interpretable model-agnostic explanations approach for computer-aided diagnosis systems
Zafar M and Khan N · 2019
Earlier work this paper cites.
Alime: Autoencoder based approach for local interpretability
Shankaranarayana S and Runje D · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Voita E, Talbot D, Moiseev F, Sennrich R, and Titov I · 2019
Earlier work this paper cites.
Attention is not explanation
Jain S and Wallace B · 2019
Earlier work this paper cites.
Is attention interpretable?
Serrano S and Smith N · 2019
Earlier work this paper cites.
Eraser: A benchmark to evaluate rationalized nlp models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C Wallace · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh V, Debut L, Chaumond J, and Wolf T · 2019
Earlier work this paper cites.
This looks like that: Deep learning for interpretable image recognition
Chen C, Li O, Tao D, Barnett A, Rudin C, and Su J · 2019
Earlier work this paper cites.
The multifactorial nature of beak and skull shape evolution in parrots and cockatoos (psittaciformes)
Jen A Bright, Jesús Marugán-Lobón, Emily J Rayfield, and Samuel N Cobb · 2019
Earlier work this paper cites.
Towards automatic conceptbased explanations
Ghorbani A, Wexler Google Brain J, Zou J, and Kim Google Brain B · 2019
Cited alongside, same era.
The (un) reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T Schütt, Sven Dähne, Dumitru Erhan, and Been Kim · 2019
Cited alongside, same era.
Perturbation sensitivity analysis to detect unintended model biases
Prabhakaran V, Hutchinson B, and Mitchell M · 2019
Cited alongside, same era.
Towards hierarchical importance attribution: Explaining compositional semantics for neural sequence models
Jin X, Wei Z, Du J, Xue X, and Ren X · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford A · 2019
Cited alongside, same era.
Reducing sentiment bias in language models via counterfactual evaluation
Huang P · 2019
Effects of explainable artificial intelligence on trust and human behavior in a high-risk decision task
Leichtmann B, Humer C, Hinterreiter A, Streit M, and Mara M · 2023
Later among the works it cites.
Revisiting the performance-explainability trade-off in explainable artificial intelligence (xai)
Crook B, Schluter M, and Speith T · 2023
Later among the works it cites.
Explainable ai (xai): Core ideas, techniques, and solutions
Dwivedi R · 2023
Later among the works it cites.
A comprehensive taxonomy for explainable artificial intelligence: a systematic survey of surveys on methods and concepts
Schwalbe G and Finzel B · 2023
Later among the works it cites.
Survey of explainable ai techniques in healthcare
Chaddad A, Peng J, Xu J, and Bouridane A · 2023
Later among the works it cites.
Concise and interpretable multi-label rule sets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention interpretability across nlp tasks
Vashishth S, Upadhyay S, Assistant G, Singh G, Research T, and Faruqui M · 2019
Cited alongside, same era.
What does bert look at? an analysis of bert’s attention
Clark K, Khandelwal U, Levy O, and Manning C · 2019
Cited alongside, same era.
Revealing the dark secrets of bert
Kovaleva O, Romanov A, Rogers A, and Rumshisky A · 2019
Cited alongside, same era.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis M · 2019
Cited alongside, same era.
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
Hewitt J and Manning C · 2019
Cited alongside, same era.
Ciaperoni M, Xiao H, and Gionis A · 2023
Later among the works it cites.
Fimf score-cam: Fast score-cam based on local multi-feature integration for visual interpretation of cnns
Li J, Zhang D, Meng B, Li Y, and Luo L · 2023
Later among the works it cites.
Abs-cam: a gradient optimization interpretable approach for explanation of convolutional neural networks
Chunyan Zeng, Kang Yan, Zhifeng Wang, Yan Yu, Shiyan Xia, and Nan Zhao · 2023
Later among the works it cites.
Algorithms to estimate shapley value feature attributions
Chen H, Covert I, Lundberg S, and Lee S · 2023
Later among the works it cites.
Harsanyinet: Computing accurate shapley values in a single forward propagation
Chen L, Lou S, Zhang K, Huang J, and Zhang Q · 2023
Later among the works it cites.
Faith-shap: The faithful shapley interaction index
Tsai C.-P, Yeh C.-K, and Ravikumar P · 2023
Later among the works it cites.
Shap-iq: Unified approximation of any-order shapley interactions
Fumagalli F, Muschalik M, Kolpaczki P, Hüllermeier E, and Hammer B · 2023
Later among the works it cites.
Application of explainable artificial intelligence in medical health: A systematic review of interpretability methods
Band S · 2023
Later among the works it cites.
Glime: General, stable and local lime explanation
Tan Z, Tian Y, and Li J · 2023
Later among the works it cites.
Faithful explanations of blackbox nlp models using llm-generated counterfactuals
Gat Y, Calderon N, Feder A, Chapanin A, Sharma A, and Reichart R · 2023
Later among the works it cites.
Performance of chatgpt and gpt-4 on neurosurgery written board examinations
Ali R · 2023
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Deep prototypical-parts ease morphological kidney stone identification and are competitively robust to photometric perturbations
Flores-Araiza D · 2023
Later among the works it cites.
Sanity checks and improvements for patch visualisation in prototype-based image classification
Xu-Darme R, Quénot G, Chihani Z, and Rousset M.-C · 2023
Later among the works it cites.
Towards trustable skin cancer diagnosis via rewriting model’s decision
Yan S · 2023
Later among the works it cites.
A closer look at the intervention procedure of concept bottleneck models
Shin S, Postech, Jo Y, Ahn S, Lee N, and Postech · 2023
Later among the works it cites.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Zhang Y · 2023
Later among the works it cites.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Huang L · 2023
Later among the works it cites.
Sequential integrated gradients: a simple but effective method for explaining language models
Enguehard J · 2023
Later among the works it cites.
Mistral 7b
Jiang A · 2023
Later among the works it cites.
Emergent analogical reasoning in large language models
Webb T, Holyoak K, and Lu H · 2023
Later among the works it cites.
Response: Emergent analogical reasoning in large language models
Hodel D and West J · 2023
Later among the works it cites.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Turpin M, Michael J, Perez E, and Bowman S · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao S · 2023
Later among the works it cites.
Decompx: Explaining transformers decisions by propagating token decomposition
Modarressi A, Fayyaz M, Aghazadeh E, Yaghoobzadeh Y, and Pilehvar M · 2023
Later among the works it cites.
Analyzing feed-forward blocks in transformers through the lens of attention maps
Kobayashi G, Kuribayashi T, Yokoi S, and Inui K · 2023
Later among the works it cites.
Probing for hyperbole in pre-trained language models
Schneidermann N, Hershcovich D, and Pedersen B · 2023
Later among the works it cites.
Finding neurons in a haystack: Case studies with sparse probing
Gurnee W, Nanda N, Pauly M, Harvey K, Troitskii D, and Bertsimas D · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models
Nanda N, Lee A, and Wattenberg M · 2023
Later among the works it cites.
Towards automated circuit discovery for mechanistic interpretability
Conmy A, Mavor-Parker A, Lynch A, Heimersheim S, Garriga-Alonso A, and Ai F · 2023
Later among the works it cites.
Attribution patching outperforms automated circuit discovery
Syed A, Rager C, and Conmy A · 2023
Later among the works it cites.
Attribution patching: Activation patching at industrial scale
Nanda N · 2023
Later among the works it cites.
Sparse autoencoders find highly interpretable features in language models
Cunningham H, Ewart A, Riggs L, Huben R, and Sharkey L · 2023
Later among the works it cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken T · 2023
Later among the works it cites.
Eliciting latent predictions from transformers with the tuned lens
Belrose N · 2023
Later among the works it cites.
Lico: Explainable models with language-image consistency
Lei Y, Li Z, Li Y, Zhang J, and Shan H · 2023
Later among the works it cites.
Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation
Lin Y · 2023
Later among the works it cites.
Huntgpt: Integrating machine learning-based anomaly detection and explainable ai with large language models (llms)
Ali T and Kostakos P · 2023
Later among the works it cites.
Explaining machine learning models with interactive natural language conversations using talktomodel
Slack D, Krishna S, Lakkaraju H, and Singh S · 2023
Later among the works it cites.
Interrolang: Exploring nlp models and datasets through dialogue-based explanations
Feldhus N, Wang Q, Anikina T, Chopra S, Oguz C, and Möller S · 2023
Later among the works it cites.
Tell me a story! narrative-driven xai with large language models
Martens D, Hinns J, Dams C, Vergouwen M, and Evgeniou T · 2023
Later among the works it cites.
Sparse linear concept discovery models
Panousis K, Ienco D, and Marcos D · 2023
Later among the works it cites.
Language in a bottle: Language model guided concept bottlenecks for interpretable image classification
Yang Y, Panagopoulou A, Zhou S, Jin D, Callison-Burch C, and Yatskar M · 2023
Later among the works it cites.
Label-free concept bottleneck models
Oikarinen T, Das S, Nguyen L, and Weng T · 2023
Later among the works it cites.
Clip-qda: An explainable concept bottleneck model
Kazmierczak R, Berthier E, Frehse G, and Franchi G · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron H · 2023
Later among the works it cites.
Learning concise and descriptive attributes for visual recognition
Yan A · 2023
Later among the works it cites.
Learning to intervene on concept bottlenecks
Steinmann D, Stammer W, Friedrich F, and Kersting K · 2023
Later among the works it cites.
Faithful vision-language interpretation via concept bottleneck models
Lai S, Hu L, Wang J, Berti-Equille L, and Wang D · 2023
Later among the works it cites.
A critical survey on fairness benefits of explainable ai
Deck L, Schoeffer J, De-Arteaga M, and Kühl N · 2024
Later among the works it cites.
Explainable ai for safe and trustworthy autonomous driving: A systematic review
Kuznietsov A, Gyevnar B, Wang C, Peters S, and Albrecht S · 2024
Later among the works it cites.
Safety implications of explainable artificial intelligence in end-to-end autonomous driving
Atakishiyev S, Salameh M, and Goebel R · 2024
Later among the works it cites.
Rethinking interpretability in the era of large language models
Singh C, Inala J, Galley M, Caruana R, and Gao J · 2024
Later among the works it cites.
Explainability for large language models: A survey
Zhao H · 2024
Later among the works it cites.
Improving deep learning with prior knowledge and cognitive models: A survey on enhancing explainability, adversarial robustness and zero-shot learning
Mumuni F and Mumuni A · 2024
Later among the works it cites.
Are linear regression models white box and interpretable?
Salih A and Wang Y · 2024
Later among the works it cites.
Automated data processing and feature engineering for deep learning and big data applications: A survey
Mumuni A and Mumuni F · 2024
Later among the works it cites.
Gaussian process neural additive models
Zhang W, Barr B, and Paisley J · 2024
Later among the works it cites.
Grad++scorecam: Enhancing visual explanations of deep convolutional networks using incremented gradient and score-weighted methods
Soomro S, Niaz A, and Choi K · 2024
Later among the works it cites.
Finding the right xai method-a guide for the evaluation and ranking of explainable ai methods in climate science
Bommer P, Kretschmer M, Hedström A, Bareeva D, M, and Höhne C · 2024
Later among the works it cites.
Poly-cam: high resolution class activation map for convolutional neural networks
Englebert A, Cornu O, and Vleeschouwer C · 2024
Later among the works it cites.
Opti-cam: Optimizing saliency maps for interpretability
Zhang H, Torres F, Sicre R, Avrithis Y, and Ayache S · 2024
Later among the works it cites.
Cape: Cam as a probabilistic ensemble for enhanced dnn interpretation
Chowdhury T · 2024
Later among the works it cites.
Gt-cam: Game theory based class activation map for gcn
Li Y, Shi T, Chen Z, Zhang L, and Xie W · 2024
Later among the works it cites.
Shap@k: Efficient and probably approximately correct (pac) identification of top-k features
Kariyappa S, Tsepenekas L, Lécué F, and Magazzeni D · 2024
Later among the works it cites.
On the failings of shapley values for explainability
Huang X and Marques-Silva J · 2024
Later among the works it cites.
Interpreting artificial intelligence models: a systematic review on the application of lime and shap in alzheimer’s disease detection
Vimbi V, Shaffi N, and Mahmud M · 2024
Later among the works it cites.
Explainable ai for intrusion detection systems: Lime and shap applicability on multi-layer perceptron
Gaspar D, Silva P, and Silva C · 2024
Later among the works it cites.
Us-lime: Increasing fidelity in lime using uncertainty sampling on tabular data
Saadatfar H, Kiani-Zadegan Z, and Ghahremani-Nezhad B · 2024
Later among the works it cites.
Shap-iq: Unified approximation of any-order shapley interactions
Fabian Fumagalli, Maximilian Muschalik, Patrick Kolpaczki, Eyke Hüllermeier, and Barbara Hammer · 2024
Later among the works it cites.
Visual attention methods in deep learning: An in-depth survey
Hassanin M, Anwar S, Radwan I, Khan F, and Mian A · 2024
Later among the works it cites.
On the faithfulness of vision transformer explanations
Junyi Wu, Weitai Kang, Hao Tang, Yuan Hong, and Yan Yan · 2024
Later among the works it cites.
Counterfactual explanations and how to find them: literature review and benchmarking
Guidotti R · 2024
Later among the works it cites.
Towards llm-guided causal explainability for blackbox text classifiers
Bhattacharjee A, Moraffah R, Garland J, and Liu H · 2024
Later among the works it cites.
Keep the faith: Faithful explanations in convolutional neural networks for case-based reasoning
Wolf T, Bongratz F, Rickmann A, Pölsterl S, and Wachinger C · 2024
Later among the works it cites.
A comprehensive survey of hallucination mitigation techniques in large language models
Tonmoy S · 2024
Later among the works it cites.
Syntaxshap: Syntax-aware explainability method for text generation
Amara K, Sevastjanova R, and El-Assady M · 2024
Later among the works it cites.
Multi-level explanations for generative language models
Paes L · 2024
Later among the works it cites.
Counterfactual explainable incremental prompt attack analysis on large language models
Shu D, Jin M, Chen T, Zhang C, and Zhang Y · 2024
Later among the works it cites.
Uncovering bias in large vision-language models with counterfactuals
Howard P, Bhiwandiwalla A, Fraser K, and Kiritchenko S · 2024
Later among the works it cites.
Using counterfactual tasks to evaluate the generality of analogical reasoning in large language models
Lewis M and Mitchell M · 2024
Later among the works it cites.
Evidence from counterfactual tasks supports emergent analogical reasoning in large language models
Webb T, Holyoak K, and Lu H · 2024
Later among the works it cites.
Unveiling factual recall behaviors of large language models through knowledge neurons
Wang Y, Chen Y, Wen W, Sheng Y, Li L, and Zeng D · 2024
Later among the works it cites.
Journey to the center of the knowledge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons
Chen Y, Cao P, Chen Y, Liu K, and Zhao J · 2024
Later among the works it cites.
Compositional chain-of-thought prompting for large multimodal models
Mitra C, Huang B, Darrell T, and Herzig R · 2024
Later among the works it cites.
How likely do llms with cot mimic human reasoning?
Bao G, Zhang H, Wang C, Yang L, and Zhang Y · 2024
Later among the works it cites.
Graph of thoughts: Solving elaborate problems with large language models
Besta M · 2024
Later among the works it cites.
Universal neurons in gpt2 language models
Gurnee W · 2024
Later among the works it cites.
Commonsensevis: Visualizing and understanding commonsense reasoning capabilities of natural language models
Wang X, Huang R, Jin Z, Fang T, and Qu H · 2024
Later among the works it cites.
Not all language model features are linear
Engels J, Michaud E, Liao I, Gurnee W, and Tegmark M · 2024
Later among the works it cites.
Enhancing machine learning model interpretability in intrusion detection systems through shap explanations and llm-generated descriptions
Khediri A, Slimi H, Yahiaoui A, Derdour M, Bendjenna H, and Ghenai C · 2024
Later among the works it cites.
Augmenting xai with llms: A case study in banking marketing recommendation
Castelnovo A · 2024
Later among the works it cites.
Qoexplainer: Mediating explainable quality of experience models with large language models
Wehner N, Feldhus N, Seufert M, Möller S, and Hoßfeld T · 2024
Later among the works it cites.
Llms for xai: Future directions for explaining explanations
Zytek A, Pidò S, and Veeramachaneni K · 2024
Later among the works it cites.
Large language model-based interpretable machine learning control in building energy systems
Zhang L and Chen Z · 2024
Later among the works it cites.
Pre-trained vision-language models learn discoverable visual concepts
Zang Y, Yun T, Tan H, Bui T, and Sun C · 2024
Later among the works it cites.
Vlg-cbm: Training concept bottleneck models with vision-language guidance
Srivastava D, Yan G, and Weng T.-W · 2024
Later among the works it cites.
Improving concept alignment in vision-language concept bottleneck models
Selvaraj N, Guo X, A, Kong .-K, and Kot A · 2024
Later among the works it cites.
Coarse-to-fine concept bottleneck models
Panousis K, Ienco D, and Marcos D · 2024
Later among the works it cites.
Sparse concept bottleneck models: Gumbel tricks in contrastive learning
Semenov A, Ivanov V, Beznosikov A, and Gasnikov A · 2024
Later among the works it cites.
The performance-interpretability trade-off: a comparative study of machine learning models
Assis A, Dantas J, and Andrade E · 2025
Closest in time.
Sparse additive models
Ravikumar P, Lafferty J, Liu H, and Wasserman L · 2025
Closest in time.
Google books
Quinlan J · 2025
Closest in time.
Interpretable neural network classification model using firstorder logic rules
Tuo H, Meng Z, Shi Z, and Zhang D · 2025
Closest in time.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Liu S · 2025
Closest in time.