Fetching the paper…
Reading the bibliography…
Large Language Models are prone to biased predictions and hallucinations, underlining the paramount importance of understanding their model-internal reasoning process.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. (2019) · 1910
Earlier work this paper cites.
A linear approximation method for the shapley value
Fatima, S. S., Wooldridge, M., and Jennings, N. R. (2008) · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Captum: A unified and generic model interpretability library for pytorch
Kokhlikyan, N., Miglani, V., Martin, M., Wang, E., Alsallakh, B., Reynolds, J., Melnikov, A., Kliushkina, N., Araya, C., Yan, S., et al. (2020) · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. (2011) · 2011
Earlier work this paper cites.
Slic superpixels compared to state-of-the-art superpixel methods
Achanta, R., Shaji, A., Smith, K., Lucchi, A., Fua, P., and Süsstrunk, S. (2012) · 2012
Earlier work this paper cites.
Deep inside convolutional networks: visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A. (2014) · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R. (2014) · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W. (2015) · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Earlier work this paper cites.
Layer-wise relevance propagation for neural networks with local renormalization layers
Binder, A., Montavon, G., Lapuschkin, S., Müller, K.-R., and Samek, W. (2016) · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C. (2016) · 2016
Earlier work this paper cites.
Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
Nguyen, A., Dosovitskiy, A., Yosinski, J., Brox, T., and Clune, J. (2016) · 2016
Earlier work this paper cites.
”why should I trust you?”: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C. (2016) · 2016
Earlier work this paper cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
Balduzzi, D., Frean, M., Leary, L., Lewis, J., Ma, K. W.-D., and McWilliams, B. (2017) · 2017
Earlier work this paper cites.
Visualizing and understanding neural machine translation
Ding, Y., Liu, Y., Luan, H., and Sun, M. (2017) · 2017
Earlier work this paper cites.
Interpretable explanations of black boxes by meaningful perturbation
Fong, R. C. and Vedaldi, A. (2017) · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S. (2017) · 2017
Earlier work this paper cites.
Explaining nonlinear classification decisions with deep taylor decomposition
Montavon, G., Lapuschkin, S., Binder, A., Samek, W., and Müller, K.-R. (2017) · 2017
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Samek, W., Binder, A., Montavon, G., Lapuschkin, S., and Müller, K.-R. (2017) · 2017
Earlier work this paper cites.
Improving the compositionality of word embeddings
Scheepers, T. (2017) · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2017) · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A. (2017) · 2017
Cited alongside, same era.
Smoothgrad: removing noise by adding noise
Smilkov, D., Thorat, N., Kim, B., Viégas, F., and Wattenberg, M. (2017) · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q. (2017) · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Explaining image classifiers by counterfactual generation
Chang, C.-H., Creager, E., Goldenberg, A., and Duvenaud, D. (2018) · 2018
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al. (2022) · 2022
Later among the works it cites.
Knowledge neurons in pretrained transformers
Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F. (2022) · 2022
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C. (2022) · 2022
Later among the works it cites.
Towards robust explanations for deep neural networks
Dombrowski, A.-K., Anders, C. J., Müller, K.-R., and Kessel, P. (2022) · 2022
Later among the works it cites.
A review of sparse expert models in deep learning
Fedus, W., Dean, J., and Zoph, B. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guidotti, R., Monreale, A., Ruggieri, S., Pedreschi, D., Turini, F., and Giannotti, F. (2018) · 2018
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for squad
Rajpurkar, P., Jia, R., and Liang, P. (2018) · 2018
Cited alongside, same era.
What does bert look at? an analysis of bert’s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D. (2019) · 2019
Cited alongside, same era.
Layer-wise relevance propagation: an overview
Montavon, G., Binder, A., Lapuschkin, S., Samek, W., and Müller, K.-R. (2019) · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019) · 2019
Cited alongside, same era.
Attention is not not explanation
Wiegreffe, S. and Pinter, Y. (2019) · 2019
Cited alongside, same era.
Root mean square layer normalization
Zhang, B. and Sennrich, R. (2019) · 2019
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Geva, M., Caciularu, A., Wang, K., and Goldberg, Y. (2022) · 2022
Later among the works it cites.
From attribution maps to human-understandable explanations through concept relevance propagation
Achtibat, R., Dreyer, M., Eisenbraun, I., Bosse, S., Wiegand, T., Samek, W., and Lapuschkin, S. (2023) · 2023
Later among the works it cites.
Feature perturbation augmentation for reliable evaluation of importance estimators in neural networks
Brocki, L. and Chung, N. C. (2023) · 2023
Later among the works it cites.
Atman: Understanding transformer predictions through memory efficient attention manipulation
Deb, M., Deiseroth, B., Weinbach, S., Schramowski, P., and Kersting, K. (2023) · 2023
Later among the works it cites.
Exploring explainability for vision transformers
Gildenblat, J. (2020. Accessed on Dec 01, 2023) · 2023
Later among the works it cites.
Quantus: An explainable ai toolkit for responsible evaluation of neural network explanations and beyond
Hedström, A., Weber, L., Krakowczyk, D., Bareeva, D., Motzkus, F., Samek, W., Lapuschkin, S., and Höhne, M. M. M. (2023) · 2023
Later among the works it cites.
Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., et al. (2023) · 2023
Later among the works it cites.
Textbooks are all you need ii: phi-1.5 technical report
Li, Y., Bubeck, S., Eldan, R., Del Giorno, A., Gunasekar, S., and Lee, Y. T. (2023) · 2023
Later among the works it cites.
Using captum to explain generative language models
Miglani, V., Yang, A., Markosyan, A., Garcia-Olano, D., and Kokhlikyan, N. (2023) · 2023
Later among the works it cites.
Optimizing explanations by network canonization and hyperparameter search
Pahde, F., Yolcu, G. Ü., Binder, A., Samek, W., and Lapuschkin, S. (2023) · 2023
Later among the works it cites.
Zeroscrolls: A zero-shot benchmark for long text understanding
Shaham, U., Ivgi, M., Efrat, A., Berant, J., and Levy, O. (2023) · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Later among the works it cites.
Neurons in large language models: Dead, n-gram, positional
Voita, E., Ferrando, J., and Nalmpantis, C. (2023) · 2023
Later among the works it cites.
Wikimedia downloads
Wikimedia Foundation (2023. Accessed on Dec 01, 2023) · 2023
Later among the works it cites.
Decoupling pixel flipping and occlusion strategy for consistent xai benchmarks
Blücher, S., Vielhaben, J., and Strodthoff, N. (2024) · 2024
Closest in time.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2024) · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. (2024) · 2024
Closest in time.