Fetching the paper…
Reading the bibliography…
Recent sequence modeling approaches using selective state space sequence models, referred to as Mamba models, have seen a surge of interest.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
How to explain individual classification decisions
D. Baehrens, T. Schroeter, S. Harmeling, M. Kawanabe, K. Hansen, and K.-R. Müller · 2010
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. Müller, and W. Samek · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
S. R. Bowman, G. Angeli, C. Potts, and C. D. Manning · 2015
Earlier work this paper cites.
Interpretable explanations of black boxes by meaningful perturbation
R. C. Fong and A. Vedaldi · 2017
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
W. Samek, A. Binder, G. Montavon, S. Lapuschkin, and K.-R. Müller · 2017
Earlier work this paper cites.
Grad-CAM: Visual explanations from deep networks via gradient-based localization
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
A. Shrikumar, P. Greenside, and A. Kundaje · 2017
Earlier work this paper cites.
SmoothGrad: removing noise by adding noise
D. Smilkov, N. Thorat, B. Kim, F. B. Viégas, and M. Wattenberg · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
M. Sundararajan, A. Taly, and Q. Yan · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Visualizing deep neural network decisions: Prediction difference analysis
L. M. Zintgraf, T. S. Cohen, T. Adel, and M. Welling · 2017
Earlier work this paper cites.
Methods for interpreting and understanding deep neural networks
G. Montavon, W. Samek, and K.-R. Müller · 2018
Earlier work this paper cites.
CARER: contextualized affect representations for emotion recognition
E. Saravia, H. T. Liu, Y. Huang, J. Wu, and Y. Chen · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Explaining and interpreting LSTMs
L. Arras, J. Arjona-Medina, M. Widrich, G. Montavon, M. Gillhofer, K.-R. Müller, S. Hochreiter, and W. Samek · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
DARPA’s explainable artificial intelligence (XAI) program
D. Gunning · 2019
Earlier work this paper cites.
Attention is not Explanation
S. Jain and B. C. Wallace · 2019
Earlier work this paper cites.
Unmasking clever hans predictors and assessing what machines really learn
S. Lapuschkin, S. Wäldchen, A. Binder, G. Montavon, W. Samek, and K.-R. Müller · 2019
Earlier work this paper cites.
Layer-wise relevance propagation: An overview
G. Montavon, A. Binder, S. Lapuschkin, W. Samek, and K.-R. Müller · 2019
Earlier work this paper cites.
Attention is not not explanation
S. Wiegreffe and Y. Pinter · 2019
Earlier work this paper cites.
Quantifying attention flow in transformers
S. Abnar and W. Zuidema · 2020
Earlier work this paper cites.
Explainable artificial intelligence (XAI): concepts, taxonomies, opportunities and challenges toward responsible AI
A. B. Arrieta, N. D. Rodríguez, J. D. Ser, A. Bennetot, S. Tabik, A. Barbado, S. García, S. Gil-Lopez, D. Molina, R. Benjamins, R. Chatila, and F. Herrera · 2020
Earlier work this paper cites.
Building and interpreting deep similarity models
O. Eberle, J. Büttner, F. Kräutli, K.-R. Müller, M. Valleriani, and G. Montavon · 2020
Cited alongside, same era.
Effective gene expression prediction from sequence by integrating long-range interactions
Ž. Avsec, V. Agarwal, D. Visentin, J. R. Ledsam, A. Grabska-Barwinska, K. R. Taylor, Y. Assael, J. Jumper, P. Kohli, and D. R. Kelley · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Cited alongside, same era.
AST: Audio Spectrogram Transformer
Y. Gong, Y.-A. Chung, and J. Glass · 2021
Cited alongside, same era.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
A. Gu, I. Johnson, K. Goel, K. Saab, T. Dao, A. Rudra, and C. Ré · 2021
Cited alongside, same era.
No robots
N. Rajani, L. Tunstall, E. Beeching, N. Lambert, A. M. Rush, and T. Wolf · 2023
Later among the works it cites.
Diagonal state space augmented transformers for speech recognition
G. Saon, A. Gupta, and X. Cui · 2023
Later among the works it cites.
Inseq: An interpretability toolkit for sequence generation models
G. Sarti, N. Feldhus, L. Sickert, O. van der Wal, M. Nissim, and A. Bisazza · 2023
Later among the works it cites.
Simplified state space layers for sequence modeling
J. T. Smith, A. Warrington, and S. Linderman · 2023
Later among the works it cites.
Selective structured state-spaces for long-form video understanding
J. Wang, W. Zhu, P. Wang, X. Yu, L. Liu, M. Omar, and R. Hamid · 2023
Later among the works it cites.
Diffusion models without attention
J. N. Yan, J. Gu, and A. M. Rush · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Explaining deep neural networks and beyond: A review of methods and applications
W. Samek, G. Montavon, S. Lapuschkin, C. J. Anders, and K.-R. Müller · 2021
Cited alongside, same era.
Informer: Beyond efficient transformer for long sequence time-series forecasting
H. Zhou, S. Zhang, J. Peng, S. Zhang, J. Li, H. Xiong, and W. Zhang · 2021
Cited alongside, same era.
XAI for transformers: Better explanations through conservative propagation
A. Ali, T. Schnake, O. Eberle, G. Montavon, K.-R. Müller, and L. Wolf · 2022
Cited alongside, same era.
Disentangled explanations of neural network predictions by finding relevant subspaces
P. Chormai, J. Herrmann, K.-R. Müller, and G. Montavon · 2022
Cited alongside, same era.
Decision S4: Efficient sequence-based rl via state spaces layers
S. B. David, I. Zimerman, E. Nachmani, and L. Wolf · 2022
Cited alongside, same era.
Towards robust explanations for deep neural networks
A.-K. Dombrowski, C. J. Anders, K.-R. Müller, and P. Kessel · 2022
Cited alongside, same era.
Hungry hungry hippos: Towards language modeling with state space models
D. Y. Fu, T. Dao, K. K. Saab, A. W. Thomas, A. Rudra, and C. Ré · 2022
Cited alongside, same era.
Later among the works it cites.
AttnLRP: Attention-aware layer-wise relevance propagation for transformers
R. Achtibat, S. M. V. Hatefi, M. Dreyer, A. Jain, T. Wiegand, S. Lapuschkin, and W. Samek · 2024
Closest in time.
The hidden attention of mamba models
A. Ali, I. Zimerman, and L. Wolf · 2024
Closest in time.
Training-free long-context scaling of large language models
C. An, F. Huang, J. Zhang, S. Gong, X. Qiu, C. Zhou, and L. Kong · 2024
Closest in time.
BlackMamba: Mixture of experts for state-space models
Q. Anthony, Y. Tokpanov, P. Glorioso, and B. Millidge · 2024
Closest in time.
Graph Mamba: Towards learning on graphs with state space models
A. Behrouz and F. Hashemi · 2024
Closest in time.
Decoupling pixel flipping and occlusion strategy for consistent XAI benchmarks
S. Blücher, J. Vielhaben, and N. Strodthoff · 2024
Closest in time.
CLEX: Continuous length extrapolation for large language models
G. Chen, X. Li, Z. Meng, S. Liang, and L. Bing · 2024
Closest in time.
Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality
T. Dao and A. Gu · 2024
Closest in time.
H. Gong, L. Kang, Y. Wang, X. Wan, and H. Li · 2024
Closest in time.
RULER: What’s the real context size of your long-context language models?
C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, Y. Zhang, and B. Ginsburg · 2024
Closest in time.
Structured state space models for in-context reinforcement learning
C. Lu, Y. Schroecker, A. Gu, E. Parisotto, J. Foerster, S. Singh, and F. Behbahani · 2024
Closest in time.
Does transformer interpretability transfer to rnns?, 2024
G. Paulo, T. Marshall, and N. Belrose · 2024
Closest in time.
MoE-Mamba: Efficient selective state space models with mixture of experts
M. Pióro, K. Ciebiera, K. Król, J. Ludziejewski, and S. Jaszczur · 2024
Closest in time.
VM-UNet: Vision mamba UNet for medical image segmentation
J. Ruan and S. Xiang · 2024
Closest in time.
Z. Wang and C. Ma · 2024
Closest in time.
SegMamba: Long-range sequential modeling mamba for 3d medical image segmentation
Z. Xing, T. Ye, Y. Yang, G. Liu, and L. Zhu · 2024
Closest in time.
Vivim: a video vision mamba for medical video object segmentation
Y. Yang, Z. Xing, and L. Zhu · 2024
Closest in time.
Vision Mamba: Efficient visual representation learning with bidirectional state space model
L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang · 2024
Closest in time.
A unified implicit attention formulation for gated-linear recurrent sequence models, 2024
I. Zimerman, A. Ali, and L. Wolf · 2024
Closest in time.