Fetching the paper…
Reading the bibliography…
We present a theoretical framework for analyzing linear attention models through matrix-valued state space models (SSMs).
Generalized inverses of block triangular matrices
C. D. Meyer · 1970
Earlier work this paper cites.
Prefix sums and their applications
G. E. Blelloch · 1990
Earlier work this paper cites.
Differential equations driven by rough signals
T. J. Lyons · 1998
Earlier work this paper cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Deep signature transforms
P. Kidger, P. Bonnier, I. Perez Arribas, C. Salvi, and T. Lyons · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
P. Tillet, H. T. Kung, and D. Cox · 2019
Earlier work this paper cites.
Sig-sdes model for quantitative finance
I. P. Arribas, C. Salvi, and L. Szpruch · 2020
Earlier work this paper cites.
Transformers are RNNs: Fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Earlier work this paper cites.
Learning associative inference using fast weight memory
I. Schlag, T. Munkhdalai, and J. Schmidhuber · 2020
Earlier work this paper cites.
Sk-tree: a systematic malware detection algorithm on streaming trees via the signature kernel
T. Cochrane, P. Foster, V. Chhabra, M. Lemercier, T. Lyons, and C. Salvi · 2021
Earlier work this paper cites.
Neural rough differential equations for long time series
J. Morrill, C. Salvi, P. Kidger, and J. Foster · 2021
Earlier work this paper cites.
Rough paths, kernels, differential equations and an algebra of functions on streams
C. Salvi · 2021
Earlier work this paper cites.
Linear transformers are secretly fast weight programmers
I. Schlag, K. Irie, and J. Schmidhuber · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces, 2022
A. Gu, K. Goel, and C. Ré · 2022
Cited alongside, same era.
Transformer quality in linear time
W. Hua, Z. Dai, H. Liu, and Q. Le · 2022
Cited alongside, same era.
Neural signature kernels as infinite-width-depth-limits of controlled resnets
N. M. Cirone, M. Lemercier, and C. Salvi · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
T. Dao · 2023
Cited alongside, same era.
Sigdiffusions: Score-based diffusion models for long time series via log-signature embeddings
B. Barancikova, Z. Huang, and C. Salvi · 2024
Later among the works it cites.
Lecture notes on rough paths and applications to machine learning
T. Cass and C. Salvi · 2024
Later among the works it cites.
Theoretical foundations of deep selective state-space models
N. M. Cirone, A. Orvieto, B. Walker, C. Salvi, and T. Lyons · 2024
Later among the works it cites.
Transformers are ssms: generalized models and efficient algorithms through structured state space duality
T. Dao and A. Gu · 2024
Later among the works it cites.
Unlocking state-tracking in linear rnns through negative eigenvalues, 2024
R. Grazzi, J. Siems, J. K. H. Franke, A. Zela, F. Hutter, and M. Pontil · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Fermanian, T. Lyons, J. Morrill, and C. Salvi · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
A. Gu and T. Dao · 2023
Cited alongside, same era.
A neural rde approach for continuous-time non-markovian stochastic control problems
M. Hoglund, E. Ferrucci, C. Hernández, A. M. Gonzalez, C. Salvi, L. Sanchez-Betancourt, and Y. Zhang · 2023
Cited alongside, same era.
Optimal stopping via distribution regression: a higher rank signature approach
B. Horvath, M. Lemercier, C. Liu, T. Lyons, and C. Salvi · 2023
Cited alongside, same era.
Scaling transnormer to 175 billion parameters
Z. Qin, D. Li, W. Sun, W. Sun, X. Shen, X. Han, Y. Wei, B. Lv, F. Yuan, X. Luo, et al · 2023
Cited alongside, same era.
A structure theorem for streamed information
C. Salvi, J. Diehl, T. Lyons, R. Preiss, and J. Reizenstein · 2023
Cited alongside, same era.
Simplified state space layers for sequence modeling
J. T. Smith, A. Warrington, and S. Linderman · 2023
Cited alongside, same era.
Later among the works it cites.
Exact gradients for stochastic spiking neural networks driven by rough signals
C. Holberg and C. Salvi · 2024
Later among the works it cites.
Non-adversarial training of neural sdes with signature kernel scores
Z. Issa, B. Horvath, M. Lemercier, and C. Salvi · 2024
Later among the works it cites.
Signature kernel conditional independence tests in causal discovery for stochastic processes
G. Manten, C. Casolo, E. Ferrucci, S. W. Mogensen, C. Salvi, and N. Kilbertus · 2024
Later among the works it cites.
A path-dependent pde solver based on signature kernels
A. Pannier and C. Salvi · 2024
Later among the works it cites.
Sparse signature coefficient recovery via kernels
D. Shmelev and C. Salvi · 2024
Later among the works it cites.
Fla: A triton-based library for hardware-efficient implementations of linear attention mechanism, Jan. 2024
S. Yang and Y. Zhang · 2024
Later among the works it cites.
N. M. Cirone and C. Salvi · 2025
Closest in time.
Deltaproduct: Improving state-tracking in linear rnns via householder products, 2025
J. Siems, T. Carstensen, A. Zela, F. Hutter, M. Pontil, and R. Grazzi · 2025
Closest in time.