Fetching the paper…
Reading the bibliography…
In this paper, we analyze the computational limitations of Mamba and State-space Models (SSMs) by using the circuit complexity framework.
Neural networks and physical systems with emergent collective computational abilities
John J Hopfield · 1982
Earlier work this paper cites.
Bounded-width polynomial-size branching programs recognize exactly those languages in nc
David A Barrington · 1986
Earlier work this paper cites.
The boolean formula value problem is in alogtime
Samuel R Buss · 1987
Earlier work this paper cites.
An optimal parallel algorithm for formula evaluation
S Buss, S Cook, Arvind Gupta, and Vijaya Ramachandran · 1992
Earlier work this paper cites.
Time, hardware, and uniformity
D Mix Barrington and Neil Immerman · 1994
Earlier work this paper cites.
Long short-term memory
S Hochreiter · 1997
Earlier work this paper cites.
Efficient threshold circuits for power series
Alexis Maciel and Denis Thérien · 1999
Earlier work this paper cites.
Introduction to circuit complexity: a uniform approach
Heribert Vollmer · 1999
Earlier work this paper cites.
Uniform constant-depth threshold circuits for division and iterated multiplication
William Hesse, Eric Allender, and David A Mix Barrington · 2002
Earlier work this paper cites.
Computational Complexity: A Modern Approach
Sanjeev Arora and Boaz Barak · 2009
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho · 2014
Earlier work this paper cites.
Long short-term memory recurrent neural network architectures for large scale acoustic modeling
H Sak, A Senior, and F Beaufays · 2014
Earlier work this paper cites.
An efficient state-space model for joint tempo and meter tracking
Florian Krebs, Sebastian Böck, and Gerhard Widmer · 2015
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
On the turing completeness of modern neural network architectures
Jorge Pérez, Javier Marinković, and Pablo Barceló · 2019
Earlier work this paper cites.
Fast training of deep lstm networks
Wen Yu, Xiaoou Li, and Jesus Gonzalez · 2019
Earlier work this paper cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Earlier work this paper cites.
Instahide’s sample complexity when mixing two private images
Baihe Huang, Zhao Song, Runzhou Tao, Junze Yin, Ruizhe Zhang, and Danyang Zhuo · 2020
Earlier work this paper cites.
Lévy state-space models for tracking and intent prediction of highly maneuverable objects
Runze Gan, Bashar I Ahmad, and Simon J Godsill · 2021
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Ré · 2021
Earlier work this paper cites.
Optimal-degree polynomial approximations for exponentials and gaussian kernel density estimation
Amol Aggarwal and Josh Alman · 2022
Earlier work this paper cites.
Improving time series forecasting using lstm and attention models
Hossein Abbasimehr and Reza Paki · 2022
Earlier work this paper cites.
Sample complexity of learning parametric quantum circuits
Haoyuan Cai, Qi Ye, and Dong-Ling Deng · 2022
Earlier work this paper cites.
A nearly optimal size coreset algorithm with nearly linear time
Yichuan Deng, Zhao Song, Yitan Wang, and Yuanyuan Yang · 2022
Earlier work this paper cites.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank · 2022
Earlier work this paper cites.
Sublinear time algorithm for online weighted bipartite matching
Hang Hu, Zhao Song, Runzhou Tao, Zhaozhuo Xu, Junze Yin, and Danyang Zhuo · 2022
Earlier work this paper cites.
A dynamic fast gaussian transform
Baihe Huang, Zhao Song, Omri Weinstein, Junze Yin, Hengjie Zhang, and Ruizhe Zhang · 2022
Earlier work this paper cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2022
Earlier work this paper cites.
A faster k k -means++ algorithm
Jiehao Liang, Somdeb Sarkhel, Zhao Song, Chenbo Yin, Junze Yin, and Danyang Zhuo · 2022
Earlier work this paper cites.
Dynamic maintenance of kernel density estimation data structure: From practice to theory
Jiehao Liang, Zhao Song, Zhaozhuo Xu, Junze Yin, and Danyang Zhuo · 2022
Earlier work this paper cites.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A Smith · 2022
Earlier work this paper cites.
A multi-layer extreme learning machine refined by sparrow search algorithm and weighted mean filter for short-term multi-step wind speed forecasting
Haochen Zhang, Zhiyun Peng, Junjie Tang, Ming Dong, Ke Wang, and Wenyuan Li · 2022
Earlier work this paper cites.
Masked hard-attention transformers and boolean rasp recognize exactly the star-free languages
Dana Angluin, David Chiang, and Andy Yang · 2023
Earlier work this paper cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Earlier work this paper cites.
Federated empirical risk minimization via second-order method
Song Bian, Zhao Song, and Junze Yin · 2023
Earlier work this paper cites.
Tighter bounds on the expressivity of transformer encoders
David Chiang, Peter Cholak, and Anand Pillay · 2023
Earlier work this paper cites.
Query complexity of active learning for function family with nearly orthogonal basis
Xiang Chen, Zhao Song, Baocheng Sun, Junze Yin, and Danyang Zhuo · 2023
Earlier work this paper cites.
Yichuan Deng, Sridhar Mahadevan, and Zhao Song · 2023
Earlier work this paper cites.
Faster robust tensor power method for arbitrary order
Yichuan Deng, Zhao Song, and Junze Yin · 2023
Earlier work this paper cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Earlier work this paper cites.
An over-parameterized exponential regression
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Earlier work this paper cites.
Yeqi Gao, Zhao Song, Weixin Wang, and Junze Yin · 2023
Earlier work this paper cites.
Yeqi Gao, Zhao Song, and Shenghao Xie · 2023
Cited alongside, same era.
Gradientcoin: A peer-to-peer decentralized large language models
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Cited alongside, same era.
An iterative algorithm for rescaled hyperbolic functions regression
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Cited alongside, same era.
On sparse modern hopfield model
Jerry Yao-Chieh Hu, Donglin Yang, Dennis Wu, Chenwei Xu, Bo-Yu Chen, and Han Liu · 2023
Cited alongside, same era.
Local convergence of approximate newton method for two layer nonlinear regression
Zhihang Li, Zhao Song, Zifan Wang, and Junze Yin · 2023
Yingyu Liang, Heshan Liu, Zhenmei Shi, Zhao Song, and Junze Yin · 2024
Closest in time.
A tighter complexity analysis of sparsegpt
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
Fast second-order method for neural networks under small treewidth setting
Xiaoyu Li, Jiangxuan Long, Zhao Song, and Tianyi Zhou · 2024
Closest in time.
Mamba4rec: Towards efficient sequential recommendation with selective state space models
Chengkai Liu, Jianghao Lin, Jianling Wang, Hanzhou Liu, and James Caverlee · 2024
Closest in time.
Uniform last-iterate guarantee for bandits and reinforcement learning
Junyan Liu, Yunfan Li, Ruosong Wang, and Lin Yang · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Cited alongside, same era.
Low-switching policy gradient with exploration via online sensitivity sampling
Yunfan Li, Yiran Wang, Yu Cheng, and Lin Yang · 2023
Cited alongside, same era.
The parallelism tradeoff: Limitations of log-precision transformers
William Merrill and Ashish Sabharwal · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Ritwik Sinha, Zhao Song, and Tianyi Zhou · 2023
Cited alongside, same era.
A unified scheme of resnet and softmax
Zhao Song, Weixin Wang, and Junze Yin · 2023
Cited alongside, same era.
Fast and efficient matching algorithm with deadline instances
Zhao Song, Weixin Wang, Chenbo Yin, and Junze Yin · 2023
Cited alongside, same era.
Closest in time.
Xyscannet: An interpretable state space model for perceptual image deblurring
Hanzhou Liu, Chengkai Liu, Jiacong Xu, Peng Jiang, and Mi Lu · 2024
Closest in time.
Achieving near-optimal regret for bandit algorithms with uniform last-iterate guarantee
Junyan Liu, Yunfan Li, and Lin Yang · 2024
Closest in time.
Chain of thought empowers transformers to solve inherently serial problems
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma · 2024
Closest in time.
Multi-layer transformers gradient can be approximated in almost linear time
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Multi-layer transformers gradient can be approximated in almost linear time
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Differential privacy of cross-attention with provable guarantee
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Tensor attention training: Provably efficient learning of higher-order transformers
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
How to inverting the leverage score distribution?
Zhihang Li, Zhao Song, Weixin Wang, Junze Yin, and Zheng Yu · 2024
Closest in time.
Inverting the leverage score gradient: An efficient approximate newton method
Chenyang Li, Zhao Song, Zhaoxing Xu, and Junze Yin · 2024
Closest in time.
On the model-misspecification in reinforcement learning
Yunfan Li and Lin Yang · 2024
Closest in time.
Introducing llama 3.1: Our most capable models to date, 2024
Meta · 2024
Closest in time.
The illusion of state in state-space models
William Merrill, Jackson Petty, and Ashish Sabharwal · 2024
Closest in time.
Hello gpt-4o, 2024
OpenAI · 2024
Closest in time.
Introducing openai o1-preview, 2024
OpenAI · 2024
Closest in time.
Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, and Shafiq Joty · 2024
Closest in time.
Solving attention kernel regression problem via pre-conditioner
Zhao Song, Junze Yin, and Lichen Zhang · 2024
Closest in time.
Fast dynamic sampling for determinantal point processes
Zhao Song, Junze Yin, Lichen Zhang, and Ruizhe Zhang · 2024
Closest in time.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Dennis Wu, Jerry Yao-Chieh Hu, Weijian Li, Bo-Yu Chen, and Han Liu · 2024
Closest in time.
Is a picture worth a thousand words? delving into spatial reasoning for vision language models
Jiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet, Xin Wang, Yixuan Li, and Neel Joshi · 2024
Closest in time.
Bishop: Bi-directional cellular learning for tabular data with generalized sparse modern hopfield model
Chenwei Xu, Yu-Chao Huang, Jerry Yao-Chieh Hu, Weijian Li, Ammar Gilani, Hsi-Sheng Goan, and Han Liu · 2024
Closest in time.
Limits of deep learning: Sequence modeling through the lens of complexity theory
Nikola Zubić, Federico Soldá, Aurelio Sulser, and Davide Scaramuzza · 2024
Closest in time.
Transformers in uniform 𝖳𝖢 0 \mathsf{TC}^{0}
David Chiang · 2025
Closest in time.
Yuefan Cao, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Jiahao Zhang · 2025
Closest in time.
Bypassing the exponential dependency: Looped transformers efficiently learn in-context by multi-step gradient descent
Bo Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Universal approximation of visual autoregressive transformers
Yifang Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Circuit complexity bounds for visual autoregressive model
Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Curse of attention: A kernel-based perspective for why transformers fail to generalize on time series forecasting and beyond
Yekun Ke, Yingyu Liang, Zhenmei Shi, Zhao Song, and Chiwun Yang · 2025
Closest in time.
Neural algorithmic reasoning for hypergraphs with looped transformers
Xiaoyu Li, Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, and Zhen Zhuang · 2025
Closest in time.
Fourier circuits in neural networks and transformers: A case study of modular arithmetic with multiple inputs
Chenyang Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Tianyi Zhou · 2025
Closest in time.
On the computational capability of graph neural networks: A circuit complexity bound perspective
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song, Wei Wang, and Jiahao Zhang · 2025
Closest in time.
Beyond linear approximations: A novel pruning approach for attention matrix
Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2025
Closest in time.
Looped relu mlps may be all you need as practical programmable computers
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2025
Closest in time.
Lazydit: Lazy learning for the acceleration of diffusion transformers
Xuan Shen, Zhao Song, Yufa Zhou, Bo Chen, Yanyu Li, Yifan Gong, Kai Zhang, Hao Tan, Jason Kuen, Henghui Ding, Zhihao Shu, Wei Niu, Pu Zhao, Yanzhi Wang, and Jiuxiang Gu · 2025
Closest in time.
Numerical pruning for efficient autoregressive models
Xuan Shen, Zhao Song, Yufa Zhou, Bo Chen, Jing Liu, Ruiyi Zhang, Ryan A. Rossi, Hao Tan, Tong Yu, Xiang Chen, Yufan Zhou, Tong Sun, Pu Zhao, Yanzhi Wang, and Jiuxiang Gu · 2025
Closest in time.
Efficient alternating minimization with applications to weighted low rank approximation
Zhao Song, Mingquan Ye, Junze Yin, and Lichen Zhang · 2025
Closest in time.
Statistical guarantees for lifelong reinforcement learning using pac-bayesian theory
Zhi Zhang, Chris Chow, Yasi Zhang, Yanchao Sun, Haochen Zhang, Eric Hanchen Jiang, Han Liu, Furong Huang, Yuchen Cui, and Oscar Hernan Madrid Padilla · 2025
Closest in time.