Fetching the paper…
Reading the bibliography…
We approach designing a state-space model for deep learning applications through its dual representation, the transfer function, and uncover a highly efficient sequence parallel inference algorithm that is state-free: unlike other proposed algorithms, state-free inference does not incur any significant memory or computational cost with an increase in state size.
On the theory of linear multi-loop feedback systems
Sandberg, I. W · 1963
Earlier work this paper cites.
Convolution algorithms
Burrus, C. S. and Parks, T · 1985
Earlier work this paper cites.
Matrix Analysis
Horn, R. A. and Johnson, C. R · 1985
Earlier work this paper cites.
Prefix sums and their applications
Blelloch, G. E · 1990
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Linear System Theory and Design
Chen, C.-T · 1998
Earlier work this paper cites.
Bounding the zeros of polynomials using the frobenius companion matrix partitioned by the cartesian decomposition
Alomari, M. W. and Chesneau, C · 1999
Earlier work this paper cites.
Discrete-Time Signal Processing
Oppenheim, A. V., Schafer, R. W., and Buck, J. R · 1999
Earlier work this paper cites.
Structured Matrices and Polynomials: Unified Superfast Algorithms
Pan, V. Y · 2001
Earlier work this paper cites.
Digital image processing
Gonzalez, R. C. and Woods, R. E · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
The ACL Anthology network corpus
Radev, D. R., Muthukrishnan, P., and Qazvinian, V · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Unitary evolution recurrent neural networks
Arjovsky, M., Shah, A., and Bengio, Y · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Language modeling with gated convolutional networks
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Cited alongside, same era.
Fast convolution and filtering
Selesnick, I. W. and Burrus, C. S · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Cited alongside, same era.
Learning long-range spatial dependencies with horizontal gated recurrent units
Linsley, D., Kim, J., Veerabadran, V., Windolf, C., and Serre, T · 2018
Diagonal state spaces are as effective as structured state spaces
Gupta, A., Gu, A., and Berant, J · 2022
Later among the works it cites.
Ckconv: Continuous kernel convolution for sequential data
Romero, D. W., Kuzina, A., Bekkers, E. J., Tomczak, J. M., and Hoogendoorn, M · 2022
Later among the works it cites.
Hungry Hungry Hippos: Towards language modeling with state space models
Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A., and Ré, C · 2023
Later among the works it cites.
A framework for few-shot language model evaluation, 12 2023
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces, 2023
Gu, A. and Dao, T · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Parallelizing linear recurrent neural nets over sequence length
Martin, E. and Cundy, C · 2018
Cited alongside, same era.
ListOps: A diagnostic dataset for latent tree learning
Nangia, N. and Bowman, S · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Cited alongside, same era.
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Neural controlled differential equations for irregular time series
Kidger, P., Morrill, J., Foster, J., and Lyons, T · 2020
Cited alongside, same era.
How to train your HIPPO: State space models with generalized orthogonal basis projections
Gu, A., Johnson, I., Timalsina, A., Rudra, A., and Re, C · 2023
Later among the works it cites.
Liquid structural state-space models
Hasani, R., Lechner, M., Wang, T.-H., Chahine, M., Amini, A., and Rus, D · 2023
Later among the works it cites.
Gaussian error linear units (gelus), 2023
Hendrycks, D. and Gimpel, K · 2023
Later among the works it cites.
Gateloop: Fully data-controlled linear recurrence for sequence modeling, 2023
Katsch, T · 2023
Later among the works it cites.
Laughing hyena distillery: Extracting compact recurrences from convolutions
Massaroli, S., Poli, M., Fu, D. Y., Kumbong, H., Parnichkun, R. N., Romero, D. W., Timalsina, A., McIntyre, Q., Chen, B., Rudra, A., Zhang, C., Re, C., Ermon, S., and Bengio, Y · 2023
Later among the works it cites.
Resurrecting recurrent neural networks for long sequences
Orvieto, A., Smith, S. L., Gu, A., Fernando, A., Gulcehre, C., Pascanu, R., and De, S · 2023
Later among the works it cites.
Sparse modular activation for efficient sequence modeling, 2023
Ren, L., Liu, Y., Wang, S., Xu, Y., Zhu, C., and Zhai, C · 2023
Later among the works it cites.
Simplified state space layers for sequence modeling
Smith, J. T., Warrington, A., and Linderman, S · 2023
Later among the works it cites.
Effectively modeling time series with simple discrete state spaces
Zhang, M., Saab, K., Poli, M., Dao, T., Goel, K., and Ré, C · 2023
Later among the works it cites.
In-context language learning: Architectures and algorithms, 2024
Akyürek, E., Wang, B., Kim, Y., and Andreas, J · 2024
Closest in time.
Never train from scratch: Fair comparison of long-sequence models requires data-driven priors
Amos, I., Berant, J., and Gupta, A · 2024
Closest in time.
Understanding in-context learning in transformers and LLMs by learning to learn discrete functions
Bhattamishra, S., Patel, A., Blunsom, P., and Kanade, V · 2024
Closest in time.
FlashFFTConv: Efficient convolutions for long sequences with tensor cores
Fu, D. Y., Kumbong, H., Nguyen, E., and Ré, C · 2024
Closest in time.