Fetching the paper…
Reading the bibliography…
State space models (SSMs) have high performance on long sequence modeling but require sophisticated initialization techniques and specialized implementations for high quality and runtime performance.
An algorithm for the machine calculation of complex fourier series
Cooley, J. W. and Tukey, J. W · 1965
Earlier work this paper cites.
Applications of digital signal processing
Oppenheim, A. V · 1978
Earlier work this paper cites.
Displacement ranks of matrices and linear equations
Kailath, T., Kung, S.-Y., and Morf, M · 1979
Earlier work this paper cites.
The fast Fourier transform and its applications
Brigham, E. O · 1988
Earlier work this paper cites.
FFTs in external or hierarchical memory
Bailey, D. H · 1990
Earlier work this paper cites.
Random Butterfly Transformations with Applications in Computational Linear Algebra
Parker, D · 1995
Earlier work this paper cites.
The scientist and engineer’s guide to digital signal processing, 1997
Smith, S. W. et al · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Inside the FFT black box: serial and parallel fast Fourier transform algorithms
Chu, E. and George, A · 1999
Earlier work this paper cites.
On a new class of structured matrices
Eidelman, Y. and Gohberg, I · 1999
Earlier work this paper cites.
Discrete-time signal processing. Vol. 2
Oppenheim, A. V., Buck, J. R., and Schafer, R. W · 2001
Earlier work this paper cites.
Parallel fft algorithms on network-on-chips
Bahn, J. H., Yang, J. S., Hu, W.-H., and Bagherzadeh, N · 2009
Earlier work this paper cites.
Pipelined parallel fft architectures via folding transformation
Ayinala, M., Brown, M., and Parhi, K. K · 2011
Earlier work this paper cites.
Freesurfer
Fischl, B · 2012
Earlier work this paper cites.
Function in the human connectome: task-fmri and individual differences in behavior
Barch, D. M., Burgess, G. C., Harms, M. P., Petersen, S. E., Schlaggar, B. L., Corbetta, M., Glasser, M. F., Curtiss, S., Dixit, S., Feldt, C., et al · 2013
Earlier work this paper cites.
Structured transforms for small-footprint deep learning
Sindhwani, V., Sainath, T., and Kumar, S · 2015
Earlier work this paper cites.
Cooley-tukey fft algorithms
Bekele, A · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Dong, X., Chen, S., and Pan, S · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2017
Earlier work this paper cites.
Runtime neural pruning
Lin, J., Rao, Y., Lu, J., and Zhou, J · 2017
Earlier work this paper cites.
Nvidia Tesla V100 GPU architecture, 2017
NVIDIA · 2017
Earlier work this paper cites.
Long-term temporal convolutions for action recognition
Varol, G., Laptev, I., and Schmid, C · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
A two-pronged progress in structured dense matrix vector multiplication
De Sa, C., Cu, A., Puttagunta, R., Ré, C., and Rudra, A · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Cited alongside, same era.
Quadrature-based features for kernel approximation
Munkhoeva, M., Kapushev, Y., Burnaev, E., and Oseledets, I · 2018
Cited alongside, same era.
Unifying orthogonal monte carlo methods
Choromanski, K., Rowland, M., Chen, W., and Weller, A · 2019
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J. G., Le, Q., and Salakhutdinov, R · 2019
Adaptive fourier neural operators: Efficient token mixers for transformers
Guibas, J., Mardani, M., Li, Z., Tao, A., Anandkumar, A., and Catanzaro, B · 2021
Later among the works it cites.
Fnet: Mixing tokens with fourier transforms
Lee-Thorp, J., Ainslie, J., Eckstein, I., and Ontanon, S · 2021
Later among the works it cites.
tcfft: Accelerating half-precision fft through tensor cores
Li, B., Cheng, S., and Lin, J · 2021
Later among the works it cites.
Evit: Expediting vision transformers via token reorganizations
Liang, Y., Chongjian, G., Tong, Z., Song, Y., Wang, J., and Xie, P · 2021
Later among the works it cites.
Deformable butterfly: A highly structured and sparse linear transform
Lin, R., Ran, J., Chiu, K. H., Chesi, G., and Wong, N · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning fast algorithms for linear transforms using butterfly factorizations
Dao, T., Gu, A., Eichhorn, M., Rudra, A., and Ré, C · 2019
Cited alongside, same era.
Stabilizing the lottery ticket hypothesis
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 2019
Cited alongside, same era.
Openwebtext corpus, 2019
Gokaslan, A., Cohen, V., Pavlick, E., and Tellex, S · 2019
Cited alongside, same era.
Functional boundaries in the human cerebellum revealed by a multi-domain task battery
King, M., Hernandez-Castillo, C. R., Poldrack, R. A., Ivry, R. B., and Diedrichsen, J · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Cited alongside, same era.
Fine-grain atlases of functional modes for fmri analysis
Dadi, K., Varoquaux, G., Machlouzarides-Shalit, A., Gorgolewski, K. J., Wassermann, D., Thirion, B., and Mensch, A · 2020
Cited alongside, same era.
Informer: Beyond efficient transformer for long sequence time-series forecasting
Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W · 2021
Later among the works it cites.
It’s raw! audio generation with state-space models
Goel, K., Gu, A., Donahue, C., and Ré, C · 2022
Later among the works it cites.
Diagonal state spaces are as effective as structured state spaces
Gupta, A., Gu, A., and Berant, J · 2022
Later among the works it cites.
Liquid structural state-space models
Hasani, R., Lechner, M., Wang, T.-H., Chahine, M., Amini, A., and Rus, D · 2022
Later among the works it cites.
Long movie clip classification with state-space video models
Islam, M. M. and Bertasius, G · 2022
Later among the works it cites.
What makes convolutional models great on long sequence modeling?
Li, Y., Cai, T., Zhang, Y., Chen, D., and Dey, D · 2022
Later among the works it cites.
Mega: moving average equipped gated attention
Ma, X., Zhou, C., Kong, X., He, J., Gui, L., Neubig, G., May, J., and Zettlemoyer, L · 2022
Later among the works it cites.
Long range language modeling via gated state spaces
Mehta, H., Gupta, A., Cutkosky, A., and Neyshabur, B · 2022
Later among the works it cites.
S4nd: Modeling images and videos as multidimensional signals with state spaces
Nguyen, E., Goel, K., Gu, A., Downs, G., Shah, P., Dao, T., Baccus, S., and Ré, C · 2022
Later among the works it cites.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Later among the works it cites.
Simplified state space layers for sequence modeling
Smith, J. T., Warrington, A., and Linderman, S. W · 2022
Later among the works it cites.
Tang, S., Dunnmon, J. A., Qu, L., Saab, K. K., Lee-Messer, C., and Rubin, D. L · 2022
Later among the works it cites.
Self-supervised learning of brain dynamics from broad neuroimaging data
Thomas, A. W., Ré, C., and Poldrack, R. A · 2022
Later among the works it cites.
Trockman, A. and Kolter, J. Z · 2022
Later among the works it cites.
Memorizing transformers
Wu, Y., Rabe, M. N., Hutchins, D., and Szegedy, C · 2022
Later among the works it cites.
Deep latent state space models for time-series generation
Zhou, L., Poli, M., Xu, W., Massaroli, S., and Ermon, S · 2022
Later among the works it cites.
Effectively modeling time series with simple discrete state spaces
Zhang, M., Saab, K. K., Poli, M., Goel, K., Dao, T., and Ré, C · 2023
Closest in time.