Fetching the paper…
Reading the bibliography…
Deep neural networks based on linear RNNs interleaved with position-wise MLPs are gaining traction as competitive approaches for sequence modeling.
Zur theorie der orthogonalen funktionensysteme
Haar, A · 1911
Earlier work this paper cites.
Über den approximationssatz von Weierstrass
Müntz, C. H · 1914
Earlier work this paper cites.
Über die approximation stetiger funktionen durch lineare aggregate von potenzen
Szász, O · 1916
Earlier work this paper cites.
Elements of physical biology
Lotka, A. J · 1925
Earlier work this paper cites.
Variations and fluctuations of the number of individuals in animal species living together
Volterra, V · 1928
Earlier work this paper cites.
On the representation of continuous functions of many variables by superposition of continuous functions of one variable and addition
Kolmogorov, A. N · 1957
Earlier work this paper cites.
Deterministic nonperiodic flow
Lorenz, E. N · 1963
Earlier work this paper cites.
Optimally conditioned vandermonde matrices
Gautschi, W · 1975
Earlier work this paper cites.
Fading memory and the problem of approximating nonlinear operators with volterra series
Boyd, S. and Chua, L · 1985
Earlier work this paper cites.
Extensions of lipschitz maps into banach spaces
Johnson, W. B., Lindenstrauss, J., and Schechtman, G · 1986
Earlier work this paper cites.
Lower bounds for the condition number of vandermonde matrices
Gautschi, W. and Inglese, G · 1987
Earlier work this paper cites.
On the approximate realization of continuous mappings by neural networks
Funahashi, K.-I · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H · 1989
Earlier work this paper cites.
Prefix sums and their applications, 1990
Blelloch, G. E · 1990
Earlier work this paper cites.
Vandermonde matrices on the circle: spectral properties and conditioning
Córdova, A., Gautschi, W., and Ruscheweyh, S · 1990
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Hornik, K · 1991
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Barron, A. R · 1993
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
The mnist database of handwritten digits
LeCun, Y · 1998
Earlier work this paper cites.
Approximation theory of the MLP model in neural networks
Pinkus, A · 1999
Earlier work this paper cites.
Learning overcomplete representations
Lewicki, M. S. and Sejnowski, T. J · 2000
Earlier work this paper cites.
Beurling–malliavin multiplier theorem: the seventh proof
Mashreghi, J., Nazarov, F., and Havin, V · 2006
Earlier work this paper cites.
Recurrent neural networks are universal approximators
Schäfer, A. M. and Zimmermann, H. G · 2006
Earlier work this paper cites.
Bayesian ranking of biochemical system models
Vyshemirsky, V. and Girolami, M. A · 2008
Earlier work this paper cites.
Fourier analysis and its applications , volume 4
Folland, G. B · 2009
Earlier work this paper cites.
Matrix analysis , volume 169
Bhatia, R · 2013
Earlier work this paper cites.
Ode parameter inference using adaptive gradient matching with gaussian processes
Dondelinger, F., Husmeier, D., Rogers, S., and Filippone, M · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
Arjovsky, M., Shah, A., and Bengio, Y · 2016
Cited alongside, same era.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Diagonal state spaces are as effective as structured state spaces
Gupta, A., Gu, A., and Berant, J · 2022
Later among the works it cites.
Liquid structural state-space models
Hasani, R., Lechner, M., Wang, T.-H., Chahine, M., Amini, A., and Rus, D · 2022
Later among the works it cites.
Mega: moving average equipped gated attention
Ma, X., Zhou, C., Kong, X., He, J., Gui, L., Neubig, G., May, J., and Zettlemoyer, L · 2022
Later among the works it cites.
S4nd: Modeling images and videos as multidimensional signals using state spaces
Nguyen, E., Goel, K., Gu, A., Downs, G. W., Shah, P., Dao, T., Baccus, S. A., and Ré, C · 2022
Later among the works it cites.
Simplified state space layers for sequence modeling
Smith, J. T., Warrington, A., and Linderman, S. W · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Cited alongside, same era.
The expressive power of neural networks: A view from the width
Lu, Z., Pu, H., Wang, F., Hu, Z., and Wang, L · 2017
Cited alongside, same era.
Parallelizing linear recurrent neural nets over sequence length
Martin, E. and Cundy, C · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Online learning of long-range dependencies, 2023b
Zucchet, N., Meier, R., Schug, S., Mujika, A., and Sacramento, J · 2017
Cited alongside, same era.
Orthogonal recurrent neural networks with scaled cayley transform
Helfrich, K., Willmott, D., and Ye, Q · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Cited alongside, same era.
Bump functions with monotone fourier transforms satisfying decay bounds
Tlas, T · 2022
Later among the works it cites.
Never train from scratch: Fair comparison of long-sequence models requires data-driven priors
Amos, I., Berant, J., and Gupta, A · 2023
Closest in time.
On the effectiveness of randomized signatures as reservoir for learning rough dynamics
Compagnoni, E. M., Scampicchio, A., Biggio, L., Orvieto, A., Hofmann, T., and Teichmann, J · 2023
Closest in time.
Formal aspects of language modeling
Cotterell, R., Svete, A., Meister, C., Liu, T., and Du, L · 2023
Closest in time.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Closest in time.
Recurrent distance-encoding neural networks for graph representation learning
Ding, Y., Orvieto, A., He, B., and Hofmann, T · 2023
Closest in time.
Hungry hungry hippos: Towards language modeling with state space models
Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A., and Re, C · 2023
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2023
Closest in time.
Gateloop: Fully data-controlled linear recurrence for sequence modeling, 2023
Katsch, T · 2023
Closest in time.
Structured state space models for in-context reinforcement learning
Lu, C., Schroecker, Y., Gu, A., Parisotto, E., Foerster, J., Singh, S., and Behbahani, F · 2023
Closest in time.
Resurrecting recurrent neural networks for long sequences
Orvieto, A., Smith, S. L., Gu, A., Fernando, A., Gulcehre, C., Pascanu, R., and De, S · 2023
Closest in time.
Rwkv: Reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., et al · 2023
Closest in time.
Simplified state space layers for sequence modeling
Smith, J. T., Warrington, A., and Linderman, S. W · 2023
Closest in time.
Retentive network: A successor to transformer for large language models
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Pretraining without attention, 2023
Wang, J., Yan, J. N., Gu, A., and Rush, A. M · 2023
Closest in time.
Wang, S. and Xue, B · 2023
Closest in time.
Gated linear attention transformers with hardware-efficient training
Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y · 2023
Closest in time.
Theoretical foundations of deep selective state-space models
Cirone, N. M., Orvieto, A., Walker, B., Salvi, C., and Lyons, T · 2024
Closest in time.
Griffin: Mixing gated linear recurrences with local attention for efficient language models, 2024
De, S., Smith, S. L., Fernando, A., Botev, A., Cristian-Muraru, G., Gu, A., Haroun, R., Berrada, L., Chen, Y., Srinivasan, S., Desjardins, G., Doucet, A., Budden, D., Teh, Y. W., Pascanu, R., Freitas, N. D., and Gulcehre, C · 2024
Closest in time.
Repeat after me: Transformers are better than state space models at copying
Jelassi, S., Brandfonbrener, D., Kakade, S. M., and Malach, E · 2024
Closest in time.