Fetching the paper…
Reading the bibliography…
Recent advances in attention-free sequence models rely on convolutions as alternatives to the attention operator at the core of Transformers.
“The curious case of neural text degeneration”, 2019
Ari Holtzman et al · 1904
Earlier work this paper cites.
“Fast transformer decoding: One write-head is all you need”, 2019
Noam Shazeer · 1911
Earlier work this paper cites.
“On the multiplication of successions of Fourier constants”
William Young · 1912
Earlier work this paper cites.
“The approximation of one matrix by another of lower rank”
Carl Eckart and Gale Young · 1936
Earlier work this paper cites.
“On the theory of linear multi-loop feedback systems”
Irwin Sandberg · 1963
Earlier work this paper cites.
“Effective construction of linear state-variable models from input/output functions”
L Ho and Rudolf Kalman · 1966
Earlier work this paper cites.
“Analytic properties of Schmidt pairs for a Hankel operator and the generalized Schur–Takagi problem”
Vadim Adamyan, Damir Arov and Mark’evich Krein · 1971
Earlier work this paper cites.
“Inequalities in Fourier analysis on Rn”
William Beckner · 1975
Earlier work this paper cites.
“Sharpness in Young’s inequality for convolution”
John Fournier · 1977
Earlier work this paper cites.
“A new identification and model reduction algorithm via singular value decomposition”
Sun-Yuan Kung · 1978
Earlier work this paper cites.
“On approximating an FIR filter using discrete orthonormal exponentials”
D Friedman · 1981
Earlier work this paper cites.
“On the approximation of FIR by IIR digital filters”
J Bednar · 1983
Earlier work this paper cites.
“Sharpness of Young’s inequality for convolution”
Tong Quek and Leonard Yap · 1983
Earlier work this paper cites.
“Linear system theory and design”
Chi-Tsong Chen · 1984
Earlier work this paper cites.
“Extensions of Lipschitz mappings into a Hilbert space”
William Johnson · 1984
Earlier work this paper cites.
“Model reduction with balanced realizations: An error bound and a frequency weighted generalization”
Dale. Enns · 1984
Earlier work this paper cites.
“Prefix sums and their applications”
Guy Blelloch · 1990
Earlier work this paper cites.
“Certain models from uncertain data: the algebraic case”
RP Guidorzi · 1991
Earlier work this paper cites.
“Approximation of FIR by IIR digital filters: An algorithm based on balanced model reduction”
Bartlomiej Beliczynski, Izzet Kale and Gerald Cain · 1992
Earlier work this paper cites.
“Essentials of robust control”
Kemin Zhou and John Doyle · 1998
Cited alongside, same era.
“System identification”
Lennart Ljung · 1998
Cited alongside, same era.
“Discrete-time signal processing”
Alan Oppenheim · 1999
Cited alongside, same era.
“Barycentric lagrange interpolation”
Jean-Paul Berrut and Lloyd Trefethen · 2004
Cited alongside, same era.
“Approximation of large-scale dynamical systems”
Athanasios Antoulas · 2005
Cited alongside, same era.
“Model order reduction: theory, research aspects and applications”
Wilhelmus Schilders, Henk Van Vorst and Joost Rommes · 2008
Cited alongside, same era.
“Concentration of measure for the analysis of randomized algorithms”
“Hippo: Recurrent memory with optimal polynomial projections”
Albert Gu et al · 2020
Later among the works it cites.
“Implicit neural representations with periodic activation functions”
Vincent Sitzmann et al · 2020
Later among the works it cites.
“Multiplicative filter networks”
Rizal Fathony, Anit Sahu, Devin Willmott and J Kolter · 2020
Later among the works it cites.
“A data-driven McMillan degree lower bound”
Jeffrey Hokanson · 2020
Later among the works it cites.
“Efficiently modeling long sequences with structured state spaces”, 2021
Albert Gu, Karan Goel and Christopher R\’e · 2021
Later among the works it cites.
“Ckconv: Continuous kernel convolution for sequential data”, 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devdatt Dubhashi and Alessandro Panconesi · 2009
Cited alongside, same era.
“Fair and Square Computation of Inverse Z-Transforms of Rational Functions”
Marcos Moreira and Jo\˜ao Basilio · 2011
Cited alongside, same era.
“Neural machine translation by jointly learning to align and translate”, 2014
Dzmitry Bahdanau, Kyunghyun Cho and Yoshua Bengio · 2014
Cited alongside, same era.
“Estimates for the numerical radius and the spectral radius of the Frobenius companion matrix and bounds for the zeros of polynomials”
Amer Abu-Omar and Fuad Kittaneh · 2014
Cited alongside, same era.
“Application of the AAK theory for sparse approximation of exponential sums”, 2016
Gerlind Plonka and Vlada Pototskaia · 2016
Cited alongside, same era.
“Using fast weights to attend to the recent past”
Jimmy Ba et al · 2016
Cited alongside, same era.
David Romero et al · 2021
Later among the works it cites.
“Flexconv: Continuous kernel convolutions with differentiable kernel sizes”, 2021
David Romero et al · 2021
Later among the works it cites.
“A framework for few-shot language model evaluation”
Leo Gao et al · 2021
Later among the works it cites.
“Diagonal state spaces are as effective as structured state spaces”
Ankit Gupta, Albert Gu and Jonathan Berant · 2022
Later among the works it cites.
“Simplified state space layers for sequence modeling”, 2022
Jimmy Smith, Andrew Warrington and Scott Linderman · 2022
Later among the works it cites.
“Holistic evaluation of language models”, 2022
Percy Liang et al · 2022
Later among the works it cites.
“Constitutional AI: Harmlessness from AI Feedback”, 2022
Yuntao Bai et al · 2022
Later among the works it cites.
“Hungry Hungry Hippos: Towards Language Modeling with State Space Models”
Daniel Fu et al · 2023
Closest in time.
“Hyena Hierarchy: Towards Larger Convolutional Language Models”, 2023
Michael Poli et al · 2023
Closest in time.
“Simple Hardware-Efficient Long Convolutions for Sequence Modeling”
Daniel. Fu et al · 2023
Closest in time.
“Resurrecting Recurrent Neural Networks for Long Sequences”, 2023
Antonio Orvieto et al · 2023
Closest in time.
“GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints”, 2023
Joshua Ainslie et al · 2023
Closest in time.
“Pythia: A suite for analyzing large language models across training and scaling”, 2023
Stella Biderman et al · 2023
Closest in time.
“Effectively Modeling Time Series with Simple Discrete State Spaces”, 2023
Michael Zhang et al · 2023
Closest in time.