Fetching the paper…
Reading the bibliography…
Attention-free language models that combine gating and convolutions are growing in popularity due to their efficiency and increasingly competitive performance.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 1904
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 1904
Earlier work this paper cites.
An algorithm for the machine calculation of complex fourier series
James W Cooley and John W Tukey · 1965
Earlier work this paper cites.
Non-holographic associative memory
David J Willshaw, O Peter Buneman, and Hugh Christopher Longuet-Higgins · 1969
Earlier work this paper cites.
Parallel models of associative memory, 1981
JA Feldman, GE Hinton, and JA Anderson · 1981
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John J Hopfield · 1982
Earlier work this paper cites.
An 0 (n log n) sorting network
Miklós Ajtai, János Komlós, and Endre Szemerédi · 1983
Earlier work this paper cites.
Multiplicative complexity, convolution, and the DFT
Michael T Heideman and C Sidney Burrus · 1988
Earlier work this paper cites.
Parallel binary search
Selim G Akl and Henk Meijer · 1990
Earlier work this paper cites.
The analysis of time series: An introduction, fifth edition
Chris Chatfield · 1995
Earlier work this paper cites.
Algebraic Complexity Theory
P. Bürgisser, T. Lickteig, M. Clausen, and A. Shokrollahi · 1996
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2006
Earlier work this paper cites.
The one-way communication complexity of hamming distance
Thathachar S Jayram, Ravi Kumar, and Dandapani Sivakumar · 2008
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Using fast weights to attend to the recent past
Jimmy Ba, Geoffrey E Hinton, Volodymyr Mnih, Joel Z Leibo, and Catalin Ionescu · 2016
Earlier work this paper cites.
The galactic dependencies treebanks: Getting more data by synthesizing new languages
Dingquan Wang and Jason Eisner · 2016
Earlier work this paper cites.
A guide to learning arithmetic circuits
Ilya Volkovich · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Wei Zhang and Bowen Zhou · 2017
Earlier work this paper cites.
Studying the inductive biases of rnns with synthetic variations of natural languages
Shauli Ravfogel, Yoav Goldberg, and Tal Linzen · 2019
Earlier work this paper cites.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin · 2019
Earlier work this paper cites.
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Scott Wen tau Yih, Sinong Wang, and Jie Tang · 2019
Earlier work this paper cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Cited alongside, same era.
Kaleidoscope: An efficient, learnable representation for all structured linear maps
Tri Dao, Nimit S Sohoni, Albert Gu, Matthew Eichhorn, Amit Blonder, Megan Leszczynski, Atri Rudra, and Christopher Ré · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and et al · 2020
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Cited alongside, same era.
Thread: circuits
Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, Ludwig Schubert, Chelsea Voss, Ben Egan, and Swee Kiat Lim · 2020
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Later among the works it cites.
Ckconv: Continuous kernel convolution for sequential data
David W. Romero, Anna Kuzina, Erik J. Bekkers, Jakub M. Tomczak, and Mark Hoogendoorn · 2022
Later among the works it cites.
Diagonal state spaces are as effective as structured state spaces, 2022
Ankit Gupta, Albert Gu, and Jonathan Berant · 2022
Later among the works it cites.
On the parameterization and initialization of diagonal state space models, 2022
Albert Gu, Ankit Gupta, Karan Goel, and Christopher Ré · 2022
Later among the works it cites.
Long range language modeling via gated state spaces, 2022
Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Condconv: Conditionally parametrized convolutions for efficient inference
Brandon Yang, Gabriel Bender, Quoc V. Le, and Jiquan Ngiam · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Ré · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Cited alongside, same era.
Shuangfei Zhai, Walter Talbott, Nitish Srivastava, Chen Huang, Hanlin Goh, Ruixiang Zhang, and Josh Susskind · 2021
Cited alongside, same era.
Examining the inductive bias of neural language models with artificial languages
Jennifer C White and Ryan Cotterell · 2021
Cited alongside, same era.
Thomas H Cormen, Charles E Leiserson, Ronald L Rivest, and Clifford Stein · 2022
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, and Jiaming et al. Kong · 2023
Closest in time.
Focus your attention (with adaptive iir filters), 2023
Shahar Lutati, Itamar Zimerman, and Lior Wolf · 2023
Closest in time.
On the computational complexity of self-attention
Feyza Duman Keles, Pruthuvi Mahesakya Wijewardena, and Chinmay Hegde · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, and Shruti Bhosale · 2023
Closest in time.
Effectively modeling time series with simple discrete state spaces
Michael Zhang, Khaled Saab, Michael Poli, Tri Dao, Karan Goel, and Christopher Ré · 2023
Closest in time.
GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch, 9 2023
Alex Andonian, Quentin Anthony, Stella Biderman, Sid Black, Preetham Gali, Leo Gao, Eric Hallahan, Josh Levy-Kramer, Connor Leahy, Lucas Nestler, Kip Parker, Michael Pieler, Jason Phang, Shivanshu Purohit, Hailey Schoelkopf, Dashiell Stander, Tri Songz, Curt Tigges, Benjamin Thérien, Phil Wang, and Samuel Weinbach · 2023
Closest in time.
Retentive network: A successor to transformer for large language models, 2023
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei · 2023
Closest in time.
Sparse modular activation for efficient sequence modeling
Liliang Ren, Yang Liu, Shuohang Wang, Yichong Xu, Chenguang Zhu, and ChengXiang Zhai · 2023
Closest in time.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Closest in time.
Physics of language models: Part 1, context-free grammar
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Closest in time.
Laughing hyena distillery: Extracting compact recurrences from convolutions
Stefano Massaroli, Michael Poli, Daniel Y Fu, Hermann Kumbong, David Romero, Rom Parnichukun, Aman Timalsina, Quinn McIntyre, Beidi Chen, Atri Rudra, Ce Zhang, Christopher Ré, Stefano Ermon, and Yoshua Bengio · 2023
Closest in time.
Simplified state space layers for sequence modeling, 2023
Jimmy T. H. Smith, Andrew Warrington, and Scott W. Linderman · 2023
Closest in time.
Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution, 2023
Eric Nguyen, Michael Poli, Marjan Faizi, Armin Thomas, Callum Birch-Sykes, Michael Wornow, Aman Patel, Clayton Rabideau, Stefano Massaroli, Yoshua Bengio, Stefano Ermon, Stephen A. Baccus, and Chris Ré · 2023
Closest in time.
URL https://www.forbes.com/sites/robtoews/2023/09/03/transformers-revolutionized-ai-what-will-replace-them/?sh=6ed698269c1f
Transformers revolutionized ai. what will replace them?, 2023 · 2023
Closest in time.
Selective structured state-spaces for long-form video understanding
Jue Wang, Wentao Zhu, Pichao Wang, Xiang Yu, Linda Liu, Mohamed Omar, and Raffay Hamid · 2023
Closest in time.
Time-parameterized convolutional neural networks for irregularly sampled time series
Chrysoula Kosma, Giannis Nikolentzos, and Michalis Vazirgiannis · 2023
Closest in time.
Redpajama: An open source recipe to reproduce llama training dataset, 2023
Together · 2023
Closest in time.