Fetching the paper…
Reading the bibliography…
We present a new layer in which dynamic (i.e.,input-dependent) Infinite Impulse Response (IIR) filters of order two are used to process the input sequence prior to applying conventional attention.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Continual learning with hypernetworks
Johannes Von Oswald, Christian Henning, Benjamin F Grewe, and João Sacramento. 2019 · 1906
Earlier work this paper cites.
A new approach to linear filtering and prediction problems
Rudolph Emil Kalman. 1960 · 1960
Earlier work this paper cites.
Gmat: Global memory augmentation for transformers
Ankit Gupta and Jonathan Berant. 2020 · 2006
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. 2020 · 2006
Earlier work this paper cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. 2020 · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. 2020 · 2011
Earlier work this paper cites.
Learning feed-forward one-shot learners
Luca Bertinetto, João F Henriques, Jack Valmadre, Philip Torr, and Andrea Vedaldi. 2016 · 2016
Earlier work this paper cites.
David Ha, Andrew Dai, and Quoc V Le. 2016 · 2016
Earlier work this paper cites.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Stefan Elfwing, Eiji Uchibe, and Kenji Doya. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Graph hypernetworks for neural architecture search
Chris Zhang, Mengye Ren, and Raquel Urtasun. 2019 · 2019
Earlier work this paper cites.
Principled weight initialization for hypernetworks
Oscar Chang, Lampros Flokas, and Hod Lipson. 2020 · 2020
Earlier work this paper cites.
Differentiable iir filters for machine learning applications
Boris Kuznetsov, Julian D Parker, and Fabián Esqueda. 2020 · 2020
Earlier work this paper cites.
Lightweight and efficient end-to-end speech recognition using low-rank transformer
Genta Indra Winata, Samuel Cahyawijaya, Zhaojiang Lin, Zihan Liu, and Pascale Fung. 2020 · 2020
Cited alongside, same era.
AdaNN: Adaptive neural network-based equalizer via online semi-supervised learning
Qingyi Zhou, Fan Zhang, and Chuanchuan Yang. 2020 · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
N Elhage, N Nanda, C Olsson, T Henighan, N Joseph, B Mann, A Askell, Y Bai, A Chen, T Conerly, et al. 2021 · 2021
Cited alongside, same era.
A practical survey on faster and lighter transformers
Quentin Fournier, Gaétan Marceau Caron, and Daniel Aloise. 2021 · 2021
Cited alongside, same era.
Ckconv: Continuous kernel convolution for sequential data
David W Romero, Anna Kuzina, Erik J Bekkers, Jakub M Tomczak, and Mark Hoogendoorn. 2021 · 2021
Cited alongside, same era.
KalmanNet: Neural network aided kalman filtering for partially known dynamics
Guy Revach, Nir Shlezinger, Xiaoyong Ni, Adria Lopez Escoriza, Ruud J. G. van Sloun, and Yonina C. Eldar. 2022 · 2022
Later among the works it cites.
Hypersound: Generating implicit neural representations of audio signals with hypernetworks
Filip Szatkowski, Karol J Piczak, Przemysław Spurek, Jacek Tabor, and Tomasz Trzciński. 2022 · 2022
Later among the works it cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2022 · 2022
Later among the works it cites.
Junxiong Wang, Jing Nathan Yan, Albert Gu, and Alexander M Rush. 2022 · 2022
Later among the works it cites.
Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting
Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wenhan Xiong, Barlas Oğuz, Anchit Gupta, Xilun Chen, Diana Liskovich, Omer Levy, Wen-tau Yih, and Yashar Mehdad. 2021 · 2021
Cited alongside, same era.
Deep anc: A deep learning approach to active noise control
Hao Zhang and DeLiang Wang. 2021 · 2021
Cited alongside, same era.
Global memory transformer for processing long documents
Arij Al Adel. 2022 · 2022
Cited alongside, same era.
Deep learning for robust adaptive inverse control of nonlinear dynamic systems: Improved settling time with an autoencoder
Nuha A. S. Alwan and Zahir M. Hussain. 2022 · 2022
Cited alongside, same era.
It’s raw! audio generation with state-space models
Karan Goel, Albert Gu, Chris Donahue, and Christopher Ré. 2022 · 2022
Cited alongside, same era.
On the parameterization and initialization of diagonal state space models
Albert Gu, Karan Goel, Ankit Gupta, and Christopher Ré. 2022 · 2022
Cited alongside, same era.
Deep learning-based joint control of acoustic echo cancellation, beamforming and postfiltering
Thomas Haubner and Walter Kellermann. 2022 · 2022
Cited alongside, same era.
Efficient long sequence modeling via state space augmented transformer
Simiao Zuo, Xiaodong Liu, Jian Jiao, Denis Charles, Eren Manavoglu, Tuo Zhao, and Jianfeng Gao. 2022 · 2022
Later among the works it cites.
Scaling transformer to 1m tokens and beyond with rmt
Aydar Bulatov, Yuri Kuratov, and Mikhail S Burtsev. 2023 · 2023
Closest in time.
Decision s4: Efficient sequence-based rl via state spaces layers
Shmuel Bar David, Itamar Zimerman, Eliya Nachmani, and Lior Wolf. 2023 · 2023
Closest in time.
Simple hardware-efficient long convolutions for sequence modeling
Daniel Y Fu, Elliot L Epstein, Eric Nguyen, Armin W Thomas, Michael Zhang, Tri Dao, Atri Rudra, and Christopher Ré. 2023 · 2023
Closest in time.
Efficient long-text understanding with short-text models
Maor Ivgi, Uri Shaham, and Jonathan Berant. 2023 · 2023
Closest in time.
Ocd: Learning to overfit with conditional diffusion models
Shahar Shlomo Lutati and Lior Wolf. 2023 · 2023
Closest in time.
Resurrecting recurrent neural networks for long sequences
Antonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De. 2023 · 2023
Closest in time.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré. 2023 · 2023
Closest in time.
Diagonal state space augmented transformers for speech recognition
George Saon, Ankit Gupta, and Xiaodong Cui. 2023 · 2023
Closest in time.
State spaces aren’t enough: Machine translation needs attention
Ali Vardasbi, Telmo Pires, Robin M. Schmidt, and Stephan Peitz. 2023 · 2023
Closest in time.
Selective structured state-spaces for long-form video understanding
Jue Wang, Wentao Zhu, Pichao Wang, Xiang Yu, Linda Liu, Mohamed Omar, and Raffay Hamid. 2023 · 2023
Closest in time.
Megabyte: Predicting million-byte sequences with multiscale transformers
Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis. 2023 · 2023
Closest in time.
Effectively modeling time series with simple discrete state spaces
Michael Zhang, Khaled K Saab, Michael Poli, Tri Dao, Karan Goel, and Christopher Ré. 2023 · 2023
Closest in time.