Fetching the paper…
Reading the bibliography…
Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module.
“Language Models are Few-shot Learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry and Amanda Askell · 1901
Earlier work this paper cites.
“A New Approach to Linear Filtering and Prediction Problems”, 1960
Rudolph Kalman · 1960
Earlier work this paper cites.
“Prefix Sums and Their Applications”
Guy Blelloch · 1990
Earlier work this paper cites.
“Untersuchungen zu dynamischen neuronalen Netzen”
Sepp Hochreiter · 1991
Earlier work this paper cites.
“Learning to control fast-weight memories: An alternative to dynamic recurrent networks”
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
“Approximation of Dynamical Systems by Continuous Time Recurrent Neural Networks”
Ken-ichi Funahashi and Yuichi Nakamura · 1993
Earlier work this paper cites.
“Long Short-Term Memory”
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
“Gradient Flow in Recurrent Nets: The Difficulty of Learning Long-term Dependencies”
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi and Jürgen Schmidhuber · 2001
Earlier work this paper cites.
“Dynamic Causal Modelling”
Karl Friston, Lee Harrison and Will Penny · 2003
Earlier work this paper cites.
“Random Features for Large-Scale Kernel Machines”
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
“Roofline: An Insightful Visual Performance Model for Multicore Architectures”
Samuel Williams, Andrew Waterman and David Patterson · 2009
Earlier work this paper cites.
“ImageNet Classification with Deep Convolutional Neural Networks”
Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton · 2012
Earlier work this paper cites.
“On the Difficulty of Training Recurrent Neural Networks”
Razvan Pascanu, Tomas Mikolov and Yoshua Bengio · 2013
Earlier work this paper cites.
“Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling”
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho and Yoshua Bengio · 2014
Earlier work this paper cites.
“Sequence to Sequence Learning with Neural Networks”
Ilya Sutskever, Oriol Vinyals and Quoc Le · 2014
Earlier work this paper cites.
“Neural Machine Translation by Jointly Learning to Align and Translate”
Dzmitry Bahdanau, Kyunghyun Cho and Yoshua Bengio · 2015
Earlier work this paper cites.
“Unitary Evolution Recurrent Neural Networks”
Martin Arjovsky, Amar Shah and Yoshua Bengio · 2016
Earlier work this paper cites.
“Using Fast Weights to Attend to the Recent Past”
Jimmy Ba, Geoffrey Hinton, Volodymyr Mnih, Joel Leibo and Catalin Ionescu · 2016
Earlier work this paper cites.
Jimmy Ba, Jamie Kiros and Geoffrey Hinton · 2016
Earlier work this paper cites.
“Strongly-typed Recurrent Neural Networks”
David Balduzzi and Muhammad Ghifary · 2016
Earlier work this paper cites.
“Quasi-recurrent Neural Networks”
James Bradbury, Stephen Merity, Caiming Xiong and Richard Socher · 2016
Earlier work this paper cites.
“Recurrent Orthogonal Networks and Long-Memory Tasks”
Mikael Henaff, Arthur Szlam and Yann LeCun · 2016
Earlier work this paper cites.
“Gaussian Error Linear Units (GELUs)”
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
“WaveNet: A Generative Model for Raw Audio”
Aaron Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
“The LAMBADA Dataset: Word Prediction Requiring a Broad Discourse Context”
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda and Raquel Fernández · 2016
Earlier work this paper cites.
“Efficient Long Sequence Modeling via State Space Augmented Transformer”
Simiao Zuo, Xiaodong Liu, Jian Jiao, Denis Charles, Eren Manavoglu, Tuo Zhao and Jianfeng Gao · 2016
Earlier work this paper cites.
“Language Modeling with Gated Convolutional Networks”
Yann Dauphin, Angela Fan, Michael Auli and David Grangier · 2017
Earlier work this paper cites.
“SampleRNN”
DeepSound · 2017
Earlier work this paper cites.
“HyperNetworks”
David Ha, Andrew Dai and Quoc. Le · 2017
Earlier work this paper cites.
“Simple Recurrent Units for Highly Parallelizable Recurrence”
Tao Lei, Yu Zhang, Sida Wang, Hui Dai and Yoav Artzi · 2017
Earlier work this paper cites.
“SampleRNN: An Unconditional End-to-End Neural Audio Generation Model”
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville and Yoshua Bengio · 2017
Earlier work this paper cites.
“Efficient Orthogonal Parametrisation of Recurrent Neural Networks using Householder Reflections”
Zakaria Mhammedi, Andrew Hellicar, Ashfaqur Rahman and James Bailey · 2017
Earlier work this paper cites.
“Swish: A Self-gated Activation Function”
Prajit Ramachandran, Barret Zoph and Quoc Le · 2017
Earlier work this paper cites.
“Attention Is All You Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan. Gomez, Lukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“On Orthogonality and Learning Recurrent Networks with Long Term Dependencies”
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury and Chris Pal · 2017
Earlier work this paper cites.
“Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge”
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick and Oyvind Tafjord · 2018
Earlier work this paper cites.
“Parallelizing Linear Recurrent Neural Nets Over Sequence Length”
Eric Martin and Chris Cundy · 2018
Earlier work this paper cites.
“Can Recurrent Neural Networks Warp Time?”
Corentin Tallec and Yann Ollivier · 2018
Earlier work this paper cites.
“Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition”
Pete Warden · 2018
Earlier work this paper cites.
“Generating Long Sequences with Sparse Transformers”
Rewon Child, Scott Gray, Alec Radford and Ilya Sutskever · 2019
Cited alongside, same era.
“Adversarial Audio Synthesis”
Chris Donahue, Julian McAuley and Miller Puckette · 2019
Cited alongside, same era.
“Deep Learning for Time Series Classification: A Review”
Hassan Ismail, Germain Forestier, Jonathan Weber, Lhassane Idoumghar and Pierre-Alain Muller · 2019
Cited alongside, same era.
“Gated Orthogonal Recurrent Units: On Learning to Forget”
Li Jing, Caglar Gulcehre, John Peurifoy, Yichen Shen, Max Tegmark, Marin Soljacic and Yoshua Bengio · 2019
Cited alongside, same era.
“Cheap Orthogonal Constraints in Neural Networks: A Simple Parametrization of the Orthogonal and Unitary Group”
Mario Lezcano-Casado and David Martínez-Rubio · 2019
Cited alongside, same era.
“An Empirical Analysis of Compute-Optimal Large Language Model Training”
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las, Lisa Hendricks, Johannes Welbl and Aidan Clark · 2022
Later among the works it cites.
“Transformer Quality in Linear Time”
Weizhe Hua, Zihang Dai, Hanxiao Liu and Quoc Le · 2022
Later among the works it cites.
“S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces”
Eric Nguyen, Karan Goel, Albert Gu, Gordon Downs, Preey Shah, Tri Dao, Stephen Baccus and Christopher Ré · 2022
Later among the works it cites.
“In-context Learning and Induction Heads” https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish and Chris Olah · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brandon Yang, Gabriel Bender, Quoc Le and Jiquan Ngiam · 2019
Cited alongside, same era.
“HellaSwag: Can a Machine Really Finish Your Sentence?”
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi and Yejin Choi · 2019
Cited alongside, same era.
“PIQA: Reasoning about Physical Commonsense in Natural Language”
Yonatan Bisk, Rowan Zellers, Jianfeng Gao and Yejin Choi · 2020
Cited alongside, same era.
“An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold and Sylvain Gelly · 2020
Cited alongside, same era.
“The Pile: An 800GB Dataset of Diverse Text for Language Modeling”
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser and Connor Leahy · 2020
Cited alongside, same era.
“HIPPO: Recurrent Memory with Optimal Polynomial Projections”
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra and Christopher Ré · 2020
Cited alongside, same era.
“Improving the Gating Mechanism of Recurrent Neural Networks”
Albert Gu, Caglar Gulcehre, Tom Paine, Matt Hoffman and Razvan Pascanu · 2020
Cited alongside, same era.
Zhen Qin, Xiaodong Han, Weixuan Sun, Dongxu Li, Lingpeng Kong, Nick Barnes and Yiran Zhong · 2022
Later among the works it cites.
“CosFormer: Rethinking Softmax in Attention”
Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong and Yiran Zhong · 2022
Later among the works it cites.
“Efficient Transformers: A Survey”
Yi Tay, Mostafa Dehghani, Dara Bahri and Donald Metzler · 2022
Later among the works it cites.
“Linear complexity randomized self-attention mechanism”
Lin Zheng, Chong Wang and Lingpeng Kong · 2022
Later among the works it cites.
“Pythia: A Suite for Analyzing Large Language Models across Training and Scaling”
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Khan, Shivanshu Purohit, USVSN Prashanth and Edward Raff · 2023
Closest in time.
“Scaling Transformer to 1M tokens and Beyond with RMT”
Aydar Bulatov, Yuri Kuratov and Mikhail Burtsev · 2023
Closest in time.
“PaLM: Scaling Language Modeling with Pathways”
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Chung, Charles Sutton and Sebastian Gehrmann · 2023
Closest in time.
“Hungry Hungry Hippos: Towards Language Modeling with State Space Models”
Tri Dao, Daniel Fu, Khaled Saab, Armin Thomas, Atri Rudra and Christopher Ré · 2023
Closest in time.
“LongNet: Scaling Transformers to 1,000,000,000 Tokens”
Jiayu Ding, Shuming Ma, Li Dong, Xingxing Zhang, Shaohan Huang, Wenhui Wang and Furu Wei · 2023
Closest in time.
Mahan Fathi, Jonathan Pilault, Pierre-Luc Bacon, Christopher Pal, Orhan Firat and Ross Goroshin · 2023
Closest in time.
“Multi-Head State Space Model for Speech Recognition”
Yassir Fathullah, Chunyang Wu, Yuan Shangguan, Junteng Jia, Wenhan Xiong, Jay Mahadeokar, Chunxi Liu, Yangyang Shi, Ozlem Kalinli, Mike Seltzer and Mark.. Gales · 2023
Closest in time.
“Simple Hardware-efficient Long Convolutions for Sequence Modeling”
Daniel Fu, Elliot Epstein, Eric Nguyen, Armin Thomas, Michael Zhang, Tri Dao, Atri Rudra and Christopher Ré · 2023
Closest in time.
“How to Train Your HIPPO: State Space Models with Generalized Basis Projections”
Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra and Christopher Ré · 2023
Closest in time.
“Liquid Structural State-Space Models”
Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang, Makram Chahine, Alexander Amini and Daniela Rus · 2023
Closest in time.
“Time-Parameterized Convolutional Neural Networks for Irregularly Sampled Time Series”
Chrysoula Kosma, Giannis Nikolentzos and Michalis Vazirgiannis · 2023
Closest in time.
“What Makes Convolutional Models Great on Long Sequence Modeling?”
Yuhong Li, Tianle Cai, Yi Zhang, Deming Chen and Debadeepta Dey · 2023
Closest in time.
“Structured State Space Models for In-Context Reinforcement Learning”
Chris Lu, Yannick Schroecker, Albert Gu, Emilio Parisotto, Jakob Foerster, Satinder Singh and Feryal Behbahani · 2023
Closest in time.
“Focus Your Attention (with Adaptive IIR Filters)”
Shahar Lutati, Itamar Zimerman and Lior Wolf · 2023
Closest in time.
“Mega: Moving Average Equipped Gated Attention”
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May and Luke Zettlemoyer · 2023
Closest in time.
“Long Range Language Modeling via Gated State Spaces”
Harsh Mehta, Ankit Gupta, Ashok Cutkosky and Behnam Neyshabur · 2023
Closest in time.
“HyenaDNA: Long-range Genomic Sequence Modeling at Single Nucleotide Resolution”
Eric Nguyen, Michael Poli, Marjan Faizi, Armin Thomas, Callum Birch-Sykes, Michael Wornow, Aman Patel, Clayton Rabideau, Stefano Massaroli and Yoshua Bengio · 2023
Closest in time.
“Resurrecting Recurrent Neural Networks for Long Sequences”
Antonio Orvieto, Samuel Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu and Soham De · 2023
Closest in time.
“RWKV: Reinventing RNNs for the Transformer Era”
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella and Kranthi GV · 2023
Closest in time.
“Hyena Hierarchy: Towards Larger Convolutional Language Models”
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon and Christopher Ré · 2023
Closest in time.
“Toeplitz Neural Network for Sequence Modeling”
Zhen Qin, Xiaodong Han, Weixuan Sun, Bowen He, Dong Li, Dongxu Li, Yuchao Dai, Lingpeng Kong and Yiran Zhong · 2023
Closest in time.
“Diagonal State Space Augmented Transformers for Speech Recognition”
George Saon, Ankit Gupta and Xiaodong Cui · 2023
Closest in time.
“Large Language Models can be Easily Distracted by Irrelevant Context”
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Chi, Nathanael Schärli and Denny Zhou · 2023
Closest in time.
“Sequence Modeling with Multiresolution Convolutional Memory”
Jiaxin Shi, Ke Wang and Emily Fox · 2023
Closest in time.
“Simplified State Space Layers for Sequence Modeling”
Jimmy Smith, Andrew Warrington and Scott Linderman · 2023
Closest in time.
“Retentive network: A successor to transformer for large language models”
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang and Furu Wei · 2023
Closest in time.
“Llama: Open and Efficient Foundation Language Models”
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro and Faisal Azhar · 2023
Closest in time.
“Selective Structured State-Spaces for Long-form Video Understanding”
Jue Wang, Wentao Zhu, Pichao Wang, Xiang Yu, Linda Liu, Mohamed Omar and Raffay Hamid · 2023
Closest in time.
“Effectively Modeling Time Series with Simple Discrete State Spaces”
Michael Zhang, Khaled Saab, Michael Poli, Tri Dao, Karan Goel and Christopher Ré · 2023
Closest in time.
“FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning”
Tri Dao · 2024
Closest in time.